Model task execution method and device, electronic equipment and readable medium
By splitting the tasks of the big model into subtasks and assigning target task execution devices, the problems of insufficient video memory and high cost during the runtime of the big model are solved, and the effects of heterogeneous computing and cost reduction are achieved.
Patent Information
- Application Number
- CN202411877221.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
AI Technical Summary
The big model relies on the CUDA platform. The graphics memory of a single graphics card is not enough to support the operation of the big model, and the cost of independent graphics cards is getting higher and higher.
The task to be executed is split into multiple sub-tasks to be executed, and each sub-task is assigned a target task execution device, and the sub-tasks are executed using the chips of these devices.
Through heterogeneous computing and task splitting, the computing cost of large models is reduced, hardware utilization is improved, and the problem of insufficient graphics memory of a single graphics card is solved.
Smart Images

Figure CN120011005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model technology, and in particular to a model task execution method, a model task execution device, an electronic device and a computer-readable medium. Background Art
[0002] With the emergence of large models (LLMs), the effectiveness of the scale law has been verified. Compared with traditional models, large models have achieved significant performance improvements and are increasingly used in fields such as human-computer interaction, text understanding, knowledge extraction, and multimodal applications. However, as the performance of large models improves, the scale of large model calculations becomes larger and larger, the cost of calculations becomes higher and higher, and the hardware limitations of calculations become more and more stringent.
[0003] In the related technologies, most general deep learning reasoning frameworks or large models rely on the CUDA (Compute Unified Device Architecture) platform. These frameworks usually use independent graphics cards to increase the computing speed. As the model size increases, the required video memory space also increases, and the video memory of a single graphics card may not be enough to support the operation of large models. Moreover, as the price of graphics cards that can run sufficiently increases, the cost of using independent graphics cards to serve large models is also increasing. Summary of the invention
[0004] The embodiments of the present invention provide a model task execution method, device, electronic device and computer-readable storage medium to solve the problem that most large models rely on the CUDA platform, and these frameworks usually use independent graphics cards to improve the calculation speed; as the model scale increases, the required video memory space also increases, and the video memory of a single graphics card may not be enough to support the operation of the large model; and as the price of graphics cards that can run sufficiently increases, the cost of using independent graphics cards to provide services for large models is also increasing.
[0005] The embodiment of the present invention discloses a task execution method of a model, which is applied to a model, wherein the model includes at least one task execution device, and the method includes:
[0006] Split the task to be executed into at least one subtask to be executed;
[0007] Allocating a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip;
[0008] The subtask to be executed is executed using the chip of the target task execution device.
[0009] Optionally, the assigning a corresponding target task execution device to the subtask to be executed includes:
[0010] By configuring a preset software development tool for the task execution device, obtaining at least one of a network address, an architecture, and a target operator of the task execution device; the target operator is an operator that can be run on the task execution device;
[0011] Based on the network address of the task execution device, the architecture, and at least one of the target operator, the target task execution device is allocated to the subtask to be executed.
[0012] Optionally, the target task execution device includes a first target task execution device; the subtask to be executed includes a first subtask to be executed; and the using the chip of the target task execution device to execute the subtask to be executed includes:
[0013] Using the first target task execution device to obtain first to-be-processed data related to the first to-be-executed subtask;
[0014] Based on the first data to be processed, the first target task execution device is used to execute the first subtask to be executed.
[0015] Optionally, the target task execution device includes a second target task execution device; the subtask to be executed includes a second subtask to be executed; if the first data to be processed includes output data of the second subtask to be executed, then when the second target task execution device executes the second subtask to be executed, the first target task execution device is in a waiting state.
[0016] Optionally, the target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device; and the using the chip of the target task execution device to execute the sub-task to be executed includes:
[0017] Controlling the host to send second to-be-processed data related to the to-be-executed sub-task to the task execution sub-device;
[0018] Based on the second data to be processed, the subtask to be executed is executed by using the chip of the task execution sub-device.
[0019] Optionally, the method comprises:
[0020] Using a preset delay calculation formula, obtain the delay cost of the model executing the task to be executed;
[0021] The delay calculation formula is:
[0022]
[0023] Among them, D total is the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is a second transmission time of the second to-be-processed data between the host and the task execution sub-device of the model, The time that the target task execution device in the model is in the waiting state.
[0024] Optionally, the chip includes at least one of a unified computing device architecture, a graphics processing unit, a tensor processing unit, and a neural network processing unit.
[0025] The embodiment of the present invention further discloses a task execution device for a model, which is applied to the model. The model includes at least one task execution device. The device includes:
[0026] A splitting module, used to split the task to be executed into at least one subtask to be executed;
[0027] An allocation module, used for allocating a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip;
[0028] An execution module is used to execute the subtask to be executed by using the chip of the target task execution device.
[0029] Optionally, the allocation module includes:
[0030] A configuration submodule, configured to obtain at least one of a network address, an architecture, and a target operator of the task execution device by configuring a preset software development tool for the task execution device; the target operator is an operator that can be run on the task execution device;
[0031] The allocation submodule is used to allocate the target task execution device to the subtask to be executed based on the network address of the task execution device, the architecture and at least one of the target operators.
[0032] Optionally, the target task execution device includes a first target task execution device; the subtask to be executed includes a first subtask to be executed; and the execution module includes:
[0033] An acquisition submodule, configured to acquire first to-be-processed data related to the first to-be-executed subtask by using the first target task execution device;
[0034] The first execution submodule is used to execute the first to-be-executed subtask based on the first to-be-processed data using the first target task execution device.
[0035] Optionally, the target task execution device includes a second target task execution device; the subtask to be executed includes a second subtask to be executed; if the first data to be processed includes output data of the second subtask to be executed, then when the second target task execution device executes the second subtask to be executed, the first target task execution device is in a waiting state.
[0036] Optionally, the target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device; and the execution module includes:
[0037] A control submodule, used for controlling the host to send the second to-be-processed data related to the to-be-executed subtask to the task execution sub-device;
[0038] The second execution submodule is used to execute the subtask to be executed based on the second data to be processed by using the chip of the task execution sub-device.
[0039] Optionally, the device comprises:
[0040] A delay cost acquisition module, used to obtain the delay cost of the model executing the task to be executed by using a preset delay calculation formula;
[0041] The delay calculation formula is:
[0042]
[0043] Among them, D total is the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is a second transmission time of the second to-be-processed data between the host and the task execution sub-device of the model, The time that the target task execution device in the model is in the waiting state.
[0044] Optionally, the chip includes at least one of a unified computing device architecture, a graphics processing unit, a tensor processing unit, and a neural network processing unit.
[0045] The embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0046] The memory is used to store computer programs;
[0047] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.
[0048] The embodiment of the present invention further discloses one or more computer-readable media on which instructions are stored. When executed by one or more processors, the processors are enabled to execute the method described in the embodiment of the present invention.
[0049] The embodiments of the present invention include the following advantages:
[0050] In an embodiment of the present invention, the model includes at least one task execution device, splits the task to be executed into at least one subtask to be executed; assigns a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip; and the subtask to be executed is executed using the chip of the target task execution device. A distributed large model service solution based on heterogeneous computing is proposed, which can use existing heterogeneous devices such as GPUs with lower performance and capacity for acceleration, and reduce the inference cost of large models and other large-scale deep learning models. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of a task execution method of a model provided in an embodiment of the present invention;
[0052] Figure 2 is a schematic diagram of a task execution device provided in an embodiment of the present invention;
[0053] Figure 3 is a schematic diagram of registration management of a task execution device provided in an embodiment of the present invention;
[0054] Figure 4 It is a schematic diagram of using multiple target task execution devices to execute tasks to be executed provided in an embodiment of the present invention;
[0055] Figure 5 It is a schematic diagram of sending an execution output result of a to-be-executed subtask to another target task execution device provided in an embodiment of the present invention;
[0056] Figure 6 is a schematic diagram of heterogeneous computing of a block provided in an embodiment of the present invention;
[0057] Figure 7 is a structural block diagram of a task execution device of a model provided in an embodiment of the present invention;
[0058] Figure 8 is a block diagram of an electronic device provided in an embodiment of the present invention;
[0059] Fig. 9 is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] To facilitate understanding of the technical solutions and technical effects of the embodiments of the present invention, the prior art of the present invention is briefly described below.
[0062] In the related art, most general deep learning reasoning frameworks or large models rely on the CUDA (Compute Unified Device Architecture) platform, and these frameworks usually use independent graphics cards to increase the computing speed. In the related art, quantization, cropping, distillation and other means can be used to optimize and compress large models to reduce the demand for computing resources for large models; however, these technical means of optimizing and compressing large models may cause the loss of accuracy of large models and the problem of error accumulation. In addition, as the scale of the model increases, the required video memory space is also increasing, and the video memory of a single graphics card may not be enough to support the operation of large models. Moreover, as the price of graphics cards that can run sufficiently increases, the cost of using independent graphics cards to provide services for large models is also increasing.
[0063] Reference Figure 1 , shows a flowchart of a task execution method of a model provided in an embodiment of the present invention, which is applied to a model, wherein the model includes at least one task execution device, and specifically may include the following steps:
[0064] Step 101, splitting the task to be executed into at least one subtask to be executed;
[0065] In an embodiment of the present invention, the model may be a large model, and the model is used to execute the task to be executed. The subtask to be executed may be a task to be calculated. The model may be a neural network model or a deep learning model. The model includes at least one layer, which is a layer of a neural network and a computing unit in deep learning. The model includes at least one task execution device.
[0066] In an embodiment of the present invention, the task to be executed by the model can be split into at least one subtask to be executed. And there may be a dependency relationship between the execution of the subtasks to be executed, that is, the data involved in a subtask to be executed is the output data of another subtask to be executed. Therefore, the execution process of the task to be executed, that is, the execution process of multiple subtasks to be executed can be abstracted as a directed acyclic graph. Each subtask to be executed can be executed by a task execution device. At this time, the task execution device used to execute each subtask to be executed can be abstracted as a computing block, that is, a block. The directed acyclic graph includes nodes and edges, the nodes represent computing blocks, and the edges represent the dependencies between the subtasks to be executed.
[0067] It should be noted that a block can be a layer of a model or a combination of multiple layers of a model. A block is a unidirectional node with only one (group) input and one (group) output.
[0068] Step 102, assigning a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip;
[0069] In an embodiment of the present invention, a corresponding target task execution device may be allocated to a subtask to be executed from at least one task execution device. The target task execution device includes a chip. The chip may be a computing chip, and the chip is used to execute the subtask to be executed.
[0070] In some embodiments of the present invention, the chip includes at least one of a unified computing device architecture, a graphics processing unit, a tensor processing unit, and a neural network processing unit.
[0071] In an embodiment of the present invention, the server may be configured with different types of heterogeneous computing hardware such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), NPU (Neural Processing Unit), TPU (Tensor Processing Unit), etc. Among them, heterogeneous computing refers to the use of different types of processing units such as CPU, GPU, NPU, TPU, etc. in a computing system to collaboratively complete computing tasks. Various computing underlying suites such as OpenCL (Open Computing Language), OpenGL (Open Graphics Library), and Vulkan all have the ability to support GPU and other heterogeneous computing. Various NPUs and TPUs also provide SDKs (Software Development Kits) for heterogeneous computing. These SDKs enable developers to use these dedicated hardware for heterogeneous computing. Therefore, the chip in the embodiment of the present invention includes at least one of a unified computing device architecture (CUDA), a graphics processing unit, a tensor processing unit, and a neural network processing unit.
[0072] Reference Figure 2 , shows a schematic diagram of a task execution device provided in an embodiment of the present invention. The task execution device includes a host and a task execution sub-device.
[0073] In the embodiment of the present invention, heterogeneous computing means that the computing process is not entirely on a single chip architecture. In the embodiment of the present invention, host refers to the CPU (central processing unit) end, and device refers to computing chips such as GPU, TPU, NPU, etc. Device is used to perform specific tasks.
[0074] In some embodiments of the present invention, the step of allocating a corresponding target task execution device to the subtask to be executed includes:
[0075] By configuring a preset software development tool for the task execution device, obtaining at least one of a network address, an architecture, and a target operator of the task execution device; the target operator is an operator that can be run on the task execution device;
[0076] Based on the network address of the task execution device, the architecture, and at least one of the target operator, the target task execution device is allocated to the subtask to be executed.
[0077] In the embodiment of the present invention, the task center of the server can be used to provide services to the outside world as a whole, and its functions include registration and management of task execution devices, registration and management of tasks to be executed, and operation splitting of tasks to be executed.
[0078] Reference Figure 3 , shows a schematic diagram of registration management of a task execution device provided in an embodiment of the present invention. The task center can execute tasks to be executed, and can register and manage task execution devices. The task center can split the tasks to be executed into multiple subtasks to be executed, including task1, task2 and task3.
[0079] In an embodiment of the present invention, the task center provides an SDK (Software Development Kit), and the task execution device needs to install and run this SDK. The task execution device that installs and runs the SDK will collect its own operation information and send this information back to the task center through the gateway. This process is usually called "registration", which means that the device is registered in the task center so that the task center can understand and manage these devices. The operation information of the device includes the IP (Internet Protocol) address of the device, the architecture of the computing operation, the supported operators, etc. Among them, the IP of the device is the network address of the device, which is used for communication between the task center and the task execution device. The architecture of the computing operation refers to the architecture of the hardware of the task execution device. Operators generally refer to operations that can be performed in the task to be executed. The supported operator list tells the task center what types of operations the task execution device can perform. Based on the network address, architecture and at least one of the target operators of the task execution device, a target task execution device can be assigned to the subtask to be executed.
[0080] Step 103: Utilize the chip of the target task execution device to execute the subtask to be executed.
[0081] In the embodiment of the present invention, the subtask to be executed may be executed by using the chip of the target task execution device.
[0082] In some embodiments of the present invention, the target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device; and the using the chip of the target task execution device to execute the sub-task to be executed includes:
[0083] Controlling the host to send second to-be-processed data related to the to-be-executed sub-task to the task execution sub-device;
[0084] Based on the second data to be processed, the subtask to be executed is executed by using the chip of the task execution sub-device.
[0085] In an embodiment of the present invention, the target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device. In an embodiment of the present invention, when the chip of the target task execution device is used to execute the sub-task to be executed, it is necessary to control the host to send the second to-be-processed data related to the sub-task to be executed to the task execution sub-device. Then, based on the second to-be-processed data, the chip of the task execution sub-device is used to execute the sub-task to be executed.
[0086] In an embodiment of the present invention, the chips of multiple target task execution devices can be used to execute corresponding subtasks to be executed, thereby realizing the execution of tasks to be executed. This process is heterogeneous computing. It should be noted that a heterogeneous computing process may have multiple parallel branches, which can be calculated in parallel on different task execution devices or on a single device. When computing resources are limited, different branches are usually placed on different devices.
[0087] Reference Figure 4 , shows a schematic diagram of using multiple target task execution devices to execute tasks to be executed provided in an embodiment of the present invention. Figure 4 In the heterogeneous computing process, there are two parallel branches, including a total of 6 target task execution devices.
[0088] In some embodiments of the present invention, the target task execution device includes a first target task execution device; the subtask to be executed includes a first subtask to be executed; and the using the chip of the target task execution device to execute the subtask to be executed includes:
[0089] Using the first target task execution device to obtain first to-be-processed data related to the first to-be-executed subtask;
[0090] Based on the first data to be processed, the first target task execution device is used to execute the first subtask to be executed.
[0091] In an embodiment of the present invention, the target task execution device includes a first target task execution device, and the subtask to be executed includes a first subtask to be executed. When the first subtask to be executed is executed by the first target task execution device, the first target task execution device is used to obtain first to-be-processed data related to the first subtask to be executed, and then based on the first to-be-processed data, the first subtask to be executed is executed by the first target task execution device.
[0092] In some embodiments of the present invention, the target task execution device includes a second target task execution device; the subtask to be executed includes a second subtask to be executed; if the first data to be processed includes output data of the second subtask to be executed, then when the second target task execution device executes the second subtask to be executed, the first target task execution device is in a waiting state.
[0093] In the embodiment of the present invention, the execution of a subtask to be executed may depend on the output results of other tasks in addition to the input, output and execution steps of the task. Therefore, the execution of a task may also include the calculation of the placeholder conditions that need to be waited for. For this reason, the execution output of the subtask to be executed may not be sent directly back to the task center, but may also be sent to another target task execution device waiting for the execution result. Figure 5 , showing a schematic diagram of sending the execution output result of a to-be-executed subtask to another target task execution device provided in an embodiment of the present invention. Figure 5 In the process, the execution output result of the subtask task2 to be executed is sent to the target task execution device corresponding to the subtask task1 to be executed.
[0094] In some embodiments of the invention, the method comprises:
[0095] Using a preset delay calculation formula, obtain the delay cost of the model executing the task to be executed;
[0096] The delay calculation formula is:
[0097]
[0098] Among them, D total is the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is a second transmission time of the second to-be-processed data between the host and the task execution sub-device of the model, The time that the target task execution device in the model is in the waiting state.
[0099] In an embodiment of the present invention, a directed acyclic graph of a task to be executed can be split into multiple forms. Since multiple subtasks to be executed need to transmit data between different target task execution devices and between the host and the task execution sub-devices when executing, the execution of a subtask to be executed may also need to wait for the execution of other subtasks to be executed. Therefore, it is necessary to consider the delay cost of the task to be executed to determine the most reasonable splitting method to ensure operation efficiency.
[0100] In the embodiment of the present invention, the delay cost of the model executing the task to be executed can be calculated using the preset delay calculation formula. The delay calculation formula is:
[0101]
[0102] Among them, D total For the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is the second transmission time of the second to-be-processed data between the host computer and the task execution sub-device of the model, The time that the device is in a waiting state for the target task execution in the model.
[0103] It should be noted that, for any target task execution device, it may need to wait for multiple other target task execution devices to complete their tasks. The time that the target task execution device is in the waiting state is the longest time required to wait for multiple other target task execution devices to complete their tasks.
[0104] In the embodiment of the present invention, the exact delay cost is usually affected by the runtime and cannot be accurately quantified. In order to quantify the calculation delay under a certain splitting condition, it is necessary to preset some cost values based on a series of experiments, as shown in Table 1.
[0105] Table 1: Examples of latency for different types, sizes, and computational blocks
[0106]
[0107] Among them, M stands for megabytes, K stands for kilobytes, and ms stands for milliseconds.
[0108] In the embodiment of the present invention, based on the preset cost, if the current scale is not in the table, interpolation processing is required. Searching for the best split is actually a typical dynamic programming problem. Modern large models have many layers and operators, and the topological space to search is large. In order to reduce the search space, the model can be divided into larger blocks first, and the structure of the model can be simplified into a simple topological graph.
[0109] In the embodiment of the present invention, single-machine heterogeneous computing includes computing implementation on different heterogeneous devices and heterogeneous computing of blocks.
[0110] The computation implementation on different heterogeneous devices is actually the operator corresponding to the time under the current computing conditions. Currently, the operators with N card CUDA as the mainstream have relatively common implementations. For integrated graphics cards and non-N card independent graphics cards, other computation implementations need to be selected. One solution is to use OpenCL, and the other solution is to use the compute shader of graphics APIs (application programming interfaces) such as OpenGL, OpenGL, and Vulkan to implement it. This requires a lot of work.
[0111] Heterogeneous computing of blocks means that if a host has multiple devices available within a computing node, further parallelization can be achieved.
[0112] Reference Figure 6 , shows a schematic diagram of heterogeneous computing of a block provided in an embodiment of the present invention. The computing block, i.e., the task execution device, includes a host and multiple task execution sub-devices. Multiple task execution sub-tasks can be executed in parallel, and multiple sub-tasks to be executed can also be occupied and waited.
[0113] In a specific example, the steps of executing the task to be executed may include: first, the task execution device integrates the SDK and registers the device information with the task center. Then, the topology of the model is entered. Then, based on the existing hardware information and computing scale, the best way to split the task to be executed is searched. Then, the task to be executed is split into different subtasks to be executed according to the splitting strategy and sent to different task execution devices. Finally, the calculation results are collected and the response is returned.
[0114] In a specific example, a self-developed model and an open source model were heterogeneously transformed and distributed heterogeneous deployment was successfully implemented. The single model of this heterogeneous deployment was run on a computing node with a dual-card 2080Ti N card and a computing node with an Intel integrated graphics card. The model was split into two parts and run on two computing nodes respectively. The operator of the integrated graphics card is implemented using the compute shader of Vulkan. The specific process is: the delay cost is calculated through the table of the pre-experiment, and the dual device can be attributed to the greedy algorithm, that is, as little network transmission as possible and as much calculation on the N card as possible will result in less final delay. It is concluded that starting from the input, the first N-2 basic blocks are placed in the N card, and the last two blocks containing transformers and the output layer are placed in the integrated graphics card. The network transmission throughput is 8192*4 bytes; the transformer and other operators are implemented on the integrated graphics card using the computer shader of Vulkan to verify that the input accuracy is within the error range. Deploy the service, register with the registration center, and generate the task id. The running results of the implementation calculation and the comparison with A100 are shown in Table 2. Considering the market price and supply of A100 and 2080ti+ integrated graphics cards, a single computing node saves about 80% of the hardware cost.
[0115] Table 2: Example of the results of the implementation calculations compared to the latency of the A100
[0116]
[0117] In an embodiment of the present invention, the model includes at least one task execution device, splits the task to be executed into at least one subtask to be executed; assigns a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip; and the subtask to be executed is executed using the chip of the target task execution device. A distributed large model service solution based on heterogeneous computing is proposed, which can use existing heterogeneous devices such as GPUs with lower performance and capacity for acceleration, and reduce the inference cost of large models and other large-scale deep learning models.
[0118] In the embodiment of the present invention, an optimal heterogeneous splitting evaluation method is proposed, which solves the problem that computing devices with low performance and capacity cannot be used in this scenario to a certain extent, and reduces service costs.
[0119] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0120] Reference Figure 7 , shows a structural block diagram of a task execution device of a model provided in an embodiment of the present invention, which is applied to a model, wherein the model includes at least one task execution device, and specifically may include the following modules:
[0121] A splitting module 701 is used to split a task to be executed into at least one subtask to be executed;
[0122] The allocation module 702 is used to allocate a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip;
[0123] The execution module 703 is used to execute the subtask to be executed by using the chip of the target task execution device.
[0124] In an optional embodiment of the present invention, the allocation module includes:
[0125] A configuration submodule, configured to obtain at least one of a network address, an architecture, and a target operator of the task execution device by configuring a preset software development tool for the task execution device; the target operator is an operator that can be run on the task execution device;
[0126] The allocation submodule is used to allocate the target task execution device to the subtask to be executed based on the network address of the task execution device, the architecture and at least one of the target operators.
[0127] In an optional embodiment of the present invention, the target task execution device includes a first target task execution device; the subtask to be executed includes a first subtask to be executed; and the execution module includes:
[0128] An acquisition submodule, configured to acquire first to-be-processed data related to the first to-be-executed subtask by using the first target task execution device;
[0129] The first execution submodule is used to execute the first to-be-executed subtask based on the first to-be-processed data using the first target task execution device.
[0130] In an optional embodiment of the present invention, the target task execution device includes a second target task execution device; the subtask to be executed includes a second subtask to be executed; if the first data to be processed includes the output data of the second subtask to be executed, then when the second target task execution device executes the second subtask to be executed, the first target task execution device is in a waiting state.
[0131] In an optional embodiment of the present invention, the target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device; the execution module includes:
[0132] A control submodule, used for controlling the host to send the second to-be-processed data related to the to-be-executed subtask to the task execution sub-device;
[0133] The second execution submodule is used to execute the subtask to be executed based on the second data to be processed by using the chip of the task execution sub-device.
[0134] In an optional embodiment of the present invention, the device comprises:
[0135] A delay cost acquisition module, used to obtain the delay cost of the model executing the task to be executed by using a preset delay calculation formula;
[0136] The delay calculation formula is:
[0137]
[0138] Among them, D total is the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is a second transmission time of the second to-be-processed data between the host and the task execution sub-device of the model, The time that the target task execution device in the model is in the waiting state.
[0139] In an optional embodiment of the present invention, the chip includes at least one of a unified computing device architecture, a graphics processing unit, a tensor processing unit, and a neural network processing unit.
[0140] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0141] In addition, an embodiment of the present invention further provides an electronic device, such as Figure 8As shown, it includes a processor 801, a communication interface 802, a memory 803 and a communication bus 804, wherein the processor 801, the communication interface 802, and the memory 803 communicate with each other through the communication bus 804.
[0142] Memory 803, used for storing computer programs;
[0143] The processor 801 is used to execute the program stored in the memory 803, and implements the following steps:
[0144] Split the task to be executed into at least one subtask to be executed;
[0145] Allocating a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip;
[0146] The subtask to be executed is executed using the chip of the target task execution device.
[0147] In an optional embodiment of the present invention, the step of allocating a corresponding target task execution device to the subtask to be executed includes:
[0148] By configuring a preset software development tool for the task execution device, obtaining at least one of a network address, an architecture, and a target operator of the task execution device; the target operator is an operator that can be run on the task execution device;
[0149] Based on the network address of the task execution device, the architecture, and at least one of the target operator, the target task execution device is allocated to the subtask to be executed.
[0150] In an optional embodiment of the present invention, the target task execution device includes a first target task execution device; the subtask to be executed includes a first subtask to be executed; and the using the chip of the target task execution device to execute the subtask to be executed includes:
[0151] Using the first target task execution device to obtain first to-be-processed data related to the first to-be-executed subtask;
[0152] Based on the first data to be processed, the first target task execution device is used to execute the first subtask to be executed.
[0153] In an optional embodiment of the present invention, the target task execution device includes a second target task execution device; the subtask to be executed includes a second subtask to be executed; if the first data to be processed includes the output data of the second subtask to be executed, then when the second target task execution device executes the second subtask to be executed, the first target task execution device is in a waiting state.
[0154] In an optional embodiment of the present invention, the target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device; and the using the chip of the target task execution device to execute the sub-task to be executed includes:
[0155] Controlling the host to send second to-be-processed data related to the to-be-executed sub-task to the task execution sub-device;
[0156] Based on the second data to be processed, the subtask to be executed is executed by using the chip of the task execution sub-device.
[0157] In an optional embodiment of the present invention, the method includes:
[0158] Using a preset delay calculation formula, obtain the delay cost of the model executing the task to be executed;
[0159] The delay calculation formula is:
[0160]
[0161] Among them, D total is the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is a second transmission time of the second to-be-processed data between the host and the task execution sub-device of the model, The time that the target task execution device in the model is in the waiting state.
[0162] In an optional embodiment of the present invention, the chip includes at least one of a unified computing device architecture, a graphics processing unit, a tensor processing unit, and a neural network processing unit.
[0163] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0164] The communication interface is used for communication between the above terminal and other devices.
[0165] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0166] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0167] like Fig. 9 As shown, in another embodiment provided by the present invention, a computer-readable storage medium 901 is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes a task execution method of a model described in the above embodiment.
[0168] In another embodiment provided by the present invention, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute a task execution method of a model described in the above embodiment.
[0169] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.
[0170] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0171] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0172] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A task execution method of a model, characterized in that: Applied to a model, the model comprising at least one task execution device, the method comprising: Split the task to be executed into at least one subtask to be executed; Allocating a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip; The subtask to be executed is executed using the chip of the target task execution device.
2. The method according to claim 1, characterized in that The step of allocating a corresponding target task execution device to the subtask to be executed includes: By configuring a preset software development tool for the task execution device, obtaining at least one of a network address, an architecture, and a target operator of the task execution device; the target operator is an operator that can be run on the task execution device; Based on the network address of the task execution device, the architecture, and at least one of the target operator, the target task execution device is allocated to the subtask to be executed.
3. The method according to claim 2, characterized in that The target task execution device includes a first target task execution device; the subtask to be executed includes a first subtask to be executed; and the using the chip of the target task execution device to execute the subtask to be executed includes: Using the first target task execution device to obtain first to-be-processed data related to the first to-be-executed subtask; Based on the first data to be processed, the first target task execution device is used to execute the first subtask to be executed.
4. The method according to claim 3, characterized in that The target task execution device includes a second target task execution device; the subtask to be executed includes a second subtask to be executed; if the first data to be processed includes the output data of the second subtask to be executed, then when the second target task execution device executes the second subtask to be executed, the first target task execution device is in a waiting state.
5. The method according to claim 4, characterized in that The target task execution device includes a host and a task execution sub-device; the chip is located in the task execution sub-device; and the using the chip of the target task execution device to execute the sub-task to be executed includes: Controlling the host to send second to-be-processed data related to the to-be-executed sub-task to the task execution sub-device; Based on the second data to be processed, the subtask to be executed is executed by using the chip of the task execution sub-device.
6. The method according to claim 5, characterized in that The method comprises: Using a preset delay calculation formula, obtain the delay cost of the model executing the task to be executed; The delay calculation formula is: Among them, D total is the delay cost, is the first transmission time of the data related to the task to be executed between the target task execution devices in the model, is a second transmission time of the second to-be-processed data between the host and the task execution sub-device of the model, The time that the target task execution device in the model is in the waiting state.
7. The method according to claim 1, characterized in that The chip includes at least one of a unified computing device architecture, a graphics processing unit, a tensor processing unit, and a neural network processing unit.
8. A task execution device of a model, characterized in that: Applied to a model, the model comprising at least one task execution device, the apparatus comprising: A splitting module, used to split the task to be executed into at least one subtask to be executed; An allocation module, used for allocating a corresponding target task execution device to the subtask to be executed; the target task execution device includes a chip; An execution module is used to execute the subtask to be executed by using the chip of the target task execution device.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 7 when executing the program stored in the memory.
10. One or more computer-readable media having instructions stored thereon, which when executed by one or more processors cause the processors to perform the method according to any one of claims 1 to 7.