Task execution method and device, computer equipment, storage medium and program product
By making the first processing unit of the artificial intelligence chip perform calculation tasks and the second processing unit performs communication tasks in the same task flow, the problem of low execution efficiency of computing tasks and communication tasks on the artificial intelligence chip is solved, and more efficient resource scheduling and shorter task execution time are achieved.
Patent Information
- Application Number
- CN202510774508.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In the prior art, the execution efficiency of computing tasks and communication tasks on artificial intelligence chips still needs to be improved, especially in the interaction between different task flows, there is a problem of low resource scheduling and switching efficiency.
In the same task flow, the computing task is performed by the first processing unit of the artificial intelligence chip, and the communication task is performed by the second processing unit, where the computing task does not depend on the communication task, the communication task depends on the corresponding computing task, and does not rely on the non-corresponding computing task, thereby realizing parallel execution of the task unit.
It improves the execution efficiency of computing and communication tasks on the artificial intelligence chip, reduces time-consuming, improves performance, and reduces the complexity of resource scheduling and handover.
Smart Images

Figure CN120295738A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence chips, and particularly to a task execution method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of artificial intelligence technology, large language models have emerged. Large language models can be used to process various natural language processing tasks, including but not limited to text generation, question answering systems, machine translation, etc. For example, DeepSeek, as a large language model, is currently widely used in scenarios such as intelligent dialogue, text generation, and code writing to help various users improve the processing efficiency of various services.
[0003] The training and inference of large language models require a large amount of computing resources and usually need to be executed on artificial intelligence chips. Among them, artificial intelligence chips are hardware chips specifically designed and optimized for artificial intelligence tasks, including but not limited to GPUs (Graphics Processing Units), NPUs (Neural Network Processing Units), and GPGPUs (General-Purpose computing on Graphics Processing Units).
[0004] When performing related calculations on artificial intelligence chips, the execution of computing tasks and communication tasks will be involved. Among them, computing tasks can be used to perform related operations such as MMA (Matrix Multiply Accumulate), and communication tasks can be used to perform related operations such as data transmission and information exchange such as Reduce.
[0005] For the execution of computing tasks and communication tasks on artificial intelligence chips, current technologies can schedule the execution of computing tasks and communication tasks through the cooperation of multiple different task streams (streams), but it requires the interaction between different task streams, and the processing efficiency of the execution of computing tasks and communication tasks on artificial intelligence chips still needs to be improved. Summary of the Invention
[0006] Based on this, it is necessary to provide a task execution method, apparatus, computer device, computer-readable storage medium, and computer program product for the above technical problems.
[0007] In a first aspect, this application provides a task execution method, including:
[0008] Obtain a task stream;
[0009] In the same task flow, cause a first processing unit of the artificial intelligence chip to execute a computing task, and cause a second processing unit of the artificial intelligence chip to execute a communication task;
[0010] Wherein, the computing task does not depend on the communication task; the communication task depends on the corresponding computing task, and the communication task does not depend on non-corresponding computing tasks.
[0011] In a second aspect, the present application further provides a task execution device, including:
[0012] An acquisition module, configured to acquire a task flow;
[0013] An execution module, configured to, in the same task flow, cause a first processing unit of the artificial intelligence chip to execute a computing task, and cause a second processing unit of the artificial intelligence chip to execute a communication task;
[0014] Wherein, the computing task does not depend on the communication task; the communication task depends on the corresponding computing task, and the communication task does not depend on non-corresponding computing tasks.
[0015] In a third aspect, the present application further provides an artificial intelligence device, including: one or more artificial intelligence chips; the artificial intelligence chips are configured to execute the steps of the above method.
[0016] In a fourth aspect, the present application further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0017] In a fifth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0018] In a sixth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0019] The above task execution method, device, computer device, computer-readable storage medium, and computer program product enable a first processing unit of an artificial intelligence chip to execute a computing task and a second processing unit of the artificial intelligence chip to execute a communication task in the same task flow. Thus, the computing task and the communication task can be executed through scheduling between different hardware units within a single task flow. Among them, the computing task does not depend on the communication task, and the communication task may not depend on a non-corresponding computing task. At this time, the first processing unit and the second processing unit can execute their respective tasks simultaneously without blocking each other. When the communication task depends on the corresponding computing task, the first processing unit can complete the computing task without being blocked, and the communication task can process the communication task after the computing task is completed. Compared with the interaction between different task flows, the resource scheduling and switching efficiency within a single task flow is higher, which improves the processing efficiency of executing the computing task and the communication task on the artificial intelligence chip, and also has shorter time consumption and better performance. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic structural diagram of an artificial intelligence device in an embodiment;
[0022] Figure 2 It is a schematic flowchart of a task execution method in an embodiment;
[0023] Figure 3 It is a schematic flowchart of data processing in an embodiment;
[0024] Figure 4 It is a schematic diagram of time consumption in an embodiment;
[0025] Figure 5 It is a structural block diagram of a task execution device in an embodiment;
[0026] Figure 6 It is an internal structural diagram of a computer device in an embodiment. Detailed Embodiments
[0027] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0028] The task execution method provided by the embodiments of the present application can be applied to, for example, Figure 1 the artificial intelligence device shown. Among them, the artificial intelligence device is a hardware device or system capable of performing artificial intelligence tasks. The artificial intelligence device may include one or more artificial intelligence chips (such as artificial intelligence chip 0, artificial intelligence chip 1,...), and these multiple artificial intelligence chips may be exactly the same, partially the same, or completely different. Among them, the artificial intelligence chip includes, but is not limited to, GPU (Graphics Processing Unit), NPU (Neural Network Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Unit).
[0029] Among them, the task execution method provided by the embodiments of the present application can be executed by the control unit of the artificial intelligence chip in the artificial intelligence device. The artificial intelligence chip may include a first processing unit, a second processing unit, a buffer unit, and a global memory. Among them, the first processing unit may be a processing unit for executing computing tasks, for example, a processing unit for processing operations related to matrix multiplication (such as matrix multiply-accumulate, general matrix multiplication, convolution, etc.) tasks; the second processing unit may be a general processing unit in the artificial intelligence chip, which can be used to process other operations except matrix multiplication, such as addition operation, Softmax operation, etc. The second processing unit can be used to execute communication tasks, implement reduction operators such as allreduce (global reduction) and reduce-scatter (reduction and scatter), and can achieve inter-card communication (between different artificial intelligence chips) through the second processing units on different cards.
[0030] For the execution of computing tasks and communication tasks on the artificial intelligence chip, the current technology can schedule the execution of computing tasks and communication tasks through the cooperation of different multiple task streams (such as the task stream of computing tasks and the task stream of communication tasks), so as to achieve, for example, the parallel execution of the task stream of computing tasks and the task stream of communication tasks. Different processing logics or different kernels of the artificial intelligence chip are running on each task stream, which requires the interaction between the task stream of computing tasks and the task stream of communication tasks. However, the scheduling granularity at the task stream level is relatively large, and the processing efficiency of the execution of computing tasks and communication tasks on the artificial intelligence chip still needs to be improved.
[0031] In this regard, the task execution method provided by the embodiments of the present application enables the first processing unit of the artificial intelligence chip to execute a computing task and enables the second processing unit of the artificial intelligence chip to execute a communication task in the same task flow. Thus, the computing task and the communication task can be executed through scheduling between different hardware units within a single task flow. Among them, the computing task does not depend on the communication task, and the communication task may not depend on the non-corresponding computing task. At this time, the first processing unit and the second processing unit can execute their respective tasks simultaneously without blocking each other. When the communication task depends on the corresponding computing task, the first processing unit can complete the computing task without being blocked, and the communication task can process the communication task after the completion of the computing task. Compared with the interaction between different task flows, the resource scheduling and switching efficiency within a single task flow is higher, which improves the processing efficiency of executing the computing task and the communication task on the artificial intelligence chip, and has shorter time consumption and better performance.
[0032] In some embodiments, a computer device may include an artificial intelligence device as Figure 1 shown. The computer device may be a server. The server may include a host and a device connected to each other. Among them, the host of the server may include a CPU (Central Processing Unit), and the device of the server may include an artificial intelligence device as Figure 1 shown. The artificial intelligence device may include one or more artificial intelligence chips, and the artificial intelligence chips include but are not limited to GPUs and GPGPUs.
[0033] In an exemplary embodiment, as Figure 2 shown, a task execution method is provided. The method may be applied to the artificial intelligence chip of the artificial intelligence device as Figure 1 shown. The method may include:
[0034] Step S201, obtain a task flow.
[0035] Step S202, in the same task flow, enable the first processing unit of the artificial intelligence chip to execute a computing task, and enable the second processing unit of the artificial intelligence chip to execute a communication task. Among them, the computing task does not depend on the communication task, the communication task depends on the corresponding computing task, and the communication task does not depend on the non-corresponding computing task.
[0036] In this embodiment, in combination with Figure 1, the control unit of the artificial intelligence chip can obtain a task flow. In this task flow, the control unit of the artificial intelligence chip causes the first processing unit to execute the calculation task of relevant data, and causes the second processing unit to execute the communication task of relevant data. Among them, the task flow can be a set of ordered task sequences, which can be used to indicate the execution order of the calculation task of the first processing unit and the communication task of the second processing unit. Among them, the calculation task does not depend on the communication task, and the communication task can depend on the corresponding calculation task and does not depend on the non-corresponding calculation task. Among them, the communication task depends on the corresponding calculation task, which means that the communication task needs to be carried out according to the calculation result obtained from the corresponding calculation task; the communication task does not depend on the non-corresponding calculation task, which means that the communication task does not need to be carried out based on the calculation result obtained from the non-corresponding calculation task. Thus, based on the first processing unit and the second processing unit of the artificial intelligence chip, the task switching and scheduling of the first processing unit and the second processing unit can be realized in a single task flow. When the communication task does not depend on the non-corresponding calculation task, the first processing unit and the second processing unit can execute their respective tasks simultaneously, that is, the calculation task of the first processing unit and the communication task of the second processing unit can be executed simultaneously without blocking each other. When the communication task depends on the corresponding calculation task, the first processing unit can complete the calculation task without being blocked, and the second processing unit can execute the corresponding communication task after the calculation task of the first processing unit is completed. Compared with the interaction between different task flows in the current technology, the resource scheduling and switching efficiency within a single task flow in this embodiment is higher, the time consumption is shorter, and the performance is better.
[0037] The task execution method of this embodiment causes the first processing unit of the artificial intelligence chip to execute the calculation task and the second processing unit of the artificial intelligence chip to execute the communication task in the same task flow, so that the calculation task and the communication task can be executed through the scheduling between different hardware units within a single task flow. Among them, the calculation task does not depend on the communication task, and the communication task can not depend on the non-corresponding calculation task. At this time, the first processing unit and the second processing unit can execute their respective tasks simultaneously without blocking each other. When the communication task depends on the corresponding calculation task, the first processing unit can complete the calculation task without being blocked, and the communication task can process the communication task after the completion of the calculation task. Compared with the interaction between different task flows, the resource scheduling and switching efficiency within a single task flow is higher, which improves the processing efficiency of executing the calculation task and the communication task on the artificial intelligence chip, and the time consumption is shorter and the performance is better.
[0038] In an exemplary embodiment, in the same task flow of step S202, causing the first processing unit of the artificial intelligence chip to execute the calculation task and causing the second processing unit of the artificial intelligence chip to execute the communication task may include:
[0039] In the same task flow, the first processing unit is made to execute a computing task on data to obtain a computing result of the data, and the second processing unit is made to execute a communication task on the computing result.
[0040] In this embodiment, in combination with Figure 1 , the control unit of the artificial intelligence chip can obtain input data, which can be tensor data or data chunks obtained by partitioning tensor data. This data can be transmitted by the control unit to the first processing unit for processing. In the same task flow, the control unit can cause the first processing unit to execute a computing task on this data to obtain a computing result of this data. For example, operations such as matrix multiply-accumulate, general matrix multiplication, and convolution can be performed on this data to obtain the computing result. The control unit of the artificial intelligence chip can also cause the second processing unit to obtain the computing result of this data obtained by the first processing unit and execute a communication task on the computing result of this data. For example, operations such as all-reduce and reduce-scatter can be performed on the computing result of this data.
[0041] The solution of this embodiment can, based on the first processing unit and the second processing unit of the artificial intelligence chip, implement the computing and communication operations of the first processing unit and the second processing unit on data in a single task flow. Compared with the interaction method between different task flows, the data processing efficiency is improved.
[0042] In an exemplary embodiment, further, the above-mentioned making the first processing unit execute a computing task on data to obtain a computing result of the data, and making the second processing unit execute a communication task on the computing result in the same task flow may include:
[0043] In the same task flow, the first processing unit is made to sequentially execute a computing task on each data chunk of the data to obtain a computing result of each data chunk, and the second processing unit is made to sequentially execute a communication task on the computing result of each data chunk obtained each time.
[0044] In this embodiment, in the same task flow, the control unit of the artificial intelligence chip can cause the first processing unit to sequentially execute the calculation tasks of each data block of the (tensor) data, so that the first processing unit can sequentially obtain the calculation results of each data block. For each calculation result of each data block sequentially obtained by the first processing unit, in the same task flow, the control unit can cause the second processing unit to sequentially execute the communication tasks of the calculation results of the data blocks obtained by the first processing unit each time. As an example, the first processing unit can sequentially execute the calculation tasks on data block 0 and data block 1, and sequentially obtain the calculation results of data block 0 and data block 1. Among them, after the first processing unit obtains the calculation result of data block 0, the first processing unit can continue to execute the calculation task of data block 1. At the same time, after the first processing unit obtains the calculation result of data block 0, the second processing unit can obtain the calculation result of this data block 0, and then the second processing unit can execute the communication task of the calculation result of this data block 0. That is to say, in a single task flow, while the first processing unit executes the calculation task on a data block, the second processing unit can execute the communication task on the calculation result of another data block. It can be understood that there will be corresponding time-consuming for the calculation and communication tasks performed on the data. The time-consuming for the calculation task performed on the data can be recorded as the calculation task time-consuming, and the time-consuming for the communication task performed on the data can be recorded as the communication task time-consuming. The solution of this embodiment can achieve partial parallelism of the calculation and communication tasks performed on the data at the single task flow level, so that part of the calculation task time-consuming and part of the communication task time-consuming cover each other, thereby improving its data processing efficiency. In the above example, when the first processing unit continues to execute the calculation task of data block 1, the second processing unit can execute the communication task of the calculation result of data block 0 in parallel, so that the time-consuming of the calculation task of data block 1 and the time-consuming of the communication task of the calculation result of data block 0 cover each other, improving the data processing efficiency. Thus, it can also be explained that the calculation tasks of data block 0 and data block 1 do not depend on the communication tasks, the communication task of the calculation result of data block 0 depends on the calculation task of data block 0 (i.e., the corresponding calculation task), the communication task of the calculation result of data block 1 depends on the calculation task of data block 1 (i.e., the corresponding calculation task), and the communication task of the calculation result of data block 0 does not depend on the calculation task of data block 1 (i.e., the non-corresponding calculation task), and the communication task of the calculation result of data block 1 does not depend on the calculation task of data block 0 (i.e., the non-corresponding calculation task).
[0045] In an exemplary embodiment, further, in the same task flow in the above method, causing the first processing unit to execute the calculation task of the data to obtain the calculation result of the data, and causing the second processing unit to execute the communication task of the calculation result may further include the following steps:
[0046] In the same task flow, enable the first processing unit to write the calculation result of each data block into the buffer unit of the artificial intelligence chip, and enable the second processing unit to read the calculation result of the data block from the buffer unit.
[0047] In this embodiment, in combination with Figure 1 , the artificial intelligence chip may include a buffer unit (Cache / Buffer), and this buffer unit may be a tensor buffer unit. In combination with Figure 3 , in the same task flow, the control unit of the artificial intelligence chip may enable the first processing unit to sequentially write the calculation result of each data block of (tensor) data into the buffer unit of the artificial intelligence chip, that is, each time the calculation task for data block i is executed, the calculation result of data block i is written into the buffer unit until all data blocks are processed, that is, i is greater than the total number of blocks, and the total number of blocks is the total number of data blocks of the data. Since the communication task of the calculation result of data block i in the second processing unit depends on the calculation task of data block i in the first processing unit, each time the first processing unit writes the calculation result of data block i into the buffer unit, the second processing unit reads the calculation result of this data block i from the buffer unit and performs the communication task on the calculation result of this data block i until the calculation results of all data blocks are processed, that is, i is greater than the total number of blocks.
[0048] In an exemplary embodiment, further, in the same task flow in the above method, enabling the first processing unit to execute the calculation task of the data to obtain the calculation result of the data, and enabling the second processing unit to execute the communication task of the calculation result may further include the following steps:
[0049] In the same task flow, enable the first processing unit to block the data to obtain each data block.
[0050] In this embodiment, in the same task flow, the control unit of the artificial intelligence chip may enable the first processing unit to sequentially execute the calculation task of each data block of the data to obtain the calculation result of each data block, and the control unit may enable the second processing unit to sequentially execute the communication task of the calculation result of each data block obtained by the first processing unit each time. Moreover, before the first processing unit sequentially executes the calculation task of each data block of the data, the control unit may further enable the first processing unit to block the data to obtain each data block, and the first processing unit may block the data according to a preset blocking method to obtain each data block, and this preset blocking method may instruct the first processing unit to block the data according to a preset granularity to obtain each data block. The solution of this embodiment can perform the blocking of the data inside a single task flow, without splitting the (tensor) data in the upper-layer software, and can save the changes in the framework layer.
[0051] In an exemplary embodiment, further, the above method may further include the following steps:
[0052] Obtain the correspondence between the total time consumption and various data chunking methods; wherein, the total time consumption includes the unit time consumption of the computing task, the unit time consumption of the communication task, and the time consumption of task overlap; according to the correspondence, determine the target chunking method; wherein, the target chunking method is the data chunking method that makes the total time consumption meet the target condition; the target chunking method is used for the first processing unit to chunk the data.
[0053] For the selection of different chunking methods (different granularities), the time consumptions of the computing task and the communication task will vary. In this embodiment, the target chunking method for the first processing unit to chunk the data can be determined according to the correspondence between the total time consumption and various data chunking methods, so that the control unit of the artificial intelligence chip can make the first processing unit chunk the (tensor) data according to the target chunking method in the same task flow to obtain each data chunk. Wherein, the target chunking method is the data chunking method that makes the total time consumption meet the target condition, and the target condition can be a condition related to the magnitude of the total time consumption, such as the total time consumption being less than a certain total time consumption threshold, the total time consumption being the smallest, etc. Thus, through the analysis and application of the target chunking method, the time consumption of data processing can be reduced.
[0054] Specifically, the total time consumption may include the unit time consumption of the computing task, the unit time consumption of the communication task, and the time consumption of task overlap. Among them, the unit time consumption of the computing task refers to the time consumption of executing the computing task of a data chunk, and the unit time consumption of the communication task refers to the time consumption of executing the communication task of the computing result of a data chunk. For the time consumption of task overlap, as described in the previous embodiment, there will be corresponding time consumptions for the computing and communication tasks executed on the data. It is possible to achieve partial parallelism between the computing and communication tasks executed on the data at the single task flow level, so that the time consumption of part of the computing task and the time consumption of part of the communication task cover each other. For example, when the first processing unit continues to execute the computing task of data chunk 1, the second processing unit can concurrently execute the communication task of the computing result of data chunk 0, so that the time consumption of the computing task of data chunk 1 and the time consumption of the communication task of the computing result of data chunk 0 cover each other. From this, it can be understood that the time consumption of task overlap refers to the time consumption of the overlapping part between the computing tasks of some data chunks and the communication tasks of the computing results of some data chunks.
[0055] For the total time consumption, as an example, such as Figure 4As shown, taking the computing tasks of four data chunks as an example for illustration, the computing tasks of the four data chunks are denoted as Computation 1 to 4, and the communication tasks of the computing results of the four data chunks are denoted as Communication 1 to 4. The total time consumption can be divided into case (a) and case (b). In case (a), the computing time consumption for a single data chunk (the unit computing time consumption of the computing task) can be greater than or equal to the communication time consumption for a single data chunk (the unit communication time consumption of the communication task), that is, Computation 1 can be greater than or equal to Communication 1, and so on. In case (b), the computing time consumption for a single data chunk is less than the communication time consumption for a single data chunk, that is, Computation 1 is less than Communication 1, and so on. In this regard, in both case (a) and case (b), the task overlapping time consumption is the time consumption during the parallel execution of Computation 2 to 4 and Communication 1 to 3. This task overlapping time consumption is determined by the larger one among the time consumptions of Computation 2 to 4 and Communication 1 to 3 during the parallel period. In case (a), the task overlapping time consumption takes the time consumption of Computation 2 to 4, and in case (b), the task overlapping time consumption takes the time consumption of Communication 1 to 3.
[0056] Based on this, as an example, assume that the size of the input matrix A (data) is [M, K], and the size of the input matrix B (data) is [K, N]. The granularity of the computing tasks for each dimension M, N, and K of the input (tensor) data chunks can be denoted as TM, TN, and TK. The size of the computing result obtained by the computing task for each data chunk is [TM, TN], and this computing result is used for the communication task.
[0057] Thus, the corresponding relationship between the total time consumption and various data chunking methods can be expressed as:
[0058] ;
[0059] Among them, T represents the total time consumption, t1 represents the unit computing time consumption of the computing task, t2 represents the unit communication time consumption of the communication task, and ceiling represents rounding up.
[0060] Specifically, as an example, the unit computing time consumption t1 of the computing task can be further expressed as:
[0061] ,
[0062] ;
[0063] Among them, load can represent the time taken for the first processing unit to read input data from the global memory. Calc can represent the time taken for the first processing unit to perform relevant operations for the computing task, such as the time taken for matrix multiplication operations. Store can represent the time taken for the first processing unit to write the calculated results of data chunks into the buffer unit. Load_bandwidth can represent the bandwidth for reading data. HW_POWER can represent the computing power of the first processing unit. Store_bandwidth can represent the bandwidth for writing out data. Dsize can represent the data size.
[0064] Specifically, as an example, the unit time consumption t2 of the communication task can be further expressed as:
[0065] ;
[0066] Among them, bandwidth can represent the bandwidth of inter-card communication.
[0067] It should be noted that the expressions for the unit time consumption of the computing task and the unit time consumption of the communication task above are only examples. In specific applications, corresponding expressions can be used for calculation according to different implementation methods and algorithms of the actual kernel.
[0068] For the same hardware, the computing power of the processing unit, the read / write bandwidth within the card, and the bandwidth of inter-card communication are usually determined. The solution of this embodiment can select the target chunking method that meets the target conditions according to the above corresponding relationship, and use the appropriate target feedback method for the first processing unit to perform chunking processing on the data, improving the data processing efficiency, and making the performance of the fusion of computing and communication tasks optimal.
[0069] In an exemplary embodiment, the first processing unit includes a tensor processing unit; the second processing unit includes a vector processing unit; the task flow indicates the execution order of the computing task of the tensor processing unit and the communication task of the vector processing unit.
[0070] In this embodiment, the artificial intelligence device may include one or more artificial intelligence chips such as GPUs. The artificial intelligence chip may include a tensor processing unit (Tensor Core) and a vector processing unit (Vector Core). Among them, the tensor processing unit can be a computing core with dedicated acceleration at the hardware level, and can be used to process computing-intensive matrix multiplication-related operations (such as matrix multiply-accumulate MMA, general matrix multiplication GEMM, convolution operation, etc.). The vector processing unit can be a general processing unit of an artificial intelligence chip such as a GPU, and can be used to process other operations except matrix multiplication, such as Add, Softmax, etc.
[0071] The solution of this embodiment can be based on the tensor processing unit and vector processing unit included in the artificial intelligence chip to achieve the switching and scheduling of the tensor processing unit and vector processing unit in a single task flow. When there is no data dependency, the tensor processing unit and vector processing unit can execute their respective tasks simultaneously without blocking each other. When there is a data dependency between the two, the tensor processing unit can complete the calculation tasks in sequence without being blocked, and the vector processing unit can perform the corresponding communication tasks after the calculation tasks of the tensor processing unit are completed. Thus, the parallelism of calculation and communication can be achieved through the scheduling and switching between different hardware units, and the execution of all the above tasks can be carried out within a single task flow, without the need to split the tensors in the upper-layer software, and the modification of the framework layer can be omitted. Moreover, compared with the interaction between different task flows, the resource scheduling and switching efficiency within a single task flow is higher, which improves the processing efficiency of the execution of calculation tasks and communication tasks on the artificial intelligence chip, and has shorter time consumption and better performance.
[0072] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.
[0073] Based on the same inventive concept, the embodiments of the present application also provide a task execution device for implementing the above-mentioned task execution method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the task execution device provided below can refer to the limitations on the task execution method in the above text, and will not be repeated here.
[0074] In an exemplary embodiment, as Figure 5 shown, a task execution device is provided. The device 500 may include:
[0075] An acquisition module 501, configured to acquire a task flow;
[0076] An execution module 502, configured to cause a first processing unit of an artificial intelligence chip to execute a computing task and cause a second processing unit of the artificial intelligence chip to execute a communication task in the same task flow;
[0077] Wherein, the computing task does not depend on the communication task; the communication task depends on the corresponding computing task, and the communication task does not depend on the non-corresponding computing task.
[0078] In an exemplary embodiment, the execution module 502 is configured to cause the first processing unit to execute a computing task on data to obtain a computing result of the data and cause the second processing unit to execute a communication task on the computing result in the same task flow.
[0079] In an exemplary embodiment, the execution module 502 is configured to cause the first processing unit to sequentially execute a computing task on each data block of the data to obtain a computing result of each data block and cause the second processing unit to sequentially execute a communication task on the computing result of each obtained data block in the same task flow.
[0080] In an exemplary embodiment, the execution module 502 is further configured to cause the first processing unit to write the computing result of each data block into a buffer unit of the artificial intelligence chip and cause the second processing unit to read the computing result of the data block obtained from the buffer unit in the same task flow.
[0081] In an exemplary embodiment, the execution module 502 is further configured to cause the first processing unit to partition the data to obtain each data block in the same task flow.
[0082] In an exemplary embodiment, the execution module 502 is further configured to obtain a correspondence between the total time consumption and various data partitioning methods; the total time consumption includes the unit time consumption of the computing task, the unit time consumption of the communication task, and the task overlapping time consumption; determine a target partitioning method according to the correspondence; wherein, the target partitioning method is a data partitioning method that enables the total time consumption to meet a target condition; the target partitioning method is used for the first processing unit to partition the data.
[0083] In an exemplary embodiment, the first processing unit includes a tensor processing unit; the second processing unit includes a vector processing unit; the task flow indicates an execution order of the computing task of the tensor processing unit and the communication task of the vector processing unit.
[0084] Each module in the above task execution device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0085] In an exemplary embodiment, an artificial intelligence device is provided. The artificial intelligence device may include one or more artificial intelligence chips, and the artificial intelligence chips can be used to execute the steps in the embodiments of the above task execution methods.
[0086] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 6 shown. The computer device may include a processor, a memory, an input / output interface, and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external devices in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a task execution method.
[0087] Those skilled in the art can understand that Figure 6 the structure shown in
[0088] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0089] In an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0090] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the foregoing method embodiments.
[0091] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0092] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0093] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0094] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A task execution method, characterized in that, The method includes: Obtaining a task flow; In the same task flow, causing a first processing unit in the artificial intelligence chip to execute a computing task, and causing a second processing unit in the artificial intelligence chip to execute a communication task; Wherein, the computing task does not depend on the communication task; the communication task depends on the corresponding computing task, and the communication task does not depend on the non-corresponding computing task; wherein, the communication task depends on the corresponding computing task, which means that the communication task is based on the computing result obtained from the corresponding computing task; the communication task does not depend on the non-corresponding computing task, which means that the communication task does not need to be based on the computing result obtained from the non-corresponding computing task; when the communication task does not depend on the non-corresponding computing task, the first processing unit and the second processing unit execute their respective tasks simultaneously; when the communication task depends on the corresponding computing task, the second processing unit executes the corresponding communication task after the computing task of the first processing unit is completed.
2. The method according to claim 1, wherein The step of, in the same task flow, causing a first processing unit in the artificial intelligence chip to execute a computing task, and causing a second processing unit in the artificial intelligence chip to execute a communication task, includes: In the same task flow, causing the first processing unit to execute a computing task on the data to obtain a computing result of the data, and causing the second processing unit to execute a communication task on the computing result.
3. The method according to claim 2, wherein The step of, in the same task flow, causing the first processing unit to execute a computing task on the data to obtain a computing result of the data, and causing the second processing unit to execute a communication task on the computing result, includes: In the same task flow, causing the first processing unit to sequentially execute a computing task on each data block of the data to obtain a computing result of each data block, and causing the second processing unit to sequentially execute a communication task on the computing result of each obtained data block.
4. The method according to claim 3, wherein The step of, in the same task flow, causing the first processing unit to execute a computing task on the data to obtain a computing result of the data, and causing the second processing unit to execute a communication task on the computing result, further includes: In the same task flow, causing the first processing unit to write the computing result of each data block into a buffer unit of the artificial intelligence chip, and causing the second processing unit to read the computing result of the data block from the buffer unit.
5. The method according to claim 3, characterized in that, The step of, in the same task flow, causing the first processing unit to execute a computing task on the data to obtain a computing result of the data, and causing the second processing unit to execute a communication task on the computing result, further includes: In the same task flow, causing the first processing unit to partition the data to obtain each data block.
6. The method according to claim 5, characterized in that It further includes: Obtaining the correspondence between the total time consumption and various data partitioning methods; The total time consumption includes the unit time consumption of the computing task, the unit time consumption of the communication task, and the time consumption of task overlap; Determining a target partitioning method according to the correspondence. Among them, the target chunking method is the data chunking method that enables the total time consumption to meet the target condition; the target chunking method is used by the first processing unit to chunk the data.
7. The method according to any one of claims 1 to 6, characterized in that, The first processing unit includes a tensor processing unit; the second processing unit includes a vector processing unit; the task flow indicates the execution order of the computing task of the tensor processing unit and the communication task of the vector processing unit.
8. An artificial intelligence device, characterized in that, The device includes: one or more artificial intelligence chips; the artificial intelligence chips are used to execute the steps of the method according to any one of claims 1 to 7.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Communication task processing method, task caching device and storage medium
CN111026539A
Large language model training method and device, equipment and storage medium
CN117057411A
Inference method and device, equipment and medium
CN119443258A
Parallel processing method and device
CN120011299A
Distributed execution method and device of large language model, medium and distributed cluster
CN120104318A
Cited By
Task allocation method and device for AI chip, chip, equipment and storage medium
CN120469819A
Task allocation method and device of AI chip, chip, equipment and storage medium
CN120469819B
Universal graphics processor, computing method, computing device, medium, and program product
CN121120361A
Artificial intelligence chip
CN121683907A
Electronic equipment and data processing method
CN122111693A