Task execution method and device, electronic equipment and storage medium
By dynamically determining the encoding mode based on storage space and data volume and generating task execution strategies, the existing compilers have solved the problem of low hardware resource utilization, and efficient computing resource allocation and task execution are achieved to adapt to diversified business needs.
Patent Information
- Application Number
- CN202510396130.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Existing artificial intelligence compilers cannot fully utilize hardware resources in dealing with complex computing scenarios, resulting in low computing resource utilization, and are prone to resource allocation imbalance and data handling bottlenecks in hybrid computing scenarios and dynamic scale scenarios.
By determining the target encoding mode based on the storage space capacity and data volume of the data to be processed, a task execution strategy is generated that is suitable for different hardware architectures, the number and parameters of processing modules are dynamically adjusted, and the allocation of computing resources is optimized.
It improves computing efficiency and performance stability, reduces user configuration complexity, realizes load balancing, supports the adaptation of diverse business scenarios and hardware architectures, and improves the throughput and real-timeness of computing-intensive tasks.
Smart Images

Figure CN120276857A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and more particularly to the fields of chip technologies and artificial intelligence compiler technologies. More specifically, the present disclosure provides a task execution method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence technologies, the applications of artificial intelligence processors are constantly increasing. The tasks executed by artificial intelligence processors can be compiled by an artificial intelligence compiler. Summary of the Invention
[0003] The present disclosure provides a task execution method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a task execution method, the method comprising: determining a target encoding mode according to the storage space capacity for the data to be processed and the data volume of the data to be processed, wherein the target encoding mode is used to indicate at least one of the number and parameters of processing modules for processing the data to be processed, and the data to be processed is obtained from the initial data of an initial task; generating a task to be executed according to the target encoding mode and the initial task; and executing the task to be executed by using the processing modules indicated by the target encoding mode.
[0005] According to another aspect of the present disclosure, there is provided a task execution apparatus, the apparatus comprising: a determination module, configured to determine a target encoding mode according to the storage space capacity for the data to be processed and the data volume of the data to be processed, wherein the target encoding mode is used to indicate at least one of the number and parameters of processing modules for processing the data to be processed, and the data to be processed is obtained from the initial data of an initial task; a generation module, configured to generate a task to be executed according to the target encoding mode and the initial task; and an execution module, configured to execute the task to be executed by using the processing modules indicated by the target encoding mode.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, the computer instructions being used to cause a computer to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a computer program product, comprising a computer program which, when executed by a processor, implements the method provided by the present disclosure.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Brief Description of the Drawings
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a flowchart of a task execution method according to an embodiment of the present disclosure;
[0012] Figure 2 is a schematic flowchart of a task execution method according to an embodiment of the present disclosure;
[0013] Figure 3 is a block diagram of a task execution device according to an embodiment of the present disclosure; and
[0014] Figure 4 is a block diagram of an electronic device that can apply the task execution method according to an embodiment of the present disclosure. Detailed Embodiments
[0015] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0016] The artificial intelligence processor can be various processors such as a general-purpose graphics processing unit (GPGPU), a neural network processing unit (NPU), and a tensor processing unit. Taking the general-purpose graphics processing unit as an example, the Triton compiler is a high-performance artificial intelligence compiler for general-purpose graphics processing units. The Triton compiler can divide the computing task into a block structure that can be executed in parallel through a tiling strategy to maximize the use of the computing resources of the graphics processing unit. When generating the tiling strategy, the Triton compiler mainly relies on developers to manually specify the tiling parameters or adopt a fixed-rule tiling scheme. Thus, the Triton compiler mainly adopts a tiling method based on artificial experience, which has certain limitations and cannot fully utilize the hardware resources of the artificial intelligence processor in complex computing scenarios.
[0017] For large-scale operators, it is rather cumbersome to configure the chunking parameters for the Triton compiler. In the case where a deep learning model involves high-dimensional tensor operations (such as three-dimensional convolution and attention mechanisms), developers can manually specify parameters such as the chunk size (Tile Size) and stride for each dimension of the tensor. If a deep learning model contains dozens of heterogeneous operators, developers need to write independent chunking rules for each operator, and the parameter configuration workload increases exponentially and is error-prone.
[0018] The static chunking strategy of the Triton compiler has poor adaptability to hardware. Graphics processors with different architectures have different numbers of stream processors and different shared memory capacities. The static chunking strategy or the above fixed chunking rules cannot dynamically adapt to different graphics processor architectures, resulting in low utilization of computing resources. On a graphics processor with limited video memory bandwidth, an overly large chunk size will cause a data transfer bottleneck, while an overly small chunk cannot fully utilize the matrix computing power of the tensor core.
[0019] In a hybrid computing scenario, the Triton compiler is likely to encounter chunking conflicts. In the case where a deep learning model includes dense computing operators and sparse computing operators, the common chunking strategies of the Triton compiler may generate resource fragmentation due to conflicts in storage access patterns. The dense computing operator can be a general matrix multiplication (GEMM) operator, which requires continuous memory access. The sparse computing operator can be a scatter operator or a gather operator, which requires random access. It is difficult to co-optimize the chunking strategies of dense computing operators and sparse computing operators.
[0020] In a dynamic shape scenario, the chunking strategy of the Triton compiler is likely to fail. The dynamic shape scenario can be a variable sequence length scenario, a variable resolution input scenario, etc. Fixed chunking rules are difficult to adjust the chunking scheme according to the actual tensor shape at runtime, resulting in an unbalanced allocation of computing resources. When the actual tensor shape cannot be evenly divided by the preset chunking parameters, a large number of invalid computing threads will be generated.
[0021] The Triton compiler is difficult to achieve multi-level storage coordination. The Triton compiler adopts different and independent chunking strategies for global memory units, shared memory units, and registers, lacking data transfer coordination across storage levels. When writing the chunked data into the shared memory unit, the Triton compiler does not consider the subsequent data reuse pattern at the register level, resulting in repeated data transfer.
[0022] Therefore, in order to improve the utilization rate of hardware resources of artificial intelligence processors, the present disclosure provides a task execution method, which will be described below.
[0023] Figure 1 It is a flowchart of a task execution method according to an embodiment of the present disclosure.
[0024] As Figure 1 shown, the method 100 may include operation S110 to operation S130.
[0025] In operation S110, a target encoding mode is determined according to the storage space capacity for the data to be processed and the data volume of the data to be processed.
[0026] In an embodiment of the present disclosure, the task to be processed may be obtained from the initial data of the initial task. The initial task may be represented as: a computation node in a high-level intermediate representation (High Level Intermediate Representation) obtained based on the developer's code. For example, the high-level intermediate representation may be a computation graph. The computation graph may include multiple computation nodes. Each computation node may correspond to an initial task. In one example, the initial task may be a row-sum task for an initial matrix. The initial matrix may serve as the initial data. The data to be processed may be a row of data of the initial matrix.
[0027] In an embodiment of the present disclosure, the storage space for the data to be processed may be a buffer storage area for the data to be processed, and may be a storage space in a multi-level cache unit. The encoding mode may be determined according to the difference between the storage space capacity for the data to be processed and the data volume of the data to be processed. For example, the storage space for the data to be processed may be a storage space in a level-3 cache unit (L3 Cache).
[0028] In an embodiment of the present disclosure, the target encoding mode is used to indicate at least one of the number and parameters of the processing modules for processing the data to be processed. The processing module may be a processor core or a thread. The parameters of the processing module may include the frequency of the processing module.
[0029] In operation S120, a task to be executed is generated according to the target encoding mode and the initial task.
[0030] For example, if the data volume of the data to be processed is consistent with the storage space capacity for the data to be processed, the target encoding mode determined according to the data volume of the data to be processed and the storage space capacity for the data to be processed may indicate that the number of processor cores for processing one row of data is one. The target encoding mode may also indicate that the processor cores operate at the default frequency. Thus, the task to be executed may indicate that one processor core processes multiple data to be processed in a loop at the default frequency. The initial task may be replaced with the task to be executed.
[0031] In operation S130, the task to be executed is executed using the processing modules indicated by the target encoding mode.
[0032] For example, after replacing the initial task in the initial computation graph with the task to be executed, an updated computation graph can be obtained. The updated computation graph can be converted into low-level computation instructions and provided to the artificial intelligence processor. Next, one or more computation instructions for the task to be executed can be executed by a processor core to cyclically process multiple pieces of data to be processed at a default frequency.
[0033] Through the embodiments of the present disclosure, an encoding mode is determined according to the storage space fusion for the data to be processed and the data volume of the data to be processed, which can improve the computing efficiency and performance stability, maximize the single-core utilization rate on the premise of avoiding memory overflow, and achieve load balancing. Compared with the fixed rule chunking strategy, the resource overhead of inter-core synchronization can be reduced, and the throughput and real-time performance of compute-intensive tasks can be improved.
[0034] In addition, through the embodiments of the present disclosure, determining the encoding mode according to the storage space fusion for the data to be processed and the data volume of the data to be processed can cover the automatic optimization of most scenarios, reduce the user configuration complexity and operation and maintenance costs. Developers do not need to manually adjust the underlying parallel parameters, realizing the design of "automation first, customization on demand", greatly reducing the threshold for tuning the chunking strategy of the artificial intelligence processor, shortening the model deployment cycle, and at the same time ensuring the accurate implementation of the chunking strategy in the compilation optimization stage, avoiding performance loss or resource waste caused by manual misconfiguration.
[0035] In addition, through the embodiments of the present disclosure, according to the initial task and the target encoding mode, the task to be executed is determined, and the task to be executed is executed by using the processing module, which can realize the adjustment of the high-level intermediate representation and the low-level computation instructions, and can be flexibly extended to a new hardware architecture (such as a memory-computation integrated chip) or a mixed-precision computing mode, and can support high-scalability service scenarios, providing underlying technical support for adapting to the upgrade of artificial intelligence computing power and diverse service requirements.
[0036] It can be understood that the method of the present disclosure is described above, and the method of the present disclosure will be further described below.
[0037] Figure 2 is a schematic flowchart of a task execution method according to an embodiment of the present disclosure.
[0038] As Figure 2 shown, before performing the above operation S110, operations S201 to S203 can be performed.
[0039] In operation S201, a high-level intermediate representation for the artificial intelligence processor is obtained.
[0040] For example, an artificial intelligence processor can be used to implement the inference of a deep learning model. In this case, the high-level intermediate representation can be a computation graph. The computation graph can include multiple computation nodes. Each computation node corresponds to an initial task.
[0041] In operation S202, obtain the data volume of the data to be processed.
[0042] For example, the data to be processed can be a row of data of an initial matrix of an initial task. The data to be processed can include multiple elements in a row of the initial matrix. In one example, the initial matrix can be a 128×128 matrix, and the scale of the data to be processed can be 1×128.
[0043] In operation S203, obtain the storage space capacity for the data to be processed.
[0044] For example, the storage space capacity allocated for the data to be processed can be obtained. It can be understood that the storage space allocated for a computation node can be used as the storage space for the data to be processed. It can be understood that operation S202 and operation S203 can be executed in parallel, but the present disclosure is not limited thereto. Operation S202 can be executed first and then operation S203. Alternatively, operation S203 can be executed first and then operation S202.
[0045] Next, in some embodiments of the above operation S110, a target encoding mode can be determined from multiple candidate encoding modes according to the storage space capacity for the data to be processed and the data volume of the data to be processed. The multiple candidate encoding modes can include a first candidate encoding mode, a second candidate encoding mode, and a third candidate encoding mode.
[0046] In some embodiments, the first candidate encoding mode can indicate that multiple sub-data to be processed of the data to be processed are respectively processed by multiple processing modules. For example, the multiple processing modules can be 4 processor cores. As described above, the scale of the data to be processed can be 1×128, and the sub-data to be processed can be obtained by splitting the data to be processed. The number of sub-data to be processed can be the same as the number of processing modules. The scale of the sub-data to be processed can be 1×32.
[0047] In some embodiments, the second candidate encoding mode can indicate that the data to be processed is processed by a processing module. For example, the data to be processed can be processed by one processor core.
[0048] In some embodiments, the third candidate encoding mode can indicate that multiple data to be processed are processed by a processing module. For example, one processor core can process multiple data to be processed.
[0049] Next, it will be described in conjunction with operations S211 to S212.
[0050] In operation S211, a comparison relationship between the storage space capacity for the data to be processed and the data volume of the data to be processed is determined. The comparison relationship can be one of a first comparison relationship, a second comparison relationship, and a third comparison relationship. The first comparison relationship can indicate that the target difference between the storage space capacity for the data to be processed and the data volume of the data to be processed is greater than a preset difference threshold and the storage space capacity for the data to be processed is less than the data volume of the data to be processed. The second comparison relationship can indicate that the target difference is less than or equal to the preset difference threshold. The third comparison relationship can indicate that the target difference is greater than the preset difference threshold and the storage space capacity for the data to be processed is greater than the data volume of the data to be processed. The target difference can be the absolute value of the target difference.
[0051] In some embodiments, in response to determining that the target difference between the storage space capacity for the data to be processed and the data volume of the data to be processed is greater than the preset difference threshold and the storage space capacity for the data to be processed is less than the data volume of the data to be processed, operation S212 can be executed.
[0052] In operation S212, the first candidate encoding mode is determined as the target encoding mode.
[0053] For example, taking the scale of the initial matrix as 128×128 and the number of processor cores as 4 as an example, each element in the initial matrix can be a 32-bit floating-point number (FP32), and the data volume can be 4 bytes (byte). The scale of the data to be processed can be 1×128, and the data volume can be 512 bytes. If the first storage space capacity for the data to be processed is 128 bytes, which is less than the data volume of the data to be processed. The target difference between the first storage space capacity for the data to be processed and the data volume of the data to be processed can be greater than the preset difference threshold. The preset difference threshold is, for example, 4 bytes. Thus, the first candidate encoding mode can be determined as the target encoding mode. A first encoding mark can be added to this initial task. In one example, in the high-level intermediate representation, the first encoding mark can be added to the computing node corresponding to this initial task. It can be understood that the first storage space for the data to be processed can be the cache space for one processor core.
[0054] Next, operation S220 can be executed to generate a task to be executed according to the target encoding mode and the initial task.
[0055] In some embodiments, an executable task is generated according to the number and parameters of processing modules indicated by a target encoding mode, and the operation mode of the data to be processed indicated by an initial task. The operation mode of the data to be processed indicated by the executable task is the same as the operation mode of the data to be processed indicated by the initial task. For example, the operation mode of the data to be processed indicated by the initial task may be row-wise summation. The operation mode indicated by the executable task may also be row-wise summation. The number of processing modules and the parameters of the processing modules may be used as the task parameters of the executable task.
[0056] In some embodiments, the task type of the executable task is one of multiple target task types. The multiple target task types include a first target task type. The executable task of the first target task type may instruct the processing modules to loop through and process the respective sub-data to be processed of multiple data to be processed. In the case where the first candidate encoding mode is used as the target encoding mode, the computation node corresponding to the initial task in the computation graph may have a first encoding mark, and an executable task of the first target task type may be generated.
[0057] For example, there may be N processing modules. The executable task of the first target task type may instruct the nth processing module to process the nth sub-data to be processed of multiple data to be processed. N may be an integer greater than 1. n may be an integer greater than or equal to 1 and less than or equal to N. Taking N = 4 and the size of the initial matrix being 128×128 as an example, the initial matrix may be split into 4 initial sub-matrices. The size of the initial sub-matrix may be 128×32. The first initial sub-matrix may include 128 first sub-data to be processed. The first sub-data to be processed may be the first sub-data to be processed among the N sub-data to be processed, and may include the first to 32nd elements of the data to be processed. The second initial sub-matrix may include 128 second sub-data to be processed. The second sub-data to be processed may be the second sub-data to be processed among the N sub-data to be processed, and may include the 33rd to 64th elements of the data to be processed. The third initial sub-matrix may include 128 third sub-data to be processed. The third sub-data to be processed may be the third sub-data to be processed among the N sub-data to be processed, and may include the 65th to 96th elements of the data to be processed. The fourth initial sub-matrix may include 128 fourth sub-data to be processed. The fourth sub-data to be processed may be the Nth sub-data to be processed among the N sub-data to be processed, and may include the 97th to 128th elements of the data to be processed. In one example, the task to be processed of the first target task type may instruct the first processing module to perform row-wise summation according to the sub-data to be processed in each cycle. The number of cycles is 128.
[0058] Next, operation S230 may be performed to execute the executable task by using the processing modules indicated by the target encoding mode.
[0059] In some embodiments, a processing module configured with parameters indicated by a target encoding pattern is utilized to execute a task to be executed. For example, for the task to be executed of the above first target task type, the processing module can be configured according to the default frequency. After converting the computational graph including the task to be executed of the first target task type into multiple computational instructions, 4 processing modules can respectively execute the multiple computational instructions of each of the 4 tasks to be executed of the first target task type, obtaining 4 intermediate execution results of 128×1. Then, one processing module is utilized to add the 4 intermediate execution results of 128×1 to obtain the target execution result.
[0060] It can be understood that the present disclosure has been described above by taking the first comparison relationship as an example, and the present disclosure will be described below in combination with the second comparison relationship.
[0061] In some other embodiments, the above operations S201 to S203 and operation S211 can be executed. The descriptions of operations S201 to S203 and operation S211 above are equally applicable to this embodiment, and the present disclosure will not be elaborated herein. It can be understood that in some other embodiments, the storage space capacity for the data to be processed obtained by executing operation S203 can be the second storage space capacity.
[0062] In some other embodiments, in response to determining that the target difference is less than the preset difference threshold, operation S213 can be executed.
[0063] In operation S213, the second candidate encoding pattern is determined as the target encoding pattern.
[0064] For example, taking the scale of the initial matrix being 128×128 and the number of processor cores being 4 as an example, each element in the initial matrix can be a 32-bit floating-point number (FP32), and the data volume can be 4 bytes. The scale of the data to be processed can be 1×128, and the data volume can be 512 bytes. If the second storage space capacity for the data to be processed is 512 bytes, which is equal to the data volume of the data to be processed. The target difference between the second storage space capacity for the data to be processed and the data volume of the data to be processed can be less than the preset difference threshold. The preset difference threshold is, for example, 4 bytes. Thus, the second candidate encoding pattern can be determined as the target encoding pattern. A second encoding mark can be added to this initial task. In one example, in the high-level intermediate representation, a second encoding mark can be added to the computational node corresponding to this initial task. It can be understood that the second storage space for the data to be processed can be the cache space for one processor core. It can also be understood that in the case where the target difference between the second storage space capacity for the data to be processed and the data volume of the data to be processed is less than the preset difference threshold, it can be determined that the second storage space capacity for the data to be processed is approximately equal to the data volume of the data to be processed.
[0065] Next, operation S220 can be performed to generate a task to be executed according to the target encoding mode and the initial task.
[0066] In some other embodiments, a task to be executed is generated according to the number and parameters of processing modules indicated by the target encoding mode, and the operation mode of the data to be processed indicated by the initial task. The operation mode of the data to be processed indicated by the task to be executed is the same as the operation mode of the data to be processed indicated by the initial task. For example, the operation mode of the data to be processed indicated by the initial task can be row-wise summation. The operation mode indicated by the task to be executed can also be row-wise summation. The number of processing modules and the parameters of the processing modules can be used as the task parameters of the task to be executed.
[0067] In some other embodiments, the task type of the task to be executed is one of multiple target task types. The multiple target task types include a second target task type. The task to be executed of the second target task type is used to instruct a processing module to cyclically process multiple data to be processed. In the case where the second candidate encoding mode is used as the target encoding mode, the computing node corresponding to the initial task in the computation graph can have a second encoding mark, and a task to be executed of the second target task type can be generated. For example, taking the scale of the initial matrix as 128×128 as an example, if the capacity of the second storage space is large, one data to be processed can be stored. In this case, the task to be processed of the second target task type can instruct a processing module to perform row-wise summation according to the data to be processed in each cycle period. The number of cycles is 128. Through the embodiments of the present disclosure, an encoding mark is set on the computing node, and the block strategy is explicitly embedded into the high-level intermediate identifier based on the encoding injection mechanism, which supports the backend to generate the parameters of the task to be executed according to different hardware characteristics and enhances the hardware resource adaptation ability. The task to be executed can correspond to a kernel function.
[0068] Next, operation S230 can be performed to execute the task to be executed by using the processing module indicated by the target encoding mode.
[0069] In some other embodiments, the task to be executed is executed by using a processing module configured with parameters indicated by the target encoding mode. For example, for the task to be executed of the above second target task type, the processing module can be configured according to the default frequency. After converting the computation graph including the task to be executed of the second target task type into multiple computation instructions, the processing module can execute the multiple computation instructions of the task to be executed of the second target task type to obtain a target execution result of 128×1.
[0070] It can be understood that the present disclosure is described above by taking the second comparison relationship as an example, and the present disclosure will be described below in combination with the third comparison relationship.
[0071] In some other embodiments, the above operations S201 to S203 and operation S211 may be performed. The descriptions of operations S201 to S203 and operation S211 above also apply to this embodiment, and the present disclosure will not repeat them here. It can be understood that in some other embodiments, the storage space capacity for the data to be processed obtained by performing operation S203 may be the third storage space capacity.
[0072] In some other embodiments, when it is determined that the target difference is greater than the preset difference threshold and the storage space capacity for the data to be processed is greater than the data volume of the data to be processed, operation S214 may be performed.
[0073] In operation S214, the third candidate encoding mode is determined as the target encoding mode.
[0074] For example, taking the scale of the initial matrix as 128×128 and the number of processor cores as 4 as an example, each element in the initial matrix may be a 32-bit floating-point number (FP32), and the data volume may be 4 bytes. The scale of the data to be processed may be 1×128, and the data volume may be 512 bytes. If the third storage space capacity for the data to be processed is 1024 bytes, which is greater than the data volume of the data to be processed. The target difference between the third storage space capacity for the data to be processed and the data volume of the data to be processed may be greater than the preset difference threshold. The preset difference threshold is, for example, 4 bytes. Thus, the third candidate encoding mode can be determined as the target encoding mode. A third encoding mark may be added to this initial task. In one example, in the high-level intermediate representation, a third encoding mark may be added to the computing node corresponding to this initial task. It can be understood that the second storage space for the data to be processed may be the cache space for one processor core.
[0075] Next, operation S220 may be performed to generate an executable task according to the target encoding mode and the initial task.
[0076] In some other embodiments, an executable task is generated according to the number and parameters of the processing modules indicated by the target encoding mode and the operation mode of the data to be processed indicated by the initial task. The operation mode of the data to be processed indicated by the executable task is consistent with the operation mode of the data to be processed indicated by the initial task. For example, the operation mode of the data to be processed indicated by the initial task may be row-wise summation. The operation mode indicated by the executable task may also be row-wise summation. The number of processing modules and the parameters of the processing modules may be used as the task parameters of the executable task.
[0077] In some other embodiments, the task type of the task to be executed is one of multiple target task types. The multiple target task types include a third target task type. The task to be executed of the third target task type is used to instruct a processing module to cyclically obtain multiple memory access merged data and cyclically process multiple pieces of to-be-processed data respectively in the multiple memory access merged data. The memory access merged data can be obtained by performing a memory access on multiple pieces of to-be-processed data together. In the case of taking the third candidate coding mode as the target coding mode, the computing node corresponding to the initial task in the computation graph can have a third coding mark, and a task to be executed of the third target task type can be generated. For example, taking the scale of the initial matrix as 128×128 as an example, if the capacity of the third storage space is large and can store 2 pieces of to-be-processed data. In this case, 2 pieces of to-be-processed data can be obtained each time a memory access is performed as the memory access merged data. The to-be-processed task of the third target task type can instruct a processing module to perform a row-wise sum according to 2 pieces of to-be-processed data in each cycle period. The number of cycles can be 64. Through the embodiments of the present disclosure, coding marks are set on the computing nodes, realizing the explicit embedding of the chunking strategy into the high-level intermediate identifier based on the coding injection mechanism, supporting the backend to generate parameters of the task to be executed according to different hardware characteristics, and enhancing the hardware resource adaptation ability. For an artificial intelligence processor with limited global storage units, the frequency of the processor core can be automatically set to overclock, so as to reduce the bandwidth pressure through memory access merging optimization, and it can be applied to heterogeneous hardware scenarios such as edge computing.
[0078] Next, operation S230 can be executed to execute the task to be executed by using the processing module indicated by the target coding mode.
[0079] In some other embodiments, the task to be executed is executed by using a processing module configured with parameters indicated by the target coding mode. For example, for the task to be executed of the above-mentioned third target task type, the processing module can be configured according to the overclock frequency. After converting the computation graph including the task to be executed of the third target task type into multiple computation instructions, the processing module can execute the task to be executed of the third target task type to obtain a target execution result of 128×1.
[0080] It can be understood that the present disclosure is described above by taking the operation mode as row-wise sum as an example. However, the present disclosure is not limited thereto, and the operation mode can be various operations, such as row-wise multiplication, row-wise subtraction, transpose, etc.
[0081] It can be understood that the method of the present disclosure is described above, and the apparatus of the present disclosure will be described below.
[0082] Figure 3 It is a block diagram of a task execution apparatus according to an embodiment of the present disclosure.
[0083] As Figure 3As shown, the apparatus 300 may include a determination module 310, a generation module 320, and an execution module 330.
[0084] The determination module 310 is configured to determine a target encoding mode according to the storage space capacity for the data to be processed and the data volume of the data to be processed. The target encoding mode is used to indicate at least one of the number and parameters of processing modules for processing the data to be processed, and the data to be processed is obtained from the initial data of the initial task.
[0085] The generation module 320 is configured to generate a task to be executed according to the target encoding mode and the initial task.
[0086] The execution module 330 is configured to execute the task to be executed by using the processing modules indicated by the target encoding mode.
[0087] In some embodiments, the determination module includes: a determination sub-module configured to determine a target encoding mode from a plurality of candidate encoding modes according to the storage space capacity for the data to be processed and the data volume of the data to be processed. The plurality of candidate encoding modes includes at least one of a first candidate encoding mode, a second candidate encoding mode, and a third candidate encoding mode. The first candidate encoding mode is used to indicate a plurality of sub-data to be processed that are respectively processed by a plurality of processing modules for the data to be processed. The second candidate encoding mode is used to indicate that the data to be processed is processed by a processing module. The third candidate encoding mode is used to indicate that a plurality of data to be processed are processed by a processing module.
[0088] In some embodiments, the determination sub-module includes at least one of the following: a first determination unit configured to determine the first candidate encoding mode as the target encoding mode in response to determining that the target difference between the storage space capacity for the data to be processed and the data volume of the data to be processed is greater than a preset difference threshold and the storage space capacity for the data to be processed is less than the data volume of the data to be processed. A second determination unit configured to determine the second candidate encoding mode as the target encoding mode in response to determining that the target difference is less than or equal to the preset difference threshold. A third determination unit configured to determine the third candidate encoding mode as the target encoding mode in response to determining that the target difference is greater than the preset difference threshold and the storage space capacity for the data to be processed is greater than the data volume of the data to be processed.
[0089] In some embodiments, the initial data includes a plurality of data to be processed. The task type of the task to be executed is one of a plurality of target task types, and the plurality of target task types includes at least one of a first target task type, a second target task type, and a third target task type. The task to be executed of the first target task type is used to instruct the processing module to loop through and process the to-be-processed sub-data of each of the plurality of data to be processed. The task to be executed of the second target task type is used to instruct the processing module to loop through and process the plurality of data to be processed. The task to be executed of the third target task type is used to instruct: the processing module to loop through and obtain a plurality of memory access merged data and loop through and process the plurality of to-be-processed data of each of the plurality of memory access merged data.
[0090] In some embodiments, the generation module includes: a generation sub-module, configured to generate a task to be executed according to the number and parameters of the processing modules indicated by the target encoding mode, and the operation mode of the data to be processed indicated by the initial task. The operation mode of the data to be processed indicated by the task to be executed is the same as the operation mode of the data to be processed indicated by the initial task.
[0091] In some embodiments, the execution module includes: an execution sub-module, configured to execute the task to be executed by using the processing modules configured with the parameters indicated by the target encoding mode.
[0092] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0093] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0094] Figure 4 FIG. shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0095] As Figure 4As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 402 or computer programs loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0096] Multiple components in device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0097] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the task execution method. For example, in some embodiments, the task execution method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the task execution method described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the task execution method in any other appropriate way (e.g., by means of firmware).
[0098] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0099] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0100] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0101] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0102] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0103] A computer system may include a client and a server. The client and the server are generally far from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other.
[0104] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added, or deleted. For example, the steps described in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0105] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A task execution method, comprising: Determining a target encoding mode according to the storage space capacity for the data to be processed and the data volume of the data to be processed, wherein the target encoding mode is used to indicate at least one of the number and parameters of processing modules for processing the data to be processed, and the data to be processed is obtained from the initial data of an initial task; Generating a task to be executed according to the target encoding mode and the initial task; Executing the task to be executed by using the processing modules indicated by the target encoding mode.
2. The method according to claim 1, wherein, The determining a target encoding mode according to the storage space capacity for the data to be processed and the data volume of the data to be processed includes: Determining a target encoding mode from multiple candidate encoding modes according to the storage space capacity for the data to be processed and the data volume of the data to be processed, wherein the multiple candidate encoding modes include at least one of a first candidate encoding mode, a second candidate encoding mode, and a third candidate encoding mode; The first candidate encoding mode is used to indicate multiple sub-data to be processed of the data to be processed, which are respectively processed by multiple processing modules; The second candidate encoding mode is used to indicate that the processing module processes the data to be processed; The third candidate encoding mode is used to indicate that the processing module processes multiple data to be processed.
3. The method according to claim 2, wherein, The determining a target encoding mode from multiple candidate encoding modes according to the storage space capacity for the data to be processed and the data volume of the data to be processed includes at least one of the following: In response to determining that the target difference between the storage space capacity for the data to be processed and the data volume of the data to be processed is greater than a preset difference threshold and the storage space capacity for the data to be processed is less than the data volume of the data to be processed, determining the first candidate encoding mode as the target encoding mode; In response to determining that the target difference is less than or equal to the preset difference threshold, determining the second candidate encoding mode as the target encoding mode; In response to determining that the target difference is greater than the preset difference threshold and the storage space capacity for the data to be processed is greater than the data volume of the data to be processed, determining the third candidate encoding mode as the target encoding mode.
4. The method according to claim 2, wherein The initial data includes multiple data to be processed; The task type of the task to be executed is one of multiple target task types, and the multiple target task types include at least one of a first target task type, a second target task type, and a third target task type; The task to be executed of the first target task type is used to indicate that the processing module circularly processes the sub-data to be processed of each of the multiple data to be processed; The task to be executed of the second target task type is used to indicate that the processing module circularly processes the multiple data to be processed; The task to be executed of the third target task type is used to indicate that the processing module circularly obtains multiple memory access merged data and circularly processes the multiple data to be processed of each of the multiple memory access merged data.
5. The method according to claim 1, wherein The generating a task to be executed according to the target encoding mode and the initial task includes: Generate the to-be-executed task according to the number and parameters of the processing modules indicated by the target coding mode, and the operation mode of the data to be processed indicated by the initial task, where the operation mode of the data to be processed indicated by the to-be-executed task is the same as the operation mode of the data to be processed indicated by the initial task.
6. The method according to claim 1, wherein The executing the to-be-executed task by using the processing modules indicated by the target coding mode includes: Execute the to-be-executed task by using the processing modules configured with the parameters indicated by the target coding mode.
7. A task execution device, comprising: A determination module, configured to determine a target coding mode according to the storage space capacity for the data to be processed and the data volume of the data to be processed, where the target coding mode is used to indicate at least one of the number and parameters of the processing modules for processing the data to be processed, and the data to be processed is obtained from the initial data of the initial task; A generation module, configured to generate a to-be-executed task according to the target coding mode and the initial task; An execution module, configured to execute the to-be-executed task by using the processing modules indicated by the target coding mode.
8. The apparatus according to claim 7, wherein, The determination module includes: A determination sub-module, configured to determine a target coding mode from multiple candidate coding modes according to the storage space capacity for the data to be processed and the data volume of the data to be processed, where the multiple candidate coding modes include at least one of a first candidate coding mode, a second candidate coding mode, and a third candidate coding mode, The first candidate coding mode is used to indicate that multiple sub-data to be processed of the data to be processed are respectively processed by multiple processing modules; The second candidate coding mode is used to indicate that the data to be processed is processed by the processing module; The third candidate coding mode is used to indicate that multiple data to be processed are processed by the processing module.
9. The apparatus according to claim 8, wherein, The determination sub-module includes at least one of the following: A first determination unit, configured to determine the first candidate coding mode as the target coding mode in response to determining that the target difference between the storage space capacity for the data to be processed and the data volume of the data to be processed is greater than a preset difference threshold and the storage space capacity for the data to be processed is less than the data volume of the data to be processed; A second determination unit, configured to determine the second candidate coding mode as the target coding mode in response to determining that the target difference is less than or equal to the preset difference threshold; A third determination unit, configured to determine the third candidate coding mode as the target coding mode in response to determining that the target difference is greater than the preset difference threshold and the storage space capacity for the data to be processed is greater than the data volume of the data to be processed.
10. The apparatus according to claim 8, wherein, The initial data includes multiple data to be processed, The task type of the to-be-executed task is one of multiple target task types, and the multiple target task types include at least one of a first target task type, a second target task type, and a third target task type, The to-be-executed task of the first target task type is used to indicate that the processing module circularly processes the sub-data to be processed of each of the multiple data to be processed; The task to be executed of the second target task type is used to instruct the processing module to process multiple pieces of the data to be processed in a loop; The task to be executed of the third target task type is used to instruct that: the processing module obtains multiple memory access merged data in a loop and processes multiple pieces of the data to be processed of each of the multiple memory access merged data in a loop.
11. The apparatus according to claim 7, wherein, The generating module includes: A generating sub-module, configured to generate the task to be executed according to the number and parameters of the processing modules indicated by the target encoding mode, and the operation mode of the data to be processed indicated by the initial task, wherein the operation mode of the data to be processed indicated by the task to be executed is the same as the operation mode of the data to be processed indicated by the initial task.
12. The device according to claim 7, wherein, The execution module includes: An execution sub-module, configured to execute the task to be executed by using the processing modules configured with the parameters indicated by the target encoding mode.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
15. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 6.