Calculation task acceleration method and device based on coarse-grained data flow architecture
By configuring the flow graph program loading instructions and base address flag bits, the multiplexing of intermediate result data in the coarse-grained data flow architecture is solved, and the problem of large overhead of data handling and flow graph program switching in the prior art is improved, and the execution efficiency of computing tasks is improved.
Patent Information
- Application Number
- CN202410063385.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
The existing coarse-grained data flow architecture cannot effectively support the multiplexing of intermediate result data when executing calculation tasks, resulting in excessive overhead of data handling and flow chart program switching, affecting the execution efficiency of computing tasks.
By configuring the flow chart program loading instructions, flow chart program base address and flow chart program base address valid flag bits, multiple computing tasks are allowed to share the same flow chart program, realizing the reuse of intermediate result data, and reducing the overhead of data handling and flow chart program switching.
When multiple computing tasks share the same streaming graph program, it supports intermediate result data reuse, reduces data transfer and flow graph program switching overhead, improves computing task execution efficiency, and flexibly allocates the storage space of on-chip cache modules.
Smart Images

Figure CN120335940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, specifically to data processing technologies based on a coarse-grained data flow architecture. More specifically, it relates to computing task processing technologies based on a coarse-grained data flow architecture, namely, a method and device for accelerating computing tasks based on a coarse-grained data flow architecture. Background Art
[0002] The coarse-grained data flow architecture is a special parallel computing architecture that uses a flow graph program to execute computing tasks. Among them, the flow graph program indicates the data transfer direction between different computing nodes (the computing nodes can be operators, subtasks within an operator, or one or more program blocks within a subtask. In the present invention, the flow graph program mainly targets the scenario where the computing node is a program block).
[0003] The coarse-grained data flow architecture is suitable for processing large-scale parallel computing tasks. In the coarse-grained data flow architecture, a program is usually divided into multiple program blocks, and the program blocks interact through data streams to form a complex and efficient flow graph program. The coarse-grained data flow architecture usually adopts an execution array for organization. The execution array includes multiple execution units (Process Element, PE), and several program blocks are executed on each execution unit. The same execution mode as that of a traditional control flow processor is adopted within a program block, that is, using a program counter to execute instructions one by one; while due to data dependency relationships between program blocks, they are driven by data for execution. The flow graph program in the coarse-grained data flow architecture is usually statically compiled by a compiler. However, due to the storage resource limitations of the execution unit, generally, the multi-layer loop program in a large-scale computing task cannot be fully expanded. Therefore, usually, one or several loop executions of the loop program are expanded into a flow graph program, and the flow graph program is iteratively executed multiple times to complete the function of the entire loop program. Usually, an iteration ID (Iteration id) is used for iteration differentiation, and the program blocks in the flow graph program are generally differentiated by a block ID (Block id). In addition, when the data scale to be processed by a parallel computing task is very large and exceeds the storage space capacity, multiple data blocks (Tile) are generally further divided, and each data block (Tile) corresponds to the processing of a part of the data.
[0004] The mapping relationship between the coarse-grained data flow architecture and the flow graph program is as Figure 1 shown, from Figure 1It can be seen that the coarse-grained data flow architecture includes a control core, a main memory, a microcontroller, an execution array, an on-chip cache module, an execution storage access module (DMA), and a bus. The program blocks in the flow graph program are mapped to the execution units in the execution array. Among them, the control core is used to configure the execution storage access module to obtain data from the main memory or the on-chip cache module. The main memory and the on-chip cache module are used to store data. The microcontroller is used to control the execution array. The execution array is used to execute the flow graph program. The bus is used for connection and data communication between the control core, the main memory, the microcontroller, and the execution storage access module.
[0005] The existing coarse-grained data flow architecture usually only supports the address access mode of fixed cache partitioning, that is, the cache space of each data block is completely independent, and only its respective cache space can be accessed, and it cannot support the reuse of intermediate result data between data blocks. Based on this access mode, when executing a computing task, there may be two situations for reusing intermediate result data. One is to continuously read the previous intermediate result data from the main memory, and the other is to start a new flow graph program to further process the intermediate result data that may be needed. Both of these situations will bring more data transfer overhead, and restarting the flow graph program will also bring more flow graph switching overhead, which is not conducive to improving the execution efficiency of computing tasks.
[0006] It should be noted that: This background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily prior art. Without evidence showing that the relevant information has been publicly disclosed before the filing date of the present invention, the relevant information should not be regarded as prior art. Summary of the Invention
[0007] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a computing task acceleration method based on a coarse-grained data flow architecture and a computing task acceleration device for a coarse-grained data flow architecture.
[0008] The object of the present invention is achieved by the following technical solutions:
[0009] According to a first aspect of the present invention, there is provided a method for accelerating a computing task based on a coarse-grained data flow architecture. The computing task corresponds to multiple operands, and there is a flow graph program corresponding to the computing task. The flow graph program is executed based on the operands to complete the computing task. The method includes: S1. Configure a flow graph program loading instruction, a flow graph program base address, and a flow graph program base address valid flag bit corresponding to the current computing task. Among them, the flow graph program loading instruction corresponding to the current computing task is used to indicate whether the flow graph program needs to be reloaded when executing the current computing task. When sharing the flow graph program with the previous computing task, the flow graph program is not reloaded; the flow graph program base address corresponding to the current computing task is used to indicate the starting address of the cache of the current computing task; the flow graph program base address valid flag bit corresponding to the current computing task is used to indicate whether the storage offset value of each operand in the current computing task needs to be accumulated with the flow graph program base address corresponding to the current computing task to address the data of each operand; S2. Based on the flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current computing task configured in step S1, obtain the storage base addresses of multiple operands in the current computing task and execute the flow graph program to complete the current computing task.
[0010] In some embodiments of the present invention, the number of bits of the flow graph program base address valid flag bit corresponding to the current computing task corresponds one-to-one to the number of operands in the current computing task, and the value of each bit in the flow graph program base address valid flag bit corresponding to the current computing task is 0 or 1. One value indicates that the storage offset value of the operand corresponding to this bit does not need to be accumulated with the flow graph program base address corresponding to the current computing task, and the other value indicates that the storage offset value of the operand corresponding to this bit needs to be accumulated with the flow graph program base address corresponding to the current computing task.
[0011] In some embodiments of the present invention, in step S1, when configuring the flow graph program base address corresponding to the current computing task, the flow graph program base address corresponding to the first executed computing task among multiple computing tasks sharing the same flow graph program is configured to 0.
[0012] In some embodiments of the present invention, in step S1, when configuring the flow graph program base address valid flag bit corresponding to the current computing task, the bit corresponding to the operand that needs to be reused in the current computing task in the flow graph program base address valid flag bit corresponding to the current computing task is set to the value indicating that it does not need to be accumulated with the flow graph program base address corresponding to the current computing task, and the storage offset value of the operand that needs to be reused in the current computing task is kept consistent with the storage offset value of the same operand that needs to be reused in other computing tasks.
[0013] According to a second aspect of the present invention, there is provided a computing task acceleration device for a coarse-grained data flow architecture. The computing task corresponds to multiple operands, and the computing task corresponds to a flow graph program. The flow graph program is executed based on the operands to complete the computing task. The device is configured to process the computing task according to the method described in the first aspect of the present invention.
[0014] In some embodiments of the present invention, the device includes a control core, a microcontroller, an execution array, and an on-chip cache module, wherein: The control core is configured to configure a flow graph program loading instruction, a flow graph program base address, and a flow graph program base address valid flag bit corresponding to the computing task and transmit them to the microcontroller. Among them, the flow graph program loading instruction corresponding to the computing task is used to indicate whether the flow graph program needs to be reloaded when executing the computing task. When sharing the flow graph program with the previous computing task, the flow graph program is not reloaded; The flow graph program base address corresponding to the computing task is used to indicate the starting address of the computing task in the on-chip cache module; The flow graph program base address valid flag bit corresponding to the computing task is used to indicate whether the storage offset value of each operand in the computing task needs to be accumulated with the flow graph program base address corresponding to the computing task to address the data of each operand on the on-chip cache module; and is used to send a start execution message to the microcontroller; The microcontroller is configured to load the flow graph program based on the flow graph program loading instruction corresponding to the computing task transmitted by the control core, and is used to store the flow graph program base address and the flow graph program base address valid flag bit corresponding to the computing task transmitted by the control core and calculate the storage base address of multiple operands on the on-chip cache module based on them, and control the execution array to execute the flow graph program based on the received start execution message; The execution array is configured to execute the flow graph program based on the data of each operand addressed by the microcontroller on the on-chip cache module to complete the computing task; The on-chip cache module is used to store the data corresponding to multiple operands respectively when executing the computing task and the data obtained after the execution array executes the flow graph program based on the operands.
[0015] In some embodiments of the present invention, a storage module and a computing logic unit are configured on the microcontroller. Among them, the storage module is used to store the flow graph program base address corresponding to the computing task transmitted by the control core and the flow graph program base address valid flag bit corresponding to the computing task; The computing logic unit is configured to calculate the storage base address of multiple operands on the on-chip cache module based on the flow graph program base address corresponding to the current computing task and the flow graph program base address valid flag bit stored in the storage unit.
[0016] In some embodiments of the present invention, the storage module includes: a flow graph program base address storage unit for storing the flow graph program base address corresponding to the computing task transmitted by the control core; a flow graph program base address valid flag storage unit for storing the flow graph program base address valid flag corresponding to the computing task transmitted by the control core.
[0017] In some embodiments of the present invention, the device is configured to execute a computing task in the following manner: configure a flow graph program loading instruction, a flow graph program base address, and a flow graph program base address valid flag corresponding to the current computing task in the control core and transmit them to the microcontroller; the microcontroller, based on the received flow graph program loading instruction, flow graph program base address, and flow graph program base address valid flag corresponding to the current computing task, starts the computing logic unit to calculate the storage base address of multiple operands in the current computing task in the on-chip cache module, and starts the execution array to execute the flow graph program to complete the current computing task.
[0018] In some embodiments of the present invention, when configuring the flow graph program base address corresponding to the current computing task in the control core, the flow graph program base address corresponding to the first executed computing task among multiple computing tasks sharing the same flow graph program is configured to be 0.
[0019] In some embodiments of the present invention, when configuring the flow graph program base address valid flag corresponding to the current computing task in the control core, the number of bits of the flow graph program base address valid flag corresponding to the current computing task corresponds one-to-one with the number of operands in the current computing task, and the value of each bit in the flow graph program base address valid flag corresponding to the current computing task is 0 or 1, and one value indicates that the storage offset value of the operand corresponding to this bit does not need to be accumulated with the flow graph program base address corresponding to the current computing task, and the other value indicates that the storage offset value of the operand corresponding to this bit needs to be accumulated with the flow graph program base address corresponding to the current computing task.
[0020] In some embodiments of the present invention, when configuring the flow graph program base address valid flag corresponding to the current computing task in the control core, the bit corresponding to the operand that needs to be reused in the current computing task in the flow graph program base address valid flag corresponding to the current computing task is set to the value indicating that it does not need to be accumulated with the flow graph program base address corresponding to the current computing task, and the storage offset value of the operand that needs to be reused in the current computing task is kept consistent with the storage offset value of the same operand that needs to be reused in other computing tasks.
[0021] Compared with the prior art, the advantages of the present invention are as follows: when multiple computing tasks share the same flow graph program, it can support the reuse of intermediate result data of multiple computing tasks, reduce the data transfer overhead and the flow graph program switching overhead, and improve the execution efficiency of computing tasks; it can also flexibly allocate the storage space of the on-chip cache module to support usage modes such as single cache, dual cache, or multi-cache of the on-chip cache module. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The following further describes embodiments of the present invention with reference to the accompanying drawings, where:
[0023] Figure 1 is a schematic diagram of the composition of an existing coarse-grained data flow architecture;
[0024] Figure 2 is a computing task acceleration device for a coarse-grained data flow architecture according to an embodiment of the present invention;
[0025] Figure 3 is a schematic diagram of the composition structure of a microcontroller according to an embodiment of the present invention;
[0026] Figure 4 is a schematic diagram of the flow of a computing task acceleration method according to an embodiment of the present invention;
[0027] Figure 5 is a computing schematic diagram taking the matrix multiplication calculation example as an example;
[0028] Figure 6 is a schematic diagram of the storage allocation of each computing task in the on-chip cache module during the C1 calculation taking the matrix multiplication calculation example as an example. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further details the present invention through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0030] For the convenience of understanding, the functions of the components of the general coarse-grained data flow architecture will be briefly introduced first. As introduced in the background art above, the coarse-grained data flow architecture includes a control core, a main memory, a microcontroller, an execution array, an on-chip cache module, an execution storage access module, and a bus. Among them, the control core is used to configure the execution storage access module to transfer the flow graph program corresponding to the computing task and the data required for executing the computing task from the main memory to the on-chip cache module, or configure the execution storage access module to transfer the data generated after the execution array executes the flow graph program stored in the on-chip cache module back to the main memory, and is also used to send a start execution message to the microcontroller through the bus to start the execution array to execute the flow graph program; the main memory is used to store the flow graph program, the data required for executing the computing task, and the data transferred by the on-chip cache module; the microcontroller is used to control the execution array to execute the flow graph program; the execution array is used to execute the flow graph program to complete the computing task; the on-chip cache module is used to store the data transferred from the main memory and the data obtained after the execution array executes the flow graph program; the execution storage access module is used to transfer the flow graph program stored in the main memory and the data required for executing the computing task to the on-chip cache module based on the configuration of the control core, or transfer the data generated after the execution array executes the flow graph program stored in the on-chip cache module back to the main memory; the bus is used for the connection and data communication between the control core, the main memory, the microcontroller, and the execution storage access module.
[0031] According to the content introduced in the above background art, the existing coarse-grained data flow architecture usually only supports the address access mode of fixed cache partitioning, that is, the cache space of each data block is completely independent, and only its respective cache space can be accessed, and it cannot support the reuse of intermediate result data between data blocks. Based on this access mode, when executing a computing task, there may be two situations for reusing intermediate result data. One is to continuously read the previous intermediate result data from the main memory, and the other is to start a new flow graph program to further process the intermediate result data that may be needed. Both of these situations will bring more data transfer overhead, and restarting the flow graph program will also bring more flow graph switching overhead, which is not conducive to improving the execution efficiency of the computing task.
[0032] To solve the above problems, the inventor analyzed the coarse-grained data flow architecture and found that: for the existing coarse-grained data flow architecture, the memory access address of data during the execution of a computing task is generally composed of the base address of the flow graph program, the iteration base address of the flow graph program, and the memory access instruction offset address within the flow graph. Among them, the base address of the flow graph program is closely related to the usage pattern of the on-chip cache module. Usually, one base address of the flow graph program corresponds to the base address of a data block (Tile), that is, the base address of the flow graph program is used to indicate the starting address of the data corresponding to the data block cached in the on-chip cache module. The iteration base address of the flow graph program is used to indicate the storage offset address of different data in a data block (Tile) relative to the base address of the flow graph program during each iteration process (processing a data block may require iteratively executing the flow graph program multiple times). The memory access instruction offset address within the flow graph is used to indicate how to obtain the specific values (offsets relative to the iteration base address of the flow graph program) contained in different data in a data block (Tile) during each iteration process. For example, a data block (Tile) corresponds to X = M * N (both M and N are k * k matrices). At this time, the base address of the flow graph program is used to indicate the starting addresses of matrices X, M, and N cached in the on-chip cache module for this data block. Assume the base address of the flow graph program is 0; the iteration base address of the flow graph program is then used to indicate the storage offset addresses of some data in matrices X, M, and N relative to the base address of the flow graph program during each iteration process. Assume that the storage offset address of some data in matrix X relative to the base address of the flow graph program during the first iteration is 0. Then, during the first iteration, the starting storage address of some data in the result matrix X is 0 + 0 in the on-chip cache module. Assume that the storage offset address of some data in matrix M relative to the base address of the flow graph program during the first iteration is x. Then, during the first iteration, the addressing of some data in matrix M starts from address 0 + x in the on-chip cache module. Assume that the storage offset address of matrix N relative to the base address of the flow graph program during the first iteration is x + m. Then, during the first iteration, the addressing of some data in matrix N starts from address 0 + x + m in the on-chip cache module. The principle of the addressing process during other iteration executions is the same as that during the first iteration and will not be elaborated here; the memory access instruction offset address within the flow graph is used to indicate how to obtain the specific values of some data in matrices X, M, and N during each iteration process; among them, the iteration base address of the flow graph program is statically generated by the compiler and stored in the address table in the on-chip cache or dynamically calculated by the microcontroller before each iteration execution.Moreover, existing coarse-grained data flow architectures usually only support address access modes with fixed cache partitioning and cannot support arbitrary cache address partitioning. Taking the dual-cache mode as an example, the storage space on the on-chip cache module is divided into two cache areas, each cache area has an independent address range, and the program block can only access the data in its corresponding cache area and cannot directly access the data in the other cache area. This address access mode with fixed cache partitioning results in completely independent cache spaces for each data block during the execution of computational tasks, and only their respective cache spaces can be accessed, making it impossible to support the direct reuse of intermediate result data between data blocks in the cache space. Based on the analysis of existing coarse-grained data flow architectures, the present invention proposes an addressing scheme for accessing the storage space of intermediate result data to achieve the direct reuse of intermediate result data. Specifically, the present invention provides hardware support by configuring a storage module and a computing logic unit in the microcontroller, and after configuring the base address of the flow graph program corresponding to the computational task and the valid flag of the base address of the flow graph program in the control core and transmitting them to the microcontroller, the microcontroller can, after receiving the base address of the flow graph program corresponding to the computational task and the valid flag of the base address of the flow graph program, implement the addressing of intermediate result data based on the hardware configured by itself. In this way, the storage base address of the intermediate result data that needs to be reused can be accessed during subsequent iterations of the flow graph program, thereby enabling the intermediate result data to be stored in the same address space, reducing the data transfer overhead and the flow graph program switching overhead, and improving the execution efficiency of computational tasks.
[0033] Before specifically introducing the embodiments of the present invention, some terms used therein are explained as follows:
[0034] The computational task is a task such as data processing and data analysis performed in a coarse-grained data flow architecture. Among them, the computational task corresponds to multiple operands, and the computational task corresponds to a flow graph program, and the flow graph program is executed based on the operands to complete the computational task.
[0035] The base address of the flow graph program is used to indicate the base address of the storage space allocated to the computational task on the on-chip cache module, that is, the base address of the flow graph program indicates the starting address of the cache of the computational task on the on-chip cache module.
[0036] The stream graph program base address valid flag is used to indicate whether the storage offset value of each operand in the computing task needs to be added to the stream graph program base address corresponding to the computing task to address the data of each operand on the on-chip cache module. It should be noted that the stream graph program base address valid flag can also be understood as being used to indicate whether the iteration base address of each operand in the computing task needs to be added to the stream graph program base address corresponding to the computing task to address the data of each operand on the on-chip cache module. The iteration base address of each operand represents the storage offset address (storage offset value) of different operands relative to the stream graph program base address in each iteration process.
[0037] To better understand the present invention, the present invention will be described in detail below in conjunction with an acceleration device and a matrix multiplication calculation embodiment.
[0038] I. Acceleration Device
[0039] According to an embodiment of the present invention, as Figure 2 shown, a computing task acceleration device for a coarse-grained data flow architecture is provided. The computing task corresponds to multiple operands, and the computing task corresponds to a stream graph program. The device includes a control core, a microcontroller, an execution array, and an on-chip cache module; the control core configures a stream graph program loading instruction, a stream graph program base address, and a stream graph program base address valid flag corresponding to the computing task and transmits them to the microcontroller. Among them, the stream graph program loading instruction corresponding to the computing task is used to indicate whether the stream graph program needs to be reloaded when executing this computing task. When sharing the stream graph program with the previous computing task, the stream graph program is not reloaded; the stream graph program base address corresponding to the computing task is used to indicate the starting address of the computing task on the on-chip cache module; the stream graph program base address valid flag corresponding to the computing task is used to indicate whether the storage offset value of each operand in the computing task needs to be added to the stream graph program base address corresponding to the computing task to address the data of each operand on the on-chip cache module; and it is used to send a start execution message to the microcontroller; the microcontroller is used to load the stream graph program based on the stream graph program loading instruction corresponding to the computing task transmitted by the control core, and is used to store the stream graph program base address and the stream graph program base address valid flag corresponding to the computing task transmitted by the control core and calculate the storage base address of multiple operands on the on-chip cache module based on them, and control the execution array to execute the stream graph program based on the received start execution message; the execution array is used to execute the stream graph program based on the data of each operand addressed by the microcontroller on the on-chip cache module to complete the computing task; the on-chip cache module is used to store the data corresponding to multiple operands respectively when executing the computing task and the data obtained after the execution array executes the stream graph program based on the operands.
[0040] According to an embodiment of the present invention, asFigure 3 As shown, a storage module and a computing logic unit are configured on the microcontroller. Among them, the storage module is used to store the base address of the flow graph program corresponding to the computing task transmitted by the control core and the valid flag bit of the base address of the flow graph program corresponding to the computing task; the computing logic unit is used to calculate the storage base addresses of multiple operands in the on-chip cache module based on the base address of the flow graph program corresponding to the computing task stored in the storage unit and the valid flag bit of the base address of the flow graph program corresponding to the computing task. According to an embodiment of the present invention, the storage module includes: a flow graph program base address storage unit for storing the base address of the flow graph program corresponding to the computing task transmitted by the control core; a flow graph program base address valid flag bit storage unit for storing the valid flag bit of the base address of the flow graph program corresponding to the computing task transmitted by the control core.
[0041] Based on the acceleration device configured in the foregoing embodiment, according to an embodiment of the present invention, as Figure 4 shown, the present invention further provides a computing task acceleration method applicable to the acceleration device, including: S1. Configure the flow graph program loading instruction, the base address of the flow graph program, and the valid flag bit of the base address of the flow graph program corresponding to the current computing task. Among them, the flow graph program loading instruction corresponding to the current computing task is used to indicate whether it is necessary to reload the flow graph program when executing the current computing task, and the flow graph program is not reloaded when sharing the flow graph program with the previous computing task; the base address of the flow graph program corresponding to the current computing task is used to indicate the starting address of the cache of the current computing task; the valid flag bit of the base address of the flow graph program corresponding to the current computing task is used to indicate whether the storage offset value of each operand in the current computing task needs to be accumulated with the base address of the flow graph program corresponding to the current computing task to address and obtain the data of each operand; S2. Based on the flow graph program loading instruction, the base address of the flow graph program, and the valid flag bit of the base address of the flow graph program corresponding to the current computing task configured in step S1, obtain the storage base addresses of multiple operands in the current computing task and execute the flow graph program to complete the current computing task.
[0042] Based on the proposed computing task acceleration method, in the actual application process, the computing task acceleration device can execute the computing task in the following manner: First, configure the flow graph program loading instruction, flow graph program base address, and flow graph program base address valid flag bit corresponding to the computing task to be executed in the control core and transmit them to the microcontroller. Then, based on the received flow graph program loading instruction, flow graph program base address, and flow graph program base address valid flag bit corresponding to the computing task to be executed, the microcontroller starts the computing logic unit to calculate the storage base addresses of multiple operands in the computing task to be executed in the on-chip cache module, and starts the execution array to execute the flow graph program to complete the computing task. It should be noted that the code for configuring the flow graph program loading instruction, flow graph program base address, and flow graph program base address valid flag bit corresponding to the computing task to be executed can be expressed as: void Kernel_Start(bool inst_reload, uint32_t base_addr, uint32_t base_addr_valid,...), where inst_reload indicates whether to reload the flow graph program, and 0 can be set to indicate that there is no need to reload the flow graph program, and 1 can be set to indicate that the flow graph program needs to be reloaded; base_addr represents the flow graph program base address; base_addr_valid represents the flow graph program base address valid flag bit, which corresponds to multiple operands of the computing task one by one from low to high.
[0043] According to an embodiment of the present invention, when configuring the flow graph program base address corresponding to the computing task to be executed in the control core, the flow graph program base address corresponding to the first executed computing task among multiple computing tasks sharing the same flow graph program is configured to 0. It should be noted that in order to ensure that the storage base address of the intermediate result data that needs to be reused can be accessed during subsequent flow graph program iterations, the flow graph program base address corresponding to the first executed computing task among multiple computing tasks sharing the same flow graph program must be configured to 0, so as to ensure that the storage base address of the intermediate result data that needs to be reused in the on-chip cache module always points to the same address space.
[0044] According to an embodiment of the present invention, when configuring the flow graph program base address valid flag bit corresponding to the computing task to be executed, the number of bits of the flow graph program base address valid flag bit corresponding to the computing task to be executed is made to correspond one by one to the number of operands in the computing task to be executed, and each bit value of the flow graph program base address valid flag bit corresponding to the computing task to be executed is 0 or 1, and one value indicates that the storage offset value of the operand corresponding to this bit does not need to be accumulated with the flow graph program base address corresponding to the computing task to be executed, and the other value indicates that the storage offset value of the operand corresponding to this bit needs to be accumulated with the flow graph program base address corresponding to the computing task to be executed.
[0045] According to an embodiment of the present invention, when the control core configures the valid flag bit of the base address of the flow graph program corresponding to the current computing task, the bit corresponding to the operand to be reused in the current computing task in the valid flag bit of the base address of the flow graph program corresponding to the current computing task is set to a value that does not need to be accumulated with the base address of the flow graph program corresponding to the current computing task, and the storage offset value of the operand to be reused in the current computing task is made consistent with the storage offset value of the same operand to be reused in other computing tasks. It should be noted that the valid flag bit of the base address of the flow graph program is used to indicate whether the base address of the flow graph program allocated to the computing task to be executed is valid. When the value set in a bit of the valid flag bit of the base address of the flow graph program indicates that the storage offset value of the operand corresponding to this bit does not need to be accumulated with the base address of the flow graph program corresponding to the computing task to be executed, it means that the base address of the flow graph program allocated to the computing task to be executed is invalid, and the data corresponding to the storage offset value of the operand corresponding to this bit is directly addressed on the on-chip cache module; when the value set in a bit of the valid flag bit of the base address of the flow graph program indicates that the storage offset value of the operand corresponding to this bit needs to be accumulated with the base address of the flow graph program corresponding to the computing task to be executed, it means that the base address of the flow graph program allocated to the computing task to be executed is valid, and the value obtained by adding the storage offset value of the operand corresponding to this bit and the base address of the flow graph program corresponding to the computing task should be used to address the corresponding data on the on-chip cache module; through such an addressing method, during the execution of the computing task, it is possible to address the intermediate result data calculated by the previous computing task on the on-chip cache module, and it is also possible to address the data required during the execution of the computing task to be executed in the on-chip cache module, so as to support usage modes such as single-cache, dual-cache, or multi-cache of the on-chip cache module.
[0046] It should be noted that in the actual application process, different computing tasks correspond to different base addresses of the flow graph program and valid flag bits of the base address of the flow graph program. Therefore, during each execution of a computing task, the control core needs to configure the corresponding base address of the flow graph program and the valid flag bit of the base address of the flow graph program according to each computing task.
[0047] II. Matrix multiplication calculation example
[0048] To better understand the present invention, take Figure 5 the matrix multiplication calculation example shown as an example to illustrate how to execute a computing task according to the acceleration device and the computing task acceleration method proposed by the present invention.
[0049] As Figure 5As shown, it shows the calculation process of matrix multiplication of matrix A and matrix B, where C = A × B. Since the data scales of matrix A and matrix B are large, matrix A and matrix B are divided into 9 sub-matrices, that is, A = {A1, A2, A3, A4, A5, A6, A7, A8, A9}, B = {B1, B2, B3, B4, B5, B6, B7, B8, B9}. After the division, matrix multiplication of matrix A and matrix B is calculated to obtain the result matrix C = {C1, C2, C3, C4, C5, C6, C7, C8, C9}. Among them, C1 = A1 × B1 + A2 × B4 + A3 × B7, C2 = A1 × B2 + A2 × B5 + A3 × B8, C3 = A1 × B3 + A2 × B6 + A3 × B9, C4 = A4 × B1 + A5 × B4 + A6 × B7, C5 = A4 × B2 + A5 × B5 + A6 × B8, C6 = A4 × B3 + A5 × B6 + A6 × B9, C7 = A7 × B1 + A8 × B4 + A9 × B7, C8 = A7 × B2 + A8 × B5 + A9 × B8, C9 = A7 × B3 + A8 × B6 + A9 × B9.
[0050] Among them, C1 = A1 × B1 + A2 × B4 + A3 × B7. The calculation process of C1 can be abstracted into three calculation tasks, that is, C1 = C1_1 + C1_2 + C1_3, C1_1 = A1 × B1, C1_2 = A2 × B4, C1_3 = A3 × B7. At this time, each calculation task corresponds to 3 operands (A, B, C), each calculation task corresponds to a flow graph program, and each calculation task shares the same flow graph program. The following takes the calculation process of C1 as an example to illustrate the execution of the calculation task using the acceleration device proposed by the present invention. The calculation processes of other result matrices are the same as that of C1, and will not be elaborated here.
[0051] When performing the first calculation task C1_1 = A1 × B1, first configure the flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current calculation task in the control core and transmit them to the microcontroller. Among them, since it is the first time to perform a calculation task and the corresponding flow graph program needs to be loaded, the flow graph program loading instruction corresponding to the current calculation task is configured to need to load the flow graph program, and at this time the flow graph program executes C* = C* + A* × B*; the flow graph program base address corresponding to the current calculation task is Tile1_base_addr, and the value of the current flow graph program base address is configured to be 0; the number of bits of the flow graph program base address valid flag bit corresponding to the current calculation task corresponds one by one to the operands in the current calculation task, corresponding to the operands C, B, A from low to high one by one. And since the value of the flow graph program base address corresponding to the current calculation task is configured to be 0, the flow graph program base address valid flag bit corresponding to the current calculation task can be configured to any value between 000 and 111, and at this time it is configured to 000. Then, based on the received flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current calculation task, the microcontroller starts the calculation logic unit to calculate the storage base addresses of multiple operands in the current calculation task in the on-chip cache module, starts the execution array to execute the flow graph program to complete the current calculation task, and stores the obtained result matrix C1_1 in the corresponding address space of the on-chip cache module. Among them, the flow graph program base address corresponding to the current calculation task being 0 means that the starting address of the cache of the current calculation task on the on-chip cache module is 0; the flow graph program base address valid flag bit corresponding to the current calculation task being 000 means that the storage offset value of each operand in the current calculation task does not need to be accumulated with the flow graph program base address corresponding to the current calculation task, that is, each operand's data is addressed in the on-chip cache module according to the storage offset value of each operand in the current calculation task. In this calculation task, the storage offset value of each operand is: [A_offset, B_offset, C_offset] -> [0, A’, A’ + B’]. At this time, the storage offset value of operand A is 0, which means starting to address the data in matrix A1 from address 0 in the on-chip cache module, the storage offset value of operand B is A’, which means starting to address the data in matrix B1 from address A’ in the on-chip cache module, and the storage offset value of operand C is A’ + B’, which means using the address A’ + B’ of the on-chip cache module as the storage base address of the result matrix C1_1. It should be noted that the above content is described by taking the example that the calculation task C1_1 can be completed by executing the flow graph program only once. However, in the actual calculation process, due to the limitation of the storage resources of the execution unit, it may be necessary to iteratively execute the flow graph program multiple times to complete the calculation task C1_1.When it is necessary to iteratively execute the flow graph program multiple times to complete the computing task C1_1, it is also necessary to combine the storage offset value (iteration base address) of each operand in each iteration process to address the corresponding partial data of each operand during each iterative calculation. For example, the base address of the flow graph program corresponding to the current computing task is Tile1_base_addr = 0, and the valid flag bit of the base address of the flow graph program corresponding to the current computing task is configured as 000. During the first iteration process, when the storage offset value of each operand is [0, A_offset, B_offset, C_offset] -> [0, 0, A’, A’ + B’], data in matrix A1 is addressed starting from address 0 from the on-chip cache module for the first iterative calculation, data in matrix B1 is addressed starting from address A’ from the on-chip cache module for the first iterative calculation, and the address A’ + B’ of the on-chip cache module is used as the storage base address of the result during the first iterative calculation; when the storage offset value of each operand during the second iteration process is [1, A_offset, B_offset, C_offset] -> [1, a1, A’ + a1, A’ + B’], data in matrix A1 is addressed starting from address a1 from the on-chip cache module for the second iterative calculation, data in matrix B1 is addressed starting from address A’ + a1 from the on-chip cache module for the second iterative calculation, and the address A’ + B’ of the on-chip cache module is still used as the storage base address of the result during the second iterative calculation. The addressing process principle in other iterative processes is the same as the foregoing, and will not be elaborated here. It should be noted that the storage offset value (iteration base address) of each operand in each iteration process is statically generated by the compiler and stored in the address table in the on-chip cache or dynamically calculated by the microcontroller before each iteration execution.
[0052] When performing the second computational task C1_2 = A2 × B4, first configure the flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current computational task in the control core and transmit them to the microcontroller. Among them, since the same flow graph program is shared with the first computational task, there is no need to load the corresponding flow graph program, that is, the flow graph program loading instruction corresponding to the current computational task is configured not to load the flow graph program; the flow graph program base address corresponding to the current computational task is Tile2_base_addr, and the value of the current flow graph program base address is configured as A’ + B’ + C1’, which means that the starting address of the cache for the current computational task is the size of the storage space occupied by the previous computational task; the number of bits of the flow graph program base address valid flag bit corresponding to the current computational task corresponds one by one to the operands in the current computational task, corresponding to the operands C, B, A from low to high, and the flow graph program base address valid flag bit corresponding to the current computational task can be configured as 011. Then, based on the received flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current computational task, the microcontroller starts the computational logic unit to calculate the storage base addresses of multiple operands in the current computational task in the on-chip cache module, and starts the execution array to execute the flow graph program to complete the current computational task, and stores the obtained result matrix value C1_1 + C1_2 in the corresponding address space of the on-chip cache module. Among them, the flow graph program base address corresponding to the current computational task is A’ + B’ + C1’, which means that the starting address of the cache for the current computational task on the on-chip cache module is A’ + B’ + C1’; the flow graph program base address valid flag bit corresponding to the current computational task is 011, which means that the storage offset value of the operand C in the current computational task does not need to be accumulated with the flow graph program base address corresponding to the current computational task, while the operands A and B in the current computational task need to be accumulated with the flow graph program base address corresponding to the current computational task; the storage offset value of each operand in the current computational task is the same as the storage offset value of each operand in the first computational task. At this time, the storage offset value of the operand A is still 0, which means that the data in matrix A2 is addressed starting from the on-chip cache module at the address A’ + B’ + C1’ + 0, the storage offset value of the operand B is still A’, which means that the data in matrix B4 is addressed starting from the on-chip cache module at the address A’ + B’ + C1’ + A, and the storage offset value of the operand C is still A’ + B’, which means that the storage base address of the result matrix C1_2 is still the address A’ + B’ of the on-chip cache module (the storage offset value of the operand C is always A’ + B’ and is not accumulated with the flow graph program base address corresponding to the current computational task, so that the finally addressed storage base address is always consistent with the storage base address of C1_1, and the address space of the result matrix C1_1 in the on-chip cache module in the first computational task is reused).
[0053] When executing the third computing task C1_3 = A3 × B7, first configure the flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current computing task in the control core. Among them, since it shares the same flow graph program with the first computing task, there is no need to load the corresponding flow graph program either. That is, the flow graph program loading instruction corresponding to the current computing task is configured not to load the flow graph program; the flow graph program base address corresponding to the current computing task is Tile3_base_addr, and the value of the current flow graph program base address is configured as A'+B'+C1'+A'+B', which means that the flow graph program base address corresponding to the current computing task uses the storage space size occupied by the previous two computing tasks as the starting address of the cache for the current computing task; the number of bits of the flow graph program base address valid flag bit corresponding to the current computing task corresponds one-to-one with the operands in the current computing task, corresponding to the operands C, B, A from low to high, and the flow graph program base address valid flag bit corresponding to the current computing task can be configured as 011. Then, based on the received flow graph program loading instruction, flow graph program base address, and flow graph program base address valid flag bit corresponding to the current computing task, the microcontroller starts the computing logic unit to calculate the storage base address of multiple operands in the current computing task in the on-chip cache module, starts the execution array to execute the flow graph program to complete the current computing task, and stores the obtained result matrix value C1_1 + C1_2 + C1_3 into the corresponding address space of the on-chip cache module.Among them, the base address of the flow graph program corresponding to the current computing task is A'+B'+C1'+A'+B', which means that the starting address of the current computing task cached in the on-chip cache module is A'+B'+C1'+A'+B'; the valid flag bit of the base address of the flow graph program corresponding to the current computing task is 011, which means that the storage offset value of the operand C in the current computing task does not need to be accumulated with the base address of the flow graph program corresponding to the current computing task, while the operands A and B in the current computing task need to be accumulated with the base address of the flow graph program corresponding to the current computing task; the storage offset value of each operand in the current computing task is the same as that of each operand in the first computing task. At this time, the storage offset value of the operand A is still 0, which means that the data in the matrix A3 is addressed starting from the on-chip cache module at the address A'+B'+C1'+A'+B'+0, the storage offset value of the operand B is still A', which means that the data in the matrix B4 is addressed starting from the on-chip cache module at the address A'+B'+C1'+A'+B', and the storage offset value of the operand C is still A'+B', which means that the storage base address of the result matrix C1_3 is still the address A'+B' of the on-chip cache module (the storage offset value of the operand C is always A'+B' and is not accumulated with the base address of the flow graph program corresponding to the current computing task, so that the finally addressed storage base address is always the same as the storage base address of C1_1, and the address space of the result matrix C1_1 in the first computing task in the on-chip cache module is reused).
[0054] To more intuitively understand the storage allocation of each computing task in the on-chip cache module during the C1 calculation process, the storage allocation of each computing task shown in Figure 6 is further described. From Figure 6It can be known that for the first computing task, the base address of the flow graph program corresponding to the on-chip cache module is Tile1_base_addr = 0. The data in matrix A1 has an addressing address of 0 in the on-chip cache module, the data in matrix B1 has an addressing address of 0 + A' in the on-chip cache module, and the addressing address space of the result matrix C1_1 in the on-chip cache module is 0 + A' + B'. For the second computing task, the base address of the flow graph program corresponding to the on-chip cache module is Tile2_base_addr = A' + B' + C1'. The data in matrix A2 has an addressing address of A' + B' + C1' + 0 in the on-chip cache module, the data in matrix B4 has an addressing address of A' + B' + C1' + A' in the on-chip cache module, and the addressing address space of the result matrix C1_2 in the on-chip cache module is still 0 + A' + B'. For the third computing task, the base address of the flow graph program corresponding to the on-chip cache module is Tile3_base_addr = A' + B' + C1' + A' + B'. The data in matrix A3 has an addressing address of A' + B' + C1' + A' + B' + 0 in the on-chip cache module, the data in matrix B7 has an addressing address of A' + B' + C1' + A' + B' + A' in the on-chip cache module, and the addressing address space of the result matrix C1_3 in the on-chip cache module is still 0 + A' + B'.
[0055] It should be noted that, according to the foregoing content, it can be known that the acceleration device proposed based on the present invention can store the intermediate result data to be reused in the same storage base address space of the on-chip storage module when executing a computing task, so that the storage base address of the intermediate result data to be reused can be accessed during subsequent process program iterations, thereby reducing the data transfer overhead and the flow graph program switching overhead, and improving the execution efficiency of the computing task. Taking the address access mode of the fixed cache partition supported by the existing coarse-grained data flow architecture as an example, still taking the calculation process of C1 as an example, there are two ways to perform calculations when executing the C1 calculation process. One is to abstract the C1 calculation process into four calculation processes, corresponding to these four calculation tasks C1_1 = A1 × B1, C1_2 = A2 × B4, C1_3 = A3 × B7, C1 = C1_1 + C1_2 + C1_3 respectively. Among them, the first three calculation tasks share the same flow graph program C* = A*x B*, and the last calculation task needs to switch the flow graph program to execute C1 = C1_1 + C1_2 + C1_3. This calculation method not only needs to transfer the result matrices C1_1, C1_2, C1_3 obtained from the previous calculation tasks from the main memory during the calculation process, but also needs to restart the flow graph program to execute the calculation C1 = C1_1 + C1_2 + C1_3, resulting in a large amount of data transfer overhead and flow graph program switching overhead during the calculation process; the other method is to abstract the C1 calculation process into three calculation processes, corresponding to C1_1 = 0 + A1 × B1, C1_2 = C1_1 + A2 × B4, C1_3 = C1_2 + A3 × B7 respectively. Among them, each calculation task shares the same flow graph program C* = C* + A*x B*. Although this calculation method does not need to restart the flow graph program to execute the calculation C1 = C1_1 + C1_2 + C1_3, it still needs to continuously obtain the previously calculated result matrices C1_1 and C1_2 from the main memory during the calculation process, resulting in a large amount of data transfer overhead during the calculation process.
[0056] The beneficial effects of the present invention are as follows: It can support the reuse of intermediate result data of multiple computing tasks when multiple computing tasks share the same flow graph program, reduce the data transfer overhead and the flow graph program switching overhead, and improve the execution efficiency of the computing task; it can also flexibly allocate the storage space of the on-chip cache module to support usage modes such as single cache, dual cache or multi-cache.
[0057] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order, as long as the required functions can be achieved.
[0058] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0059] A computer-readable storage medium may be a tangible device that stores and holds instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing.
[0060] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the field to understand the embodiments disclosed herein.
Claims
1. A method for accelerating a computing task based on a coarse-grained data flow architecture, where the computing task corresponds to multiple operands, and the computing task corresponds to a flow graph program, and the flow graph program is executed based on the operands to complete the computing task, characterized in that, The method includes: S1. Configure a flow graph program loading instruction, a flow graph program base address, and a flow graph program base address valid flag bit corresponding to the current computing task. Among them, the flow graph program loading instruction corresponding to the current computing task is used to indicate whether it is necessary to reload the flow graph program when executing the current computing task, and the flow graph program is not reloaded when sharing the flow graph program with the previous computing task; the flow graph program base address corresponding to the current computing task is used to indicate the starting address of the cache of the current computing task; the flow graph program base address valid flag bit corresponding to the current computing task is used to indicate whether the storage offset value of each operand in the current computing task needs to be accumulated with the flow graph program base address corresponding to the current computing task to address the data of each operand. S2. Based on the flow graph program loading instruction, the flow graph program base address, and the flow graph program base address valid flag bit corresponding to the current computing task configured in step S1, obtain the storage base addresses of multiple operands in the current computing task and execute the flow graph program to complete the current computing task.
2. The method according to claim 1, wherein The number of bits of the flow graph program base address valid flag bit corresponding to the current computing task corresponds one-to-one with the number of operands in the current computing task, and the value of each bit in the flow graph program base address valid flag bit corresponding to the current computing task is 0 or 1, and one value indicates that the storage offset value of the operand corresponding to this bit does not need to be accumulated with the flow graph program base address corresponding to the current computing task, and the other value indicates that the storage offset value of the operand corresponding to this bit needs to be accumulated with the flow graph program base address corresponding to the current computing task.
3. The method according to claim 2, wherein In step S1, when configuring the flow graph program base address corresponding to the current computing task, configure the flow graph program base address corresponding to the first executed computing task among multiple computing tasks sharing the same flow graph program to 0.
4. The method according to claim 3, wherein In step S1, when configuring the flow graph program base address valid flag bit corresponding to the current computing task, set the bit corresponding to the operand that needs to be reused in the current computing task in the flow graph program base address valid flag bit corresponding to the current computing task to the value indicating that it does not need to be accumulated with the flow graph program base address corresponding to the current computing task, and make the storage offset value of the operand that needs to be reused in the current computing task consistent with the storage offset value of the same operand that needs to be reused in other computing tasks.
5. A computing task acceleration device for a coarse-grained data flow architecture, where the computing task corresponds to multiple operands, and there is a flow graph program corresponding to the computing task. The flow graph program is executed based on the operands to complete the computing task, and is characterized in that, The device is configured to process computing tasks according to the method described in any one of claims 1-4.
6. The device according to claim 5, wherein The device includes a control core, a microcontroller, an execution array, and an on-chip cache module, where: The control core is used to configure the flow graph program loading instruction, the base address of the flow graph program, and the valid flag bit of the base address of the flow graph program corresponding to the computing task and transmit them to the microcontroller. Among them, the flow graph program loading instruction corresponding to the computing task is used to indicate whether the flow graph program needs to be reloaded when executing the computing task. When sharing the flow graph program with the previous computing task, the flow graph program is not reloaded; the base address of the flow graph program corresponding to the computing task is used to indicate the starting address of the computing task in the on-chip cache module; the valid flag bit of the base address of the flow graph program corresponding to the computing task is used to indicate whether the storage offset value of each operand in the computing task needs to be accumulated with the base address of the flow graph program corresponding to the computing task to address the data of each operand on the on-chip cache module; and it is used to send a start execution message to the microcontroller. The microcontroller is used to load the flow graph program based on the flow graph program loading instruction corresponding to the computing task transmitted by the control core, and is used to store the base address of the flow graph program corresponding to the computing task and the valid flag bit of the base address of the flow graph program transmitted by the control core, and calculate the storage base addresses of multiple operands on the on-chip cache module based on them, and control the execution array to execute the flow graph program based on the received start execution message. The execution array is used to execute the flow graph program based on the data of each operand addressed by the microcontroller on the on-chip cache module to complete the computing task. The on-chip cache module is used to store the data corresponding to multiple operands respectively when executing the computing task and the data obtained after the execution array executes the flow graph program based on the operands.
7. The device according to claim 6, characterized in that, The microcontroller is configured with a storage module and a computing logic unit. Among them, the storage module is used to store the base address of the flow graph program corresponding to the computing task and the valid flag bit of the base address of the flow graph program transmitted by the control core; the computing logic unit is used to calculate the storage base addresses of multiple operands on the on-chip cache module based on the base address of the flow graph program corresponding to the current computing task and the valid flag bit of the base address of the flow graph program stored in the storage unit.
8. The device according to claim 7, characterized in that, The storage module includes: A base address storage unit of the flow graph program, which is used to store the base address of the flow graph program corresponding to the computing task transmitted by the control core; A valid flag bit storage unit of the base address of the flow graph program, which is used to store the valid flag bit of the base address of the flow graph program corresponding to the computing task transmitted by the control core.
9. The device according to claim 8, characterized in that, The device is configured to execute the computing task in the following manner: Configure the flow graph program loading instruction, the base address of the flow graph program, and the valid flag bit of the base address of the flow graph program corresponding to the current computing task in the control core and transmit them to the microcontroller; Based on the received flow graph program loading instruction, the base address of the flow graph program, and the valid flag bit of the base address of the flow graph program corresponding to the current computing task, the microcontroller starts the computing logic unit to calculate the storage base addresses of multiple operands in the current computing task on the on-chip cache module, and starts the execution array to execute the flow graph program to complete the current computing task.
10. The device according to claim 9, characterized in that, When the control core configures the base address of the flow graph program corresponding to the current computing task, the base address of the flow graph program corresponding to the first executed computing task among multiple computing tasks sharing the same flow graph program is configured to 0.
11. The device according to claim 10, characterized in that, When the control core configures the valid flag bit of the base address of the flow graph program corresponding to the current computing task, the number of bits of the valid flag bit of the base address of the flow graph program corresponding to the current computing task corresponds one-to-one with the number of operands in the current computing task, and the value of each bit in the valid flag bit of the base address of the flow graph program corresponding to the current computing task is 0 or 1. One value indicates that the storage offset value of the operand corresponding to this bit does not need to be accumulated with the base address of the flow graph program corresponding to the current computing task, and the other value indicates that the storage offset value of the operand corresponding to this bit needs to be accumulated with the base address of the flow graph program corresponding to the current computing task.
12. The device according to claim 11, characterized in that, When the control core configures the valid flag bit of the base address of the flow graph program corresponding to the current computing task, the bit corresponding to the operand that needs to be reused in the current computing task in the valid flag bit of the base address of the flow graph program corresponding to the current computing task is set to the value indicating that it does not need to be accumulated with the base address of the flow graph program corresponding to the current computing task, and the storage offset value of the operand that needs to be reused in the current computing task is kept consistent with the storage offset value of the same operand that needs to be reused in other computing tasks.
13. A computer-readable storage medium, characterized in that, On which a computer program is stored, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1 to 4.
14. An electronic device, characterized in that, Comprising: One or more processors; And a memory for storing executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1 to 4 by executing the executable instructions stored in the memory.