Task Processing Method, Device, Equipment and Medium Based on Data Flow Core Architecture
By adopting the cross-tree interconnection method and task decomposition method in the core architecture of data flow, the problem of low data processing task efficiency and load balancing in the object detection deep neural network model is solved, and more efficient task processing and load balancing are achieved.
Patent Information
- Application Number
- CN202510337391.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In the core architecture of data flow, how to improve the efficiency and load balancing of data processing tasks of the deep neural network model of object detection, especially in matrix multiplication tasks, there are problems such as insufficient flexibility and low utilization of computing units.
The cross-tree interconnection method is used to interconnect the data stream cores, and the data processing tasks in the target detection model are decomposed according to the cache area of each data stream core. After the decomposition, the task is broadcast between the data stream cores through external product, and the processing results are finally spliced and written back to DRAM.
By decomposing tasks and cross-tree interconnection, the load on each data stream core is reduced, load balancing and communication efficiency is improved, the number of accesses to DRAM and access pressure is reduced, and the overall task completion efficiency is improved.
Smart Images

Figure CN119861974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-performance computing technology, and particularly to a task processing method, device, equipment and medium based on a data flow core architecture. Background Art
[0002] The data flow core architecture and the instruction set architecture are two main ways of chip implementation. Different from the instruction set architecture, the data flow core architecture does not need to use instructions to control the specific order of data transfer and calculation. The core of the data flow architecture is driven by the availability of data. When the input data is ready, the computing unit immediately executes, and after being ready, it will immediately flow into the next pipeline computing unit, avoiding resource idleness caused by instruction dependencies in the instruction flow architecture and simplifying the control logic overhead of instructions. Due to the characteristic of being driven by data availability, which is very suitable for large-scale parallel computing, AI (Artificial Intelligence) custom accelerators such as TPU (Tensor Processing Unit) adopt this architecture. However, the single large computing array adopted on chips such as TPU has problems such as insufficient flexibility and low utilization rate of computing units. To solve this problem, the data flow core architecture came into being. This architecture effectively solves the above problems by integrating multiple processing cores, deploying data flow operation units in each core, and using instruction control to achieve coarse-grained task allocation and data exchange between cores.
[0003] In the deep neural network in the field of object detection, matrix multiplication operations are performed between the input vector matrix and the weight matrix of the fully connected layer for feature transformation. For the convolutional layer for feature extraction, its convolutional operation can also be transformed into a matrix multiplication operation. That is to say, the matrix multiplication task is the core operation task in the data processing tasks of the object detection deep neural network model, and its operation has high parallelism and locality. Therefore, there are still many challenges in how to run these operations on the data flow core architecture.
[0004] In summary, how to improve the efficiency and load balancing degree of the data flow core architecture to complete data processing tasks is a problem to be solved in this field. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a task processing method, device, equipment and medium based on a data flow core architecture, so as to improve the efficiency and load balancing degree of the data flow core architecture to complete data processing tasks. The specific solutions are as follows:
[0006] In the first aspect, the present application discloses a task processing method based on a data flow core architecture, and the interconnection mode between the data flow cores in the data flow core architecture is a cross-tree interconnection mode; the method includes:
[0007] Decompose the data processing tasks in the target detection model obtained from the DRAM according to the cache area of each data stream core to obtain the decomposed data processing tasks; the data processing tasks include matrix multiplication tasks; the target detection model is a target detection model constructed based on a neural network.
[0008] Broadcast the decomposed data processing tasks to each data stream core based on the cross-tree interconnection method to obtain the decomposition task processing results output by the data stream core through the outer product method.
[0009] Stitch together the decomposition task processing results of each data stream core to obtain the target processing result, and write the target processing result back to the DRAM.
[0010] Optionally, each data stream core is respectively a root node, a relay node, and a leaf node, and the relay node is connected to the root node, and the leaf node is connected to the corresponding relay node.
[0011] Optionally, the task processing method based on the data stream core architecture further includes:
[0012] Obtain the two-dimensional grid of the data stream core architecture, and determine the starting data stream core in the two-dimensional grid as the root node;
[0013] Determine the target row and target column where the root node is located in the two-dimensional grid, and determine the data stream cores other than the root node in the target row and the target column as the relay nodes;
[0014] Determine the data stream cores other than the root node and the relay nodes in the two-dimensional grid as the leaf nodes.
[0015] Optionally, the relay node includes a first relay node and a second relay node, where the first relay node is located in the target row, and the second relay node is located in the target column;
[0016] Correspondingly, broadcasting the decomposed data processing tasks to each data stream core based on the cross-tree interconnection method includes:
[0017] Input the left matrix and the right matrix in the decomposed data processing tasks into the root node, the first relay node, and the second relay node respectively; where the left matrix of the first relay node is the same as the left matrix of the root node, and the right matrix of the second relay node is the same as the right matrix of the root node.
[0018] Broadcast the right matrix to the corresponding leaf nodes through the first relay node, and broadcast the left matrix to the corresponding leaf nodes through the second relay node.
[0019] Optionally, the broadcasting of the decomposed data processing tasks to each of the data stream cores based on the cross-tree interconnection method includes:
[0020] Control the leaf nodes to send ready signals to the corresponding relay nodes to update the ready signal quantity of the relay nodes;
[0021] When the ready signal quantity of the relay node is equal to the number of the leaf nodes corresponding to the relay node, reset the ready signal quantity of the relay node;
[0022] Control the relay nodes to broadcast the decomposed data processing tasks and transmission completion signals to the corresponding leaf nodes to update the completion signal quantity of the leaf nodes;
[0023] When the completion signal quantity of the leaf node meets a preset threshold, determine that the leaf node is in a ready state, reset the completion signal quantity of the leaf node, and re-jump to the step of controlling the leaf node to send a ready signal to the corresponding relay node.
[0024] Optionally, the data stream core includes a computing unit and an SRAM, and the computing unit includes a destination register;
[0025] Correspondingly, the data stream core outputs the decomposition task processing result by the outer product method, including:
[0026] The data stream core divides the received decomposed data processing tasks to obtain each subtask, loads each subtask into the SRAM, controls the computing unit to obtain each subtask from the SRAM and obtain the processing result of each subtask by the outer product method, sequentially writes the processing results of each subtask into the destination register, and accumulates the processing results of each subtask in the destination register to obtain the decomposition task processing result.
[0027] Optionally, the SRAM includes a first cache area and a second cache area;
[0028] Correspondingly, the loading of each subtask into the SRAM, controlling the computing unit to obtain each subtask from the SRAM and obtain the processing result of each subtask by the outer product method includes:
[0029] Alternately determine the first cache area and the second cache area as the current cache area;
[0030] Load the current subtask into the current cache area;
[0031] After the computing unit obtains the processing result of the previous subtask through the outer product method, control the computing unit to obtain the current subtask from the current cache area and obtain the processing result of the current subtask through the outer product method.
[0032] In a second aspect, the present application discloses a task processing device based on a data flow core architecture, and the interconnection method between data flow cores in the data flow core architecture is a cross-tree interconnection method; the device includes:
[0033] A task decomposition module, configured to decompose the data processing task in the target detection model obtained from the DRAM according to the cache area of each data flow core to obtain the decomposed data processing task; the data processing task includes a matrix multiplication task; the target detection model is a target detection model constructed based on a neural network;
[0034] A result acquisition module, configured to broadcast the decomposed data processing task to each data flow core based on the cross-tree interconnection method to obtain the data processing result output by the data flow core through the outer product method;
[0035] A result splicing module, configured to splice each data processing result to obtain a target processing result and write the target processing result back to the DRAM.
[0036] In a third aspect, the present application discloses an electronic device, including:
[0037] A memory, configured to store a computer program;
[0038] A processor, configured to execute the computer program to implement the steps of the foregoing disclosed task processing method based on a data flow core architecture.
[0039] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing disclosed task processing method based on a data flow core architecture are implemented.
[0040] The beneficial effects of this application are as follows: In the data stream core architecture of this application, the interconnection method between data stream cores is a cross-tree interconnection method; the data processing tasks in the target detection model obtained from DRAM are decomposed according to the cache area of each data stream core to obtain the decomposed data processing tasks; the data processing tasks include matrix multiplication tasks; the target detection model is a target detection model constructed based on a neural network; based on the cross-tree interconnection method, the decomposed data processing tasks are broadcast to each data stream core to obtain the data processing results output by the data stream cores through the outer product method; the data processing results are spliced to obtain the target processing result, and the target processing result is written back to the DRAM. It can be seen that in this application, the data processing tasks in the target detection model obtained from DRAM are decomposed according to the cache area of each data stream core. Then, the decomposed data processing tasks that each data stream core needs to process are smaller than the original data processing tasks. Therefore, the load of each data stream core is reduced, making the load balanced. Further, the interconnection method between data stream cores in the data stream core architecture is a cross-tree interconnection method. Broadcasting the decomposed data processing tasks to each data stream core based on the cross-tree interconnection method can improve communication efficiency, and broadcasting the decomposed data processing tasks can reduce the number of accesses to DRAM and the memory access pressure. Multiple data stream cores process the decomposed data processing tasks in parallel to improve the task completion efficiency. Finally, the processing results of each decomposed task are spliced to obtain the target processing result, and the target processing result is written back to the DRAM, that is, the data processing task is completed. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0042] Figure 1 It is a flowchart of a task processing method based on a data stream core architecture disclosed in this application;
[0043] Figure 2 It is a specific node connection diagram disclosed in this application;
[0044] Figure 3 It is a specific two-dimensional grid diagram disclosed in this application;
[0045] Figure 4 It is a specific task processing diagram disclosed in this application;
[0046] Figure 5 A specific schematic diagram of inter-core broadcast of data stream disclosed in this application;
[0047] Figure 6 A specific schematic diagram of single data stream core task processing disclosed in this application;
[0048] Figure 7 A schematic diagram of the structure of a task processing device based on the data stream core architecture disclosed in this application;
[0049] Figure 8 A structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] The data stream core architecture and the instruction set architecture are two main ways of chip implementation. Different from the instruction set architecture, the data stream core architecture does not need to use instructions to control the specific order of data transfer and calculation. The core of the data stream architecture is driven by the availability of data. The computing unit immediately executes when the input data is ready, and will immediately flow into the next pipeline computing unit after being ready, avoiding resource idleness caused by instruction dependencies in the instruction stream architecture and simplifying the control logic overhead of instructions. Due to the characteristics driven by data availability being very suitable for large-scale parallel computing, AI custom accelerators such as TPU all adopt this architecture. However, the single large computing array adopted on chips such as TPU has problems such as insufficient flexibility and low utilization rate of computing units. To solve this problem, the data stream core architecture came into being. This architecture effectively solves the above problems by integrating multiple processing cores, deploying data stream operation units in each core, and using instruction control to achieve coarse-grained task allocation and data exchange between cores.
[0052] In the deep neural network in the field of object detection, matrix multiplication operations are performed between the input vector matrix and the weight matrix of the fully connected layer of feature transformation. For the convolutional layer of feature extraction, its convolutional operation can also be transformed into a matrix multiplication operation. That is to say, the matrix multiplication task is the core operation task in the data processing tasks of the object detection deep neural network model, and its operation has high parallelism and locality. Therefore, there are still many challenges in how to run these operations on the data stream core architecture.
[0053] Therefore, the present application correspondingly provides a data processing task processing solution to improve the efficiency and load balancing degree of the data flow core architecture in completing data processing tasks.
[0054] See Figure 1 As shown, an embodiment of the present application discloses a task processing method based on a data flow core architecture. The interconnection mode between each data flow core in the data flow core architecture is a cross-tree interconnection mode. The method includes:
[0055] Step S11: Decompose the data processing tasks in the target detection model obtained from the DRAM according to the cache area of each data flow core to obtain the decomposed data processing tasks. The data processing tasks include matrix multiplication tasks. The target detection model is a target detection model constructed based on a neural network.
[0056] In this embodiment, each data flow core is respectively a root node, a relay node, and a leaf node, and the relay node is connected to the root node, and the leaf node is connected to the corresponding relay node. For example Figure 2 As shown in a specific node connection schematic diagram, the root node is specifically the data flow core (0, 0), the relay nodes are the data flow cores (0, 2), the data flow core (0, 1), the data flow core (1, 0), the data flow core (2, 0), and the leaf nodes are the data flow core (2, 1), the leaf node is the data flow core (1, 2), the leaf node is the data flow core (2, 2), the leaf node is the data flow core (1, 1). Among them, all relay nodes are connected to the root node, and the leaf nodes are connected to the relay nodes in their own rows and columns. Therefore, the data flow core (2, 1) is connected to the data flow cores (0, 1) and (2, 0), the data flow core (1, 1) is connected to the data flow cores (0, 1) and (1, 0), the data flow core (1, 2) is connected to the data flow cores (0, 2) and (1, 0), and the data flow core (2, 2) is connected to the data flow cores (0, 2) and (2, 0). Among them, (x, y) of the data flow core respectively represents the coordinates in the two-dimensional grid of the data flow core architecture, that is, (x, y) represents the data flow core in the x-th row and y-th column.
[0057] In this embodiment, it further includes: obtaining the two-dimensional grid of the data flow core architecture, and determining the starting data flow core in the two-dimensional grid as the root node; determining the target row and target column where the root node is located in the two-dimensional grid, and determining the data flow cores other than the root node in the target row and the target column as the relay nodes; determining the data flow cores other than the root node and the relay nodes in the two-dimensional grid as the leaf nodes.
[0058] Further, it is necessary to pre-determine whether each data flow core belongs to a root node, a relay node, or a leaf node. Specifically, a two-dimensional grid of the data flow core architecture is obtained. For example Figure 3 As shown in a specific schematic diagram of a two-dimensional grid, the data flow core architecture in this embodiment includes multiple data flow cores. Among them, the starting data flow core in the two-dimensional grid is determined as the root node. The core located at the upper leftmost position in the two-dimensional grid of data flow cores serves as the first layer of the on-chip network tree structure, that is, the data flow core (0, 0) is the root node. Further, the target row and target column where the root node is located in the two-dimensional grid are determined, that is, the target row is 0 and the target column is 0. And the data flow cores other than the root node in the target row and target column are determined as relay nodes. The cores located in the uppermost row and the leftmost column except the leftmost upper core in the two-dimensional grid of data flow cores serve as the second layer of the on-chip network tree structure. Therefore, the relay nodes are the data flow core (0, 2), the data flow core (0, 1), the data flow core (1, 0), and the data flow core (2, 0). Next, the data flow cores other than the root node and relay nodes in the two-dimensional grid are determined as leaf nodes. The cores except the uppermost row and the leftmost column in the two-dimensional network of data flow cores serve as the third layer of the on-chip network tree structure. Therefore, the leaf nodes are the data flow core (2, 1), the leaf node is the data flow core (1, 2), the leaf node is the data flow core (2, 2), and the leaf node is the data flow core (1, 1).
[0059] It can be understood that the pre-determined root node (i.e., the first layer), relay node (i.e., the second layer), and leaf node (i.e., the third layer) are obtained. And each relay node is connected to the root node, and each leaf node is connected to the relay nodes in its row and column. Data can be directly transmitted between the connected nodes, and data can be transmitted between the unconnected nodes through multiple forwarding via the on-chip network. In this way, a cross-tree interconnection among multiple data flow cores is constructed. The cross-tree interconnection method avoids complex routing algorithms between nodes, and the interconnection method has the characteristics of low cost and strong scalability.
[0060] Such as Figure 3As shown in the figure, the data stream core architecture includes a DRAM (Dynamic Random Access Memory) interface. All data stream cores can access the data in the same DRAM through the DRAM interface. It also includes a PCIE (peripheral component interconnect express, a high-speed serial computer expansion bus standard) interface, which is responsible for transmitting data and instructions with the host. The data processing tasks in the target detection model obtained from the DRAM through the DRAM interface include matrix multiplication tasks. The matrix multiplication task can specifically be the task of performing matrix multiplication between the input vector matrix and the weight matrix in the fully connected layer of feature transformation, or it can also be the convolution operation task in the convolutional layer of feature extraction.
[0061] Decompose the data processing tasks according to the cache area of each data stream core to obtain the decomposed data processing tasks. Determine the amount of tasks contained in each decomposed data processing task according to the cache area of each data stream core. And each decomposed data processing is divided into a left matrix and a right matrix. Or rather, the entire data processing task is divided into a left matrix and a right matrix. The rows and columns of the left and right matrices constitute each decomposed data processing. For example Figure 4 As shown in a specific task processing schematic diagram, each row in the left matrix is x1, x2, x3 respectively, and each column in the right matrix is y1, y2, y3 respectively.
[0062] Step S12: Broadcast the decomposed data processing tasks to each of the data stream cores based on the cross-tree interconnection method to obtain the processing results of the decomposed tasks output by the data stream cores through the outer product method.
[0063] In this embodiment, the relay nodes include a first relay node and a second relay node. Among them, the first relay node is located in the target row, and the second relay node is located in the target column. Further, in order to more carefully distinguish the relay nodes, the relay nodes in the row where the root node is located can be determined as the first relay node, and the relay nodes in the column where the root node is located can be determined as the second relay node. That is to say, the first relay node is located in the target row, and the second relay node is located in the target column. As Figure 4 shown, data stream core (0, 2) and data stream core (0, 1) are the first relay nodes, and data stream core (1, 0) and data stream core (2, 0) are the second relay nodes.
[0064] In this embodiment, the broadcasting of the decomposed data processing tasks to each of the data stream cores based on the cross-tree interconnection method includes: inputting the left matrix and the right matrix in the decomposed data processing tasks into the root node, the first relay node, and the second relay node respectively; wherein, the left matrix of the first relay node is the same as the left matrix of the root node, and the right matrix of the second relay node is the same as the right matrix of the root node; broadcasting the right matrix to the corresponding leaf nodes through the first relay node, and broadcasting the left matrix to the corresponding leaf nodes through the second relay node.
[0065] As Figure 4 shown, input x1 of the left matrix and y1 of the right matrix in the decomposed data processing task into the root node, input x2 of the left matrix and y1 of the right matrix in the decomposed data processing task into the data stream core (1,0), input x3 of the left matrix and y1 of the right matrix in the decomposed data processing task into the data stream core (2,0), input x1 of the left matrix and y2 of the right matrix in the decomposed data processing task into the data stream core (0,1), input x1 of the left matrix and y3 of the right matrix in the decomposed data processing task into the data stream core (0,2). In this way, the relay nodes obtain the corresponding left and right matrices. It can be understood that the left matrix of the first relay node is the same as the left matrix of the root node, and the right matrix of the second relay node is the same as the right matrix of the root node; further, broadcast the right matrix to the corresponding leaf nodes through the first relay node, that is, broadcast the right matrix to the leaf nodes in its own column through the first relay node. Specifically, the data stream core (0,1) broadcasts y2 of the right matrix to the data stream cores (1,1) and (2,1), and the data stream core (0,2) broadcasts y3 of the right matrix to the data stream cores (1,2) and (2,2), and broadcast the left matrix to the corresponding leaf nodes through the second relay node, that is, broadcast the left matrix to the leaf nodes in its own row through the second relay node. Specifically, the data stream core (1,0) broadcasts x2 of the left matrix to the data stream cores (1,1) and (1,2), and the data stream core (2,0) broadcasts x3 of the left matrix to the data stream cores (2,1) and (2,2). In this way, each data stream core has received its corresponding decomposed data processing task.
[0066] In this embodiment, the process of broadcasting the decomposed data processing tasks to each of the data stream cores based on the cross-tree interconnection method includes: controlling the leaf nodes to send ready signals to the corresponding relay nodes to update the ready signal quantity of the relay nodes; when the ready signal quantity of the relay node is equal to the number of the leaf nodes corresponding to the relay node, reset the ready signal quantity of the relay node; controlling the relay node to broadcast the decomposed data processing tasks and transmission completion signals to the corresponding leaf nodes to update the completion signal quantity of the leaf nodes; when the completion signal quantity of the leaf node meets the preset threshold, it is determined that the leaf node is in a ready state, reset the completion signal quantity of the leaf node, and then jump back to the step of controlling the leaf node to send a ready signal to the corresponding relay node.
[0067] The relay node broadcasts the left matrix or the right matrix to the leaf nodes in the same row or the same column. During the broadcast process, the relay node acts as the sender, and the corresponding leaf node acts as the receiver. For example Figure 5 A specific schematic diagram of the broadcast between data stream cores is shown as follows. The specific process is as follows:
[0068] 1) Control the leaf nodes (i.e., the receivers) to send ready signals to the corresponding relay nodes (i.e., the senders) to update the ready signal quantity of the relay nodes. That is, for each ready signal received by the relay node, the ready signal quantity of the relay node is incremented by 1.
[0069] 2) When the ready signal quantity of the relay node is equal to the number of the leaf nodes corresponding to the relay node, reset the ready signal quantity of the relay node to 0. For example, if the relay node needs to broadcast the left matrix to two leaf nodes in the same row, when the ready signal quantity is 2, it means that all the leaf nodes to be broadcast have been prepared. Resetting the ready signal quantity of the relay node is for the next broadcast.
[0070] 3) Control the relay node to broadcast the decomposed data processing tasks and transmission completion signals to the corresponding leaf nodes to update the completion signal quantity of the leaf nodes. That is to say, for each transmission completion signal received by the leaf node, the completion signal quantity is incremented by 1.
[0071] 4) When the completion signal quantity of the leaf node meets the preset threshold, it indicates that all the leaf nodes have received the matrix to be received, that is, the current broadcast is completed. Then it is determined that the leaf node is in a ready state, reset the completion signal quantity of the leaf node to 0, and then jump back to the step of controlling the leaf node to send a ready signal to the corresponding relay node.
[0072] In this embodiment, the data stream core includes a computing unit and an SRAM, and the computing unit includes a destination register. As Figure 3As shown, each data flow core includes a computing unit, an SRAM (Static Random-Access Memory), and also includes a network-on-chip interface for data transmission with other data flow cores, where the computing unit includes destination registers.
[0073] In this embodiment, the data flow core outputs the decomposition task processing result by means of outer product, including: the data flow core divides the received decomposed data processing task to obtain each sub-task, loads each sub-task into the SRAM, controls the computing unit to obtain each sub-task from the SRAM and obtain the processing result of each sub-task by means of outer product, writes the processing results of each sub-task into the destination register in sequence, and accumulates the processing results of each sub-task in the destination register to obtain the decomposition task processing result.
[0074] The data flow core divides the received decomposed data processing task again to obtain each sub-task. Each sub-task can be specifically divided into B blocks. Each sub-task is loaded into the SRAM, the computing unit is controlled to obtain each sub-task from the SRAM, and the computing unit is controlled to perform an outer product on the sub-task to obtain the processing result of the sub-task; each time the processing result of a sub-task is obtained, the processing result of the sub-task is written into the destination register, and the processing results of each sub-task in the destination register are accumulated to obtain the decomposition task processing result. It should be noted that the decomposition task processing result is the processing result of the decomposed data processing task assigned to its data flow core, rather than the processing result of the entire data processing task.
[0075] In this embodiment, the SRAM includes a first cache area and a second cache area. Further, the decomposed data processing task assigned to the data flow core is divided into a left matrix and a right matrix. Among them, the size of the left matrix is M*N, and the size of the right matrix is K*N. The decomposed data processing task can be specifically divided into B blocks. Then, the size of the left matrix of each sub-task loaded each time is M*(K / B), and the size of the right matrix is (K / B)*N. Among them, the first cache area is used to store the relatively earlier left matrix sub-tasks and right matrix sub-tasks, and the second cache area is used to store the relatively later left matrix sub-tasks and right matrix sub-tasks. That is to say, the SRAM reserves a space of 2*M*(K / B) for the left matrix and a space of 2*(K / B)*N for the right matrix. In this way, a double buffer area is formed.
[0076] In this embodiment, the steps of loading each of the subtasks into the SRAM, controlling the computing unit to obtain each of the subtasks from the SRAM, and obtaining the processing results of each of the subtasks through the outer product method include: alternately determining the first cache area and the second cache area as the current cache area; loading the current subtask into the current cache area; after the computing unit obtains the processing result of the previous subtask through the outer product method, controlling the computing unit to obtain the current subtask from the current cache area and obtaining the processing result of the current subtask through the outer product method.
[0077] Since the size of the first cache region is twice the sub-task size of the left matrix and the size of the second cache region is twice the sub-task size of the right matrix, the first cache region and the second cache region are alternately determined as the current cache region, and then the current sub-task is loaded into the current cache region. After the processing result of the previous sub-task is obtained by the computing unit through the outer product method, that is, after the previous sub-task is processed, the computing unit is controlled to obtain the current sub-task from the current cache region and obtain the processing result of the current sub-task through the outer product method. For example, if the first cache region has currently loaded sub-task n and the second cache region has currently loaded sub-task n+1, the first cache region is used as the current cache region at this time, and the computing unit is controlled to obtain sub-task n from the first cache region, and then the computing unit is controlled to perform an outer product on sub-task n to obtain the processing result of sub-task n. Moreover, when the computing unit is controlled to obtain sub-task n from the first cache region, the first cache region is controlled to load sub-task n+2; after the processing result of sub-task n is obtained, the second cache region is used as the new current cache region, and the computing unit is controlled to obtain sub-task n+1 from the second cache region, and the computing unit is controlled to perform an outer product on sub-task n+1 to obtain the processing result of sub-task n+1. Moreover, when the computing unit is controlled to obtain sub-task n+1 of the right matrix from the second cache region, the second cache region is controlled to load sub-task n+3; next, after the processing result of sub-task n+1 is obtained, the first cache region is used as the new current cache region, and the computing unit is controlled to obtain sub-task n+2 of the left matrix from the first cache region, and the computing unit is controlled to perform an outer product on sub-task n+2 to obtain the processing result of sub-task n+2. Moreover, when the computing unit is controlled to obtain sub-task n+2 of the left matrix from the first cache region, the first cache region is controlled to load sub-task n+4, and so on; that is, it is loaded into the first half of the double buffer during the first load, and loaded into the second half of the double buffer during the second load. At the same time, the computing unit can also read and calculate the data in the first half of the double buffer. Each subsequent load and calculation switches the target position for operation in the double buffer at the same time, so that the loading and calculation of data can form a pipelined parallelism. In this way, the double buffer strategy ensures that the computing unit can directly obtain the sub-task each time, reduces the waiting time, and improves the data processing efficiency.
[0078] Step S13: Concatenate the processing results of the decomposed tasks to obtain the target processing result, and write the target processing result back to the DRAM.
[0079] Since the data processing tasks are decomposed so that each data processing core processes different decomposed data tasks, after all data processing cores output their respective decomposed task processing results, they need to be concatenated. Only in this way can the target processing result, that is, the processing result of the complete data processing task, be obtained and written back to the DRAM.
[0080] The beneficial effects of this application are as follows: In the data stream core architecture of this application, the interconnection method between the data stream cores is a cross-tree interconnection method; the data processing tasks in the target detection model obtained from the DRAM are decomposed according to the cache area of each data stream core to obtain decomposed data processing tasks; the target detection model is a target detection model constructed based on a neural network; the data processing tasks include matrix multiplication tasks; based on the cross-tree interconnection method, the decomposed data processing tasks are broadcast to each data stream core to obtain the data processing results output by the data stream cores through the outer product method; the data processing results are concatenated to obtain the target processing result and written back to the DRAM. It can be seen that in this application, the data processing tasks in the target detection model obtained from the DRAM are decomposed according to the cache area of each data stream core. Then, the decomposed data processing tasks that each data stream core needs to process are smaller than the original data processing tasks. Therefore, the load of each data stream core is reduced, making the load balanced. Further, the interconnection method between the data stream cores in the data stream core architecture is a cross-tree interconnection method. Broadcasting the decomposed data processing tasks to each data stream core based on the cross-tree interconnection method can improve communication efficiency, and broadcasting the decomposed data processing tasks can reduce the number of accesses to the DRAM and the memory access pressure. Multiple data stream cores process the decomposed data processing tasks in parallel to improve the task completion efficiency. Finally, the processing results of each decomposed task are concatenated to obtain the target processing result and written back to the DRAM, that is, the data processing task is completed.
[0081] The following uses Figure 6 a specific schematic diagram of the task processing of a single data stream core shown to illustrate the present application accordingly. The current data stream core receives the decomposed data processing task. It can be understood that the decomposed data processing task is divided into left and right matrices. The received decomposed data processing task is divided to obtain each subtask. Next, the data flow is specifically as follows:
[0082] 1) Determine whether the current data stream core is a relay node;
[0083] 2.1) If the current data stream core is a relay node, first broadcast the left matrix or the right matrix to the corresponding leaf node through the on-chip network, and then execute step 3);
[0084] 2.2) If the current data stream core is not a relay node, directly execute step 3);
[0085] 3) Load the current subtask into the SRAM, control the computing unit to obtain the current subtask from the SRAM, and obtain the processing result of the current subtask in the form of an outer product;
[0086] 4) Write the processing result of the current subtask into the destination register;
[0087] 5) Accumulate the processing results of each subtask in the destination register;
[0088] 6) Determine whether the processing results of all subtasks in the current data stream core have been accumulated;
[0089] 7.1) If the processing results of all subtasks have been accumulated, execute step 8);
[0090] 7.2) If the processing results of all subtasks have not been accumulated, execute step 3);
[0091] 8) Determine the accumulated result of all subtasks in the destination register as the processing result of the decomposed data processing task, that is, obtain the decomposed task processing result, load the decomposed task processing result into the SRAM, and then clear the destination register.
[0092] See Figure 7 As shown, an embodiment of the present application discloses a task processing device based on a data stream core architecture. The interconnection method between the data stream cores in the data stream core architecture is a cross-tree interconnection method; the device includes:
[0093] A task decomposition module 11, configured to decompose the data processing task in the target detection model obtained from the DRAM according to the cache area of each data stream core to obtain a decomposed data processing task; the data processing task includes a matrix multiplication task; the target detection model is a target detection model constructed based on a neural network;
[0094] A result acquisition module 12, configured to broadcast the decomposed data processing task to each data stream core based on the cross-tree interconnection method to obtain the data processing result output by the data stream core in the form of an outer product;
[0095] A result splicing module 13, configured to splice each data processing result to obtain a target processing result, and write the target processing result back to the DRAM.
[0096] The beneficial effects of this application are as follows: In the data stream core architecture of this application, the interconnection method between each data stream core is a cross-tree interconnection method; the data processing tasks in the target detection model obtained from the DRAM are decomposed according to the cache area of each data stream core to obtain the decomposed data processing tasks; the data processing tasks include matrix multiplication tasks; the target detection model is a target detection model constructed based on a neural network; based on the cross-tree interconnection method, the decomposed data processing tasks are broadcast to each data stream core to obtain the data processing results output by the data stream core through the outer product method; the data processing results are spliced to obtain the target processing result, and the target processing result is written back to the DRAM. It can be seen that in this application, the data processing tasks in the target detection model obtained from the DRAM are decomposed according to the cache area of each data stream core. Then, the decomposed data processing tasks that each data stream core needs to process are smaller than the original data processing tasks. Therefore, the load of each data stream core is reduced, making the load balanced. Further, the interconnection method between each data stream core in the data stream core architecture is a cross-tree interconnection method. Broadcasting the decomposed data processing tasks to each data stream core based on the cross-tree interconnection method can improve communication efficiency. And broadcasting the decomposed data processing tasks can reduce the number of accesses to the DRAM and the memory access pressure. Multiple data stream cores process the decomposed data processing tasks in parallel to improve the task completion efficiency. Finally, the processing results of each decomposed task are spliced to obtain the target processing result, and the target processing result is written back to the DRAM, that is, the data processing task is completed.
[0097] Furthermore, an embodiment of this application also provides an electronic device. Figure 8 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of this application.
[0098] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the task processing method based on the data stream core architecture executed by the electronic device disclosed in any of the foregoing embodiments.
[0099] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations are not imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application requirements, and specific limitations are not imposed thereon here.
[0100] Among them, the processor 21 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 21 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0101] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222, data 223, etc., and the storage method may be temporary storage or permanent storage.
[0102] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the task processing method based on the data flow core architecture executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks. The data 223 may include not only the data transmitted by the external device received by the electronic device, but also the data collected by its own input / output interface 25, etc.
[0103] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the foregoing disclosed task processing method based on the data flow core architecture. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0104] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for the relevant parts.
[0105] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application. The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable EPROM (Erasable Programmable Read Only Memory), electrically erasable programmable EEPROM (Electrically Erasable Programmable read only memory), registers, hard disks, removable disks, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the art.
[0106] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0107] The above has introduced in detail a task processing method, apparatus, device and medium based on a data flow core architecture provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A task processing method based on a data flow core architecture, characterized in that: The interconnection mode between the data flow cores in the data flow core architecture is a cross tree interconnection mode; the method comprises: The data processing tasks in the target detection model obtained from the DRAM are decomposed according to the cache area of each data flow core to obtain the decomposed data processing tasks; the data processing tasks include matrix multiplication tasks; the target detection model is a target detection model built based on a neural network; each of the data flow cores is a root node, a relay node and a leaf node, and the relay node is connected to the root node, and the leaf node is connected to the corresponding relay node; Broadcasting the decomposed data processing task to each of the data stream cores based on the cross-tree interconnection method to obtain the decomposed task processing result output by the data stream core in an outer product method; splicing the processing results of each of the decomposed tasks to obtain a target processing result, and writing the target processing result back to the DRAM; The task processing method based on the data flow core architecture further includes: Acquire a two-dimensional grid of the data flow core architecture, and determine the starting data flow core in the two-dimensional grid as the root node; determine the target row and target column where the root node is located in the two-dimensional grid, and determine each data flow core in the target row and the target column except the root node as the relay node; determine each data flow core in the two-dimensional grid except the root node and the relay node as the leaf node; The relay nodes include a first relay node and a second relay node, wherein the first relay node is located in the target row and the second relay node is located in the target column; Accordingly, broadcasting the decomposed data processing task to each of the data flow cores based on the cross tree interconnection method includes: The left matrix and the right matrix in the decomposed data processing task are input to the root node, the first relay node and the second relay node respectively; wherein the left matrix of the first relay node is the same as the left matrix of the root node, and the right matrix of the second relay node is the same as the right matrix of the root node; the right matrix is broadcast to the corresponding leaf nodes through the first relay node, and the left matrix is broadcast to the corresponding leaf nodes through the second relay node.
2. The task processing method based on the data flow core architecture according to claim 1 is characterized in that: The method of broadcasting the decomposed data processing task to each of the data flow cores based on the cross tree interconnection method includes: Controlling the leaf node to send a ready signal to the corresponding relay node to update the ready signal quantity of the relay node; When the ready signal quantity of the relay node is equal to the number of the leaf nodes corresponding to the relay node, the ready signal quantity of the relay node is reset; Controlling the relay node to broadcast the decomposed data processing task and the transmission completion signal to the corresponding leaf node to update the completion signal quantity of the leaf node; When the completion signal quantity of the leaf node meets the preset threshold, the leaf node is determined to be in a ready state, the completion signal quantity of the leaf node is reset, and the process jumps again to the step of controlling the leaf node to send a ready signal to the corresponding relay node.
3. The task processing method based on the data flow core architecture according to claim 1 is characterized in that: The data flow core includes a computing unit and an SRAM, and the computing unit includes a destination register; Accordingly, the data flow core outputs the decomposition task processing result by outer product method, including: The data flow core divides the received decomposed data processing task to obtain subtasks, loads each subtask into the SRAM, controls the computing unit to obtain each subtask from the SRAM and obtains the processing result of each subtask by outer product, writes the processing result of each subtask into the destination register in sequence, and accumulates the processing results of each subtask in the destination register to obtain the decomposed task processing result.
4. The task processing method based on the data flow core architecture according to claim 3 is characterized in that: The SRAM includes a first cache area and a second cache area; Accordingly, the loading of each of the subtasks into the SRAM, controlling the computing unit to obtain each of the subtasks from the SRAM, and obtaining the processing results of each of the subtasks by means of outer products include: Alternately determining the first cache area and the second cache area as current cache areas; Loading the current subtask into the current cache area; After the computing unit obtains the processing result of the previous subtask by the outer product method, the computing unit is controlled to obtain the current subtask from the current cache area and obtain the processing result of the current subtask by the outer product method.
5. A task processing device based on a data flow core architecture, characterized in that: The interconnection mode between the data flow cores in the data flow core architecture is a cross tree interconnection mode; the device includes: A task decomposition module is used to decompose the data processing tasks in the target detection model obtained from the DRAM according to the cache area of each data flow core to obtain the decomposed data processing tasks; the data processing tasks include matrix multiplication tasks; the target detection model is a target detection model built based on a neural network; each of the data flow cores is a root node, a relay node and a leaf node, and the relay node is connected to the root node, and the leaf node is connected to the corresponding relay node; A result acquisition module, used for broadcasting the decomposed data processing task to each of the data stream cores based on the cross tree interconnection method, so as to obtain the data processing result output by the data stream core in the outer product method; A result splicing module, used for splicing the data processing results to obtain a target processing result, and writing the target processing result back to the DRAM; The task processing device based on the data flow core architecture is specifically used for: Acquire a two-dimensional grid of the data flow core architecture, and determine the starting data flow core in the two-dimensional grid as the root node; determine the target row and target column where the root node is located in the two-dimensional grid, and determine each data flow core in the target row and the target column except the root node as the relay node; determine each data flow core in the two-dimensional grid except the root node and the relay node as the leaf node; The relay nodes include a first relay node and a second relay node, wherein the first relay node is located in the target row and the second relay node is located in the target column; Accordingly, the result acquisition module is specifically used for: The left matrix and the right matrix in the decomposed data processing task are input to the root node, the first relay node and the second relay node respectively; wherein the left matrix of the first relay node is the same as the left matrix of the root node, and the right matrix of the second relay node is the same as the right matrix of the root node; the right matrix is broadcast to the corresponding leaf nodes through the first relay node, and the left matrix is broadcast to the corresponding leaf nodes through the second relay node.
6. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the steps of the task processing method based on the data flow core architecture as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the task processing method based on the data flow core architecture as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Method, system and device for improving chip computing performance and medium
CN111176731A
Storage medium, task execution management device, and task execution management method
US20210026685A1