A Mapping System and Mapping Method for CGRA Multi-Task Dynamic Resource Allocation

By designing II prediction module, lightweight layout recommendation module, layout routing module and configuration generation module in the CGRA multi-task dynamic resource allocation system, the problem of high mapping time consumption in the existing technology is solved, and efficient and automatic CGRA multi-task resource allocation and mapping is achieved.

CN117992216BActive Publication Date: 2025-05-27HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410001779.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-05-27
Estimated Expiration
2044-01-02

AI Technical Summary

Technical Problem

In the prior art, the mapping process time consumes a lot of CGRA multi-task dynamic resource allocation, resulting in low overall throughput of multi-tasks and lacks a hardware dynamic mapping system with high mapping speed and quality.

Method used

A mapping system for multi-task dynamic resource allocation in CGRA is designed, including II prediction module, lightweight layout recommendation module, layout and routing module and configuration generation module, which are executed through pipelines to achieve fast mapping and high-quality layout.

Benefits of technology

The mapping process is significantly accelerated, and the time consumption is reduced by 3 to 4 orders of magnitude, reaching or approaching high mapping quality, while achieving low on-chip memory footprint and automatic mapping without software programming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117992216B_ABST
    Figure CN117992216B_ABST
Patent Text Reader

Abstract

The present invention discloses a mapping system and a mapping method for CGRA multi-task dynamic resource allocation, including: a CGRA processing architecture, a direct memory access unit, an off-chip memory, and a CPU processor; both the CPU processor and the direct memory access unit are signal-connected to the CGRA processing architecture; the CGRA processing architecture is integrated with an on-chip memory, an array of processing units, a dynamic mapper, and a mapping result broadcaster; the on-chip memory, the dynamic mapper, and the mapping result broadcaster are all bidirectionally signal-connected to the direct memory access unit; the on-chip memory is also connected to the array of processing units; the dynamic mapper includes: an II prediction module, a lightweight layout recommendation module, a layout and routing module, and a configuration generation module. The technical solution of the present invention can achieve automatic mapping after CGRA multi-task resource allocation, without the need for software personnel to program, and at the same time has the characteristics of high mapping speed and high mapping quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dynamic resource allocation, and particularly relates to a mapping system and a mapping method for CGRA multi-task dynamic resource allocation. Background Art

[0002] Multi-tasking refers to the ability to execute multiple applications simultaneously, which is of great significance for supercomputers, cloud computing, and embedded computing to meet the user service quality (QoS). Mainstream computing architectures including CPUs, GPUs, and DSPs have provided rich support for multi-tasking. In recent years, due to its flexibility and energy efficiency and performance close to ASICs, the coarse-grained reconfigurable architecture (CGRA) has gradually been recognized by the academic community as a strong competitor to mainstream computing architectures. However, due to the essential differences between CGRA and traditional computing architectures in computing logic and hardware architecture, the research on CGRA multi-tasking is still in its infancy. The multi-task resource allocation strategy is simply set to be static, either space partitioning, time multiplexing, or a combination of both. Since the creation, destruction of multi-tasks, and the changes in task input data characteristics will all lead to changes in the hardware resource requirements of different tasks, the dynamic resource allocation of CGRA multi-tasks is necessary for improving the overall throughput. Among all dynamic resource allocation operations, the dynamic mapper is responsible for mapping the data flow graph (DFG) of a task to the newly allocated CGRA resources, which is one of the most important processes determining the subsequent execution efficiency.

[0003] Since the mapping process is a combinatorial optimization problem (COP), most existing mappers are based on heuristic algorithms, meta-heuristic algorithms, exact algorithms, or machine learning algorithms. However, these static mappers are either specific to a particular CGRA architecture or difficult to achieve high mapping quality with low mapping time consumption. There are mainly two reasons for the high mapping time consumption. First, due to the lack of a reasonable mathematical model of the relationship between the characteristics of CGRA, DFG, and mapping quality, existing static mappers may start from an inappropriate initiation interval (II) and gradually reach the optimal II through repeated mapping attempts, which introduces unnecessary computational overhead and thus increases the mapping time consumption. Second, since the mapping process is mostly an NP-complete problem in most cases, complex layout recommendation phases such as the random and iterative changes in meta-heuristic algorithms, the systematic backtracking in integer linear programming (ILP), and the learning process in ML are necessary to ensure the quality of the solution, so the high mapping time consumption is inevitable.

[0004] In summary, directly adopting these static mappers in the CGRA multi-task scenario requires a long time to remap the DFG to the newly allocated CGRA resources, which will seriously affect the overall throughput rate of multi-tasks. In addition, the dynamic mapper should also have dedicated hardware to avoid consuming host computing resources and additional software programming work.

[0005] Therefore, finding a general solution to repeated mapping and complex layout recommendation, and establishing a CGRA multi-task dynamic resource allocation hardware dynamic mapping system with both high mapping speed and mapping quality is the challenge that the present invention aims to solve. Summary of the Invention

[0006] The object of the present invention is to provide a mapping system and mapping method for CGRA multi-task dynamic resource allocation to solve the problem of slow overall throughput rate of multi-task allocation in the prior art.

[0007] On the one hand, to achieve the above object, the present invention provides a mapping system for CGRA multi-task dynamic resource allocation, including: a CGRA processing architecture, a direct memory access unit, an off-chip memory, and a CPU processor;

[0008] The direct memory access unit, the off-chip memory, and the CPU processor are connected by two-way signals in sequence; both the CPU processor and the direct memory access unit are signal-connected to the CGRA processing architecture;

[0009] The CGRA processing architecture integrates an on-chip memory, a processing unit array, a dynamic mapper, and a mapping result broadcaster; the on-chip memory, the dynamic mapper, and the mapping result broadcaster are all connected to the direct memory access unit by two-way signals; the on-chip memory is also connected to the processing unit array;

[0010] The dynamic mapper includes: an II prediction module, a lightweight layout recommendation module, a layout and routing module, and a configuration generation module;

[0011] Among them, the II prediction module, the lightweight layout recommendation module, the layout and routing module, and the configuration generation module are connected by two-way signals in sequence; the II prediction module, the lightweight layout recommendation module, the layout and routing module, and the configuration generation module are executed in a pipeline manner.

[0012] Optionally, adjacent modules of the II prediction module, the lightweight layout recommendation module, the layout and routing module, and the configuration generation module are connected by two-way handshake signals.

[0013] Optionally, the dynamic mapper further includes a common register file;

[0014] Adjacent modules of the II prediction module, the lightweight layout recommendation module, the layout and routing module, and the configuration generation module are all signal-connected to the common register file.

[0015] On the other hand, to achieve the above object, the present invention provides a mapping method for CGRA multi-task dynamic resource allocation, comprising:

[0016] Step 1: Initialize each module and set parameters;

[0017] Step 2: Obtain the task data flow diagram and current task data, input the current task data into the II prediction module for mapping prediction, and obtain mapping prediction data;

[0018] Step 3: Based on the lightweight layout recommendation module, an initial layout recommendation list of each node of the task data flow graph is obtained, and the initial layout recommendation list is adjusted twice to obtain a final layout recommendation list;

[0019] Step 4: Perform layout based on the mapping prediction data and the final layout recommendation list through the layout and routing module, and update the module routing resource map;

[0020] Step 5: The configuration generation module reads the updated module routing resource map from the placement and routing module to obtain configuration information, writes the configuration information into the off-chip memory through the direct memory access unit, and broadcasts it to each processing unit.

[0021] Optionally, the current task data is input into the II prediction module for mapping prediction, specifically including:

[0022] Construct a mapping relationship model, input the mapping relationship model and current task data into the II prediction module for mapping prediction, and obtain mapping prediction data;

[0023] Among them, the mapping relationship model is used to reflect the relationship between prediction data, task data flow diagram, CGRA processing architecture and task CGRA resource allocation.

[0024] Optionally, the step of obtaining the initial layout recommendation list in step 3 includes:

[0025] The mapping prediction data is prioritized based on the Fan I / O matching criterion to obtain an initial layout recommendation list; the initial layout recommendation list is an initial result after the processing units assigned to the current task are prioritized according to the Fan I / O criterion.

[0026] Optionally, the initial layout recommendation list is adjusted again to obtain a final layout recommendation list, including:

[0027] Calculate the number of Bypass nodes required by the processing units in the initial layout recommendation list and the processing units where the corresponding parent nodes are located in turn;

[0028] Sort the array of processing units with the same FanI / O in the initial layout recommendation list according to the BypassLess criterion in ascending order of the required number of Bypass nodes to obtain the final layout recommendation list.

[0029] Optionally, write the configuration information to the off-chip memory through the direct memory access unit, specifically including:

[0030] Generate the configuration information of the nodes and corresponding parent nodes of the current task data flow graph based on the updated module routing resource graph, and write the configuration information to the off-chip memory through the direct memory access unit; continuously read the configuration information from the off-chip memory through the direct memory access unit and broadcast it to each processing unit.

[0031] The technical effects of the present invention are as follows:

[0032] (1) The II prediction module and the lightweight layout recommendation module provided by the present invention can solve the high mapping time consumption caused by the existing static mapper due to repeated mapping and complex layout recommendation, accelerate the mapping process by 3 to 4 orders of magnitude, that is, 1000 to 10000 times, while achieving or approaching high mapping quality and having extremely low on-chip memory occupancy; the present invention can realize the automatic mapping after the multi-task resource allocation of the CGRA, without the need for software personnel to program, and has the characteristics of high mapping speed and high mapping quality.

[0033] (2) On the basis of the traditional CGRA architecture, the present invention expands the dynamic mapper and the mapping result broadcaster. Through simulation by EDA tools, the area of the dynamic mapper and the mapping result broadcaster only accounts for 25.41% of the area of the 10×10 CGRA chip, and can operate at a frequency of 800 MHz under the 45nm process, with a power consumption of 735.6 mW. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0035] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0036] Figure 1 It is the overall structure diagram of the dynamic mapper in the embodiment of the present invention;

[0037] Figure 2Schematic diagram of the mapping relationship model in the embodiments of the present invention;

[0038] Figure 3 Schematic diagram of the process for generating a layout recommendation list in the embodiments of the present invention;

[0039] Figure 4 Schematic diagram of the Placement & Routing Module in the embodiments of the present invention;

[0040] Figure 5 Schematic diagram of the Configuration Generation Module in the embodiments of the present invention;

[0041] Figure 6 Mapping flowchart in the embodiments of the present invention.

[0042] Figure 1 In (a), in the Architecture of CGRATile, N is north, S is south, W is west, E is east, Xbar is a crossbar switch. Figure 2 In, CGRALinks is the data path between each Tile in CGRA; Combinations is combination, abbreviated as Comb, T1 - T6 are the labels of Tiles, F1 - F3 are the Fan I / O capabilities of each Tile, and II is the initial interval. Figure 3 In, DataFlow Graph is a data flow graph, A - E are example nodes in the data flow graph, Predecessor is the parent node, Current DFGNode is the current DFG node, Mapping is mapping, Static mapping results are static mapping results, closet is the closest, Fan - I / O is the fan - in and fan - out capabilities of the Tile, Initial / Final Placement recommendation list is the initial / final layout recommendation list, BypassNodes are bypass nodes. Haveplaced on is already placed, data is data, and switch is a switch. Figure 4 In, Cycle is the clock cycle, T1 - T6 are Tile numbers, BP is the Bypass node, and II is the initial interval; in Figure 5Among them, SrcCycle is the source clock cycle, TgtCycle is the destination clock cycle, SrcTile is the source Tile, TgtTile is the destination Tile, Type1 to Type9 are configuration instruction types, BP, CMP, Add, LD are arithmetic operations, Mux is a multiplexer, Packing is a packer, Crossbar configuration is cross-switch configuration information, and Functional Unit is functional unit configuration information. Embodiment

[0043] A variety of exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and implementation schemes of the present invention.

[0044] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe this application in detail with reference to the accompanying drawings and in combination with the embodiments.

[0045] As Figure 1 - Figure 6 As shown, in this embodiment, a mapping system for CGRA multi-task dynamic resource allocation is provided, including: a CGRA processing architecture, a direct memory access unit, an off-chip memory, and a CPU processor; the direct memory access unit, the off-chip memory, and the CPU processor are sequentially bidirectionally signal-connected; both the CPU processor and the direct memory access unit are signal-connected to the CGRA processing architecture; the CGRA processing architecture integrates an on-chip memory, a processing unit array, a dynamic mapper, and a mapping result broadcaster; the on-chip memory, the dynamic mapper, and the mapping result broadcaster are all bidirectionally signal-connected to the direct memory access unit; the on-chip memory is also connected to the processing unit array; the dynamic mapper includes: an II prediction module, a lightweight layout recommendation module, a layout and routing module, and a configuration generation module; among them, the II prediction module, the lightweight layout recommendation module, the layout and routing module, and the configuration generation module are sequentially bidirectionally signal-connected; the II prediction module, the lightweight layout recommendation module, the layout and routing module, and the configuration generation module are executed in a pipeline manner.

[0046] A mapping method for CGRA multi-task dynamic resource allocation includes: initializing each module and setting parameters; obtaining a task data flow graph and current task data, inputting the current task data into the II prediction module for mapping prediction to obtain mapping prediction data; obtaining an initial layout recommendation list for each node of the task data flow graph based on the lightweight layout recommendation module, and performing secondary adjustment on the initial layout recommendation list to obtain a final layout recommendation list; performing layout based on the mapping prediction data and the final layout recommendation list through the layout and routing module, and updating the module routing resource graph; the configuration generation module reads the updated module routing resource graph from the layout and routing module to obtain configuration information, and writes the configuration information into the off-chip memory through the direct memory access unit, and broadcasts it to each processing unit.

[0047] The purpose of this embodiment is to find a general method to solve the problems of repeated mapping and complex layout recommendation, establish a CGRA multi-task dynamic resource allocation dynamic mapper with high mapping speed and mapping quality, and at the same time provide dedicated hardware architecture support.

[0048] The overall architecture of the dynamic mapper applied to CGRA multi-task dynamic resource allocation is as Figure 1 shown.

[0049] In Figure 1 In the CGRA multi-task dynamic resource allocation scenario shown, the Dynamic Mapper and the Broadcaster of the mapping result are extensions of the traditional CGRA architecture, and are both connected to the off-chip memory (Mem) through the direct memory access unit (DMAUnit). During the CGRA multi-task dynamic resource allocation process, the Dynamic Mapper takes turns to execute the mapping of each task (Task1~Task4) to their newly allocated resources, and writes the mapping result into the Mem through the DMA. After the mapping of all tasks is completed, the BroadCaster continuously reads the mapping result from the Mem through the DMA, and then broadcasts it to each Tile in the CGRA through the Links specified for broadcasting.

[0050] RoutReg is the steering register; LocalReg is the local register; ComputeReg is the arithmetic register; Configuration memory is the configuration memory; ResourceAllocator is the dynamic resource allocator; Configuration Receiver is the configuration receiver; Functional Unit is the functional unit;

[0051] (b) In the architecture of Multi-Tasks CGRA, the Mini-Banked Data Memory is on-chip memory,

[0052] (c) In the architecture and workflow, AXI is the bus protocol, Allocated Tiles are the CGRATiles allocated for the current task, Static mapping results are the static mapping results of the current task, and DFG is the data flow graph;

[0053] (d) In the High-Level Timing Diagram, Stage is the stage and DFG Node is the data flow graph node;

[0054] The internal structure of the processing unit (Tile) in the traditional CGRA architecture has been modified to support the operations of the Dynamic Mapper and the Broadcaster. As Figure 1 shown, the newly added internal structure of the Tile includes a dedicated data transfer network (N#, S#, W#, E#) for the Bypass node, which can reduce the situation where the Bypass node and the DFG node preempt the data transfer port during the mapping process and increase the probability of successful mapping.

[0055] For the Broadcaster, there is a corresponding Configuration Receiver in each Tile to identify whether the ID of the incoming mapping result matches the ID of the Tile. If they match, the mapping result is updated to the Configuration Memory. In addition, there is a Resource Allocator that uses the Ant Colony Optimization (ACO) algorithm to determine which task the current Tile should be allocated to.

[0056] The internal structure of the Dynamic Mapper is as Figure 1As shown, the dynamic mapper consists of A.II Prediction Module, B. Lightweight Placement Recommendation Module, C. Placement & Routing Module, and D. Config Generation Module. The four modules share a Public Register File for exchanging the calculation results of the modules. To connect adjacent modules, handshake signals (vld, rdy) are designed so that the mapping of each DFG node can be executed in a pipelined manner, which further improves the working efficiency of the proposed Dynamic Mapper. Each module in the Dynamic Mapper consists of a finite state machine (FSM), a set of computing units (Unit), and a Private Register File. Among them, the Unit is designed to achieve the specified algorithm target. The Private Register File is used to store intermediate calculation results within the module. The FSM is responsible for generating the enable and input signals of the Unit, the read and write signals of the Public / Private Register File, and the handshake signals between modules.

[0057] The entire mapping process starts from the Prediction Module A.II. This module takes the proposed mathematical model as input and predicts II by analyzing the shape of the target CGRA. It also reads the inputs required by other modules into the on-chip memory in advance. Finally, the predicted II is written into the Public Register File. After predicting II, the Placement Recommendation Module B starts from the first node in the DFG, prioritizes the Tiles in the target CGRA according to specific principles including Fan I / O matching and BypassLess, and writes the layout recommendation list into the Public Register File. Then, the Placement & Routing Module C reads the predicted II and the layout recommendation list from the Public RegisterFile, performs placement and routing for the current DFG node and possible Bypass nodes, and finally updates the Module Routing Resource Graph (MRRG). Finally, the Configuration Generation Module D reads the updated MRRG from the Placement & Routing Module Public RegisterFile, generates the configuration information required for the current DFG Node and its parent nodes, and finally writes the mapping result into the Mem through DMA. Modules B, C, and D are designed to execute in a pipelined manner, enabling the mapping processes of different DFG nodes to overlap, thereby improving the mapping efficiency. The IIPrediction Module and the Placement Recommendation Module can solve the high mapping time consumption caused by repeated mapping and complex layout recommendations in existing static mappers, accelerating the mapping process by 3 to 4 orders of magnitude (1000 to 10000 times), while achieving or approaching high mapping quality and having extremely low on-chip memory occupancy; through the simulation results of EDA tools, the area of the Dynamic Mapper and the Broadcaster only accounts for 25.41% of the 10×10 CGRA chip area, and it can operate at a frequency of 800 MHz in a 45nm process with a power consumption of 735.6 mW.

[0058] Next, taking the last DFG node E of Task1 in Figure 1 as an example, the specific implementation manners of the four modules A, B, C, and D will be specifically introduced.

[0059] A.II Prediction Module: The purpose of the IIPrediction Module is to accurately predict the II that can directly achieve successful mapping based on the CGRA resources allocated for the current task. The subsequent Placement & Routing Module will use the predicted value as the initial II, reducing the number of repeated mappings and thus the mapping time consumption. As Figure 2 shown, in this embodiment, a mathematical model that can reflect the relationship between the mapping quality II, DFG, and CGRA architecture is first baked offline. After the current task DFG is allocated to new CGRA resources, the II prediction module first calculates the number of CGRALinks it has (12 in this embodiment), reads the corresponding mathematical model from the Mem, and then substitutes the number of Links to obtain the corresponding predicted II (3 in this embodiment).

[0060] B.Placement Recommendation Module: The Placement Recommendation Module is responsible for generating a placement recommendation list (Placement Recommendation List) for the current DFG nodes. The subsequent Placement & Routing Module will alternately try to map the current DFG nodes to each Tile in the placement recommendation list, which greatly reduces the complexity of mapping. The placement recommendation list is a sorted version of the Tiles (Allocated Tiles) in the target CGRA, following the principles of Fan-I / O Match and BypassLess to significantly reduce the mapping time consumption. The definitions and working processes of the two principles are introduced below.

[0061] As Figure 3As shown, assume that the Placement Recommendation Module is currently determining the Placement Recommendation List for its last DFG node E (Current DFG Node). The Allocated Tiles are not regular squares but irregular shapes, consisting of {Tile1, Tile2, Tile3, Tile4, Tile5, Tile6}. The parent node D (Predecessor) of node E has been mapped onto Tile4. First, obtain the layout Tile of DFG node E in the existing mapping results, which is Tile5, with a Fan I / O = 3. Then, obtain the Tiles on the target CGRA with the closest Fan I / O, which are {Tile3, Tile5}, also with a Fan I / O = 3, and both are taken as the primary recommendations. Finally, in the order of decreasing Fan I / O, {Tile2, Tile6, Tile1, Tile4} are taken as secondary recommendations. Therefore, the initial placement recommendation list (Initial Placement Recommendation List) obtained by the Fan I / O matching criterion is {Tile3, Tile5, Tile2, Tile6, Tile1, Tile4}. Calculate the number of Bypass nodes required for the Tiles in the Initial Placement Recommendation List and the Tile4 where the parent node D is located in turn, which are {2, 0, 1, 1, 3, 0}. Sort the Tiles with the same Fan I / O according to the number of Bypass nodes from small to large by the BypassLess criterion. The final placement recommendation list (Final Placement Recommendation List) after the second adjustment is {Tile5, Tile3, Tile2, Tile6, Tile4, Tile1}, and the corresponding number of Bypass nodes is {0, 2, 1, 1, 0, 3}. At this point, the Placement Recommendation List for DFG node E has been generated, which is {Tile5, Tile3, Tile2, Tile6, Tile4, Tile1}. Subsequently, the Placement & Routing Module will attempt to place DFG node E on these Tiles in turn according to the order in the list.

[0062] C. Placement&Routing Module: The Placement&Routing Module accesses the prediction II and the Placement Recommendation List in the Public Register File as inputs. The purpose of this module is to find the appropriate execution clock cycle and Tile on the CGRA for the current DFG node E, and generate Bypass nodes and their mappings when necessary to meet the communication requirements between the layout Tile of DFG node E and the layout Tile of DFG node D. For each Tile in the Placement Recommendation List, calculate its spatial distance to the layout Tile4 of DFG node D, and then generate the minimum number of Bypass nodes and their layout Tiles in the modulo routing resource graph (MRRG) that do not conflict with the existing occupancy of the CGRA Links. Subsequently, starting from the cycle when the Bypass nodes are placed in the MRRG, check the occupancy status of the current Tile in the MRRG from cycle % II to (cycle + II) % II. If it is not occupied, place the DFG node on the corresponding cycle of the current Tile in the MRRG. Otherwise, switch to the next Tile in the Placement Recommendation List and retry. As Figure 4 shown, the first 5 Tiles in the Placement Recommendation List are occupied by the modulo cycle in the MRRG, and only Tile1 in cycle17 and cycle20 is not occupied. Therefore, considering the routing legality, DFG node E is placed on MRRG Cycle20 Tile1 because the nearest Bypass node is MRRG Cycle18 Tile3, and MRRG cycle19 Tile1 has already been occupied in the modulo cycle.

[0063] D. Configuration Generation Module: The Configuration Generation Module accesses the Public Register File to obtain the layout results of the current DFG node E, its parent node DFG node D, and the Bypass nodes in the MRRG as inputs, and generates the corresponding configuration signals (Configuration) for the functional units and crossbars of the relevant Tiles and cycles.

[0064] As Figure 5As shown, the configuration of the functional units within a Tile is generated according to the actual functions of different DFG nodes. For example, 0x00 is used for the Bypass node, 0x01 is used for the Cmp node, etc. The configuration of the crossbar within a Tile is generated based on the spatial and temporal differences between the layouts of adjacent DFG nodes among the Tiles in the MRRG. Since the functions of the three types of registers within a Tile are different, the configuration of the crossbar can be divided into 9 types. By matching different types of configurations according to the spatial and temporal differences, the final configuration is obtained. Each time a configuration is generated, the occupancy status of the registers used is updated to avoid subsequent register conflicts. The generated configuration is packaged together with the ID of the Tile and then written to the Mem through DMA.

[0065] As described above, the above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A CGRA multi-task dynamic resource allocation mapping system, characterized in that: include: CGRA processing architecture, direct memory access unit, off-chip memory and CPU processor; The direct memory access unit, the off-chip memory and the CPU processor are connected in a bidirectional signal manner in sequence; the CPU processor and the direct memory access unit are both connected to the CGRA processing architecture signal; The CGRA processing architecture integrates an on-chip storage, a processing unit array, a dynamic mapper and a mapping result broadcaster; the on-chip storage, the dynamic mapper and the mapping result broadcaster are all bidirectionally signal-connected to the direct memory access unit; the on-chip storage is also connected to the processing unit array; The dynamic mapper includes: an II prediction module, a lightweight layout recommendation module, a layout and routing module, and a configuration generation module; The II prediction module, the lightweight layout recommendation module, the layout and routing module and the configuration generation module are connected with each other in sequence by bidirectional signals; the II prediction module, the lightweight layout recommendation module, the layout and routing module and the configuration generation module are executed in a pipeline manner.

2. A CGRA multi-task dynamic resource allocation mapping system according to claim 1, characterized in that: The II prediction module, the lightweight layout recommendation module, the layout and routing module and the adjacent modules of the configuration generation module are bidirectionally connected through handshake signals.

3. A CGRA multi-task dynamic resource allocation mapping system according to claim 1, characterized in that: The dynamic mapper also includes a common register file; The II prediction module, the lightweight layout recommendation module, the layout and routing module and the adjacent modules of the configuration generation module are all connected to the common register file signal.

4. A mapping method for CGRA multi-task dynamic resource allocation, applied to a mapping system for CGRA multi-task dynamic resource allocation as claimed in any one of claims 1 to 3, characterized in that: include: Step 1: Initialize each module and set parameters; Step 2: Obtain a task data flow diagram and current task data, input the current task data into the II prediction module for mapping prediction, and obtain mapping prediction data; Step 3: Based on the lightweight layout recommendation module, an initial layout recommendation list of each node of the task data flow graph is obtained, and the initial layout recommendation list is adjusted twice to obtain a final layout recommendation list; Step 4: Performing layout based on the mapping prediction data and the final layout recommendation list through the layout and routing module, and updating the module routing resource map; Step 5: The configuration generation module reads the updated module routing resource map from the placement and routing module to obtain configuration information, writes the configuration information into the off-chip memory through a direct memory access unit, and broadcasts it to each processing unit.

5. The mapping method according to claim 4, characterized in that: Inputting the current task data into the II prediction module for mapping prediction specifically includes: Constructing a mapping relationship model, inputting the mapping relationship model and the current task data into the II prediction module for mapping prediction, and obtaining mapping prediction data; The mapping relationship model is used to reflect the relationship between the prediction data, the task data flow diagram, the CGRA processing architecture and the task CGRA resource allocation.

6. A mapping method for CGRA multi-task dynamic resource allocation according to claim 4, characterized in that: The steps for obtaining the initial layout recommendation list in step 3 include: The mapping prediction data is prioritized based on the Fan I / O matching criterion to obtain an initial layout recommendation list; the initial layout recommendation list is an initial result after the processing units assigned to the current task are prioritized according to the Fan I / O criterion.

7. A mapping method for CGRA multi-task dynamic resource allocation according to claim 4, characterized in that: The initial layout recommendation list is adjusted twice to obtain a final layout recommendation list, which specifically includes: Calculating in turn the number of Bypass nodes required by the processing units in the initial layout recommendation list and the processing units where the corresponding parent nodes are located; According to the BypassLess criterion, the processing unit arrays with the same FanI / O in the initial layout recommendation list are sorted from small to large according to the number of Bypass nodes to be generated, so as to obtain a final layout recommendation list.

8. A mapping method for CGRA multi-task dynamic resource allocation according to claim 4, characterized in that: and writing the configuration information into the off-chip memory through a direct memory access unit, specifically comprising: Based on the updated module routing resource graph, the configuration information of the nodes and corresponding parent nodes of the current task data flow graph is generated, and the configuration information is written into the off-chip memory through the direct memory access unit; the configuration information is continuously read from the off-chip memory through the direct memory access unit, and broadcast to each processing unit.

Citation Information

Patent Citations

  • A deep learning accelerator system and methods thereof

    CN111630505A

  • Compiling method for reducing multi-class memory access conflicts for coarse-grained reconfigurable structure

    CN112306500A