A context-switching-oriented embedded CGRA secondary memory design method
By designing an embedded CGRA secondary memory for context switching and using a finite state machine to control alternating read and write operations, the problem of task stalling caused by context switching in the CGRA architecture is solved, achieving a significant throughput improvement.
Patent Information
- Application Number
- CN202411253760.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-09-09
AI Technical Summary
When the existing CGRA architecture dynamically allocates resources for multiple tasks, context switching causes task stagnation and affects throughput.
Design an embedded CGRA secondary memory for context switching. Use PyMTL3 language to describe the CGRA hardware architecture and add components to implement alternating read and write operations. Use finite state machines to hide context switching delays.
Improved task throughput. Context switching no longer causes task stalls. Throughput increased by 15.8 times to 23.2 times, and the target throughput for resource allocation decisions reached 98.2%.
Smart Images

Figure CN119336670B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a context switching-oriented embedded CGRA secondary memory design method, belonging to the technical field of coarse-grained reconfigurable architecture memory design. Background Art
[0002] A coarse-grained reconfigurable architecture (CGRA) is defined as an array of interconnected processing units (TileArray) whose computing functions and routing directions can be flexibly reconfigured along spatial and temporal dimensions, cooperating to achieve specified computing functions. CGRA not only provides computing performance and energy efficiency close to that of ASICs, but its coarse-grained reconfigurability (Word-Level) also leads to development efficiency superior to that of FPGAs. With the end of Dennard scaling and the decline of Moore's Law, embedded fields such as aerospace, measurement, autonomous driving, and healthcare are increasingly focusing on processor performance and energy efficiency. At the same time, to accommodate the rapid iterative updates of algorithms, processors should also have high programming efficiency. Therefore, CGRA has good development prospects in the embedded field.
[0003] Processors in the embedded field often face workloads that involve a mix of multiple tasks. For example, deep learning-based facial recognition in access control systems may consist of multiple convolution, activation, batch normalization, and pooling layers, with different layers executing in a pipelined fashion as data-dependent tasks. Automotive electronic control units (ECSs) must concurrently process independent tasks such as video, audio, and data analysis from sensors surrounding the vehicle. Therefore, efficiently executing multiple tasks on processors has been a focus of extensive research. Because the throughput of some tasks varies with input density and tasks have random initialization and destruction times, dynamic resource allocation has become a key technology for maximizing processor resource utilization and improving task throughput. Accordingly, CGRA, as an emerging architecture in the embedded field, has also garnered significant attention in recent years for its multi-task dynamic resource allocation methods.
[0004] In the embedded CGRA architecture, dynamic resource allocation for multiple tasks can lead to large-scale mismatches between the configuration instructions and data stored in CGRATiles. For example, before dynamic resource allocation, the current task is allocated three tiles in an "L" shape (3-L), while after dynamic resource allocation, the current task is allocated four tiles in a "field" shape (4-S). Under the static planning of the CGRA compiler, the configuration instructions and data for these two resource allocation scenarios are stored in different tiles. This prevents 4-S from directly using the configuration instructions and data of 3-L. Therefore, after dynamic resource allocation, the configuration instructions and data required by 4-S must be reloaded into the configuration and data memories of the corresponding four tiles before execution can begin. This step, known as a "context switch," is required for all tasks experiencing resource allocation changes. However, existing CGRA multitasking research primarily focuses on compilers and schedulers, lacking efficient memory design for context switching. This causes each task to stall after resource allocation until the corresponding configuration instructions and data are reloaded, severely limiting task throughput. Summary of the Invention
[0005] The present invention aims to solve the technical problem of task stagnation during context switching, and further proposes an embedded CGRA secondary memory design method for context switching.
[0006] The technical solution adopted by the present invention to solve the above problems is: the present invention proposes an embedded CGRA secondary memory design method for context switching, comprising:
[0007] Step 1: Determine the basic CGRA hardware architecture based on actual multi-tasking requirements;
[0008] Step 2: Using the open source CGRA modeling tool OpenCGRA, based on the selected basic CGRA hardware architecture, describe the CGRA hardware architecture using the PyMTL3 language;
[0009] Step 3: Modify the basic component library Mem provided by OpenCGRA, add components using the PyMTL3 language, complete the secondary design of the data memory in each tile, and obtain the embedded CGRA;
[0010] Step 4: Design the sequential logic for state transfer and state output, and the combinational logic for state switching for all working states of the finite state machine in the CGRA in the secondary memory, and describe them in PyMTL3.
[0011] Step 5: Use OpenCGRA to generate synthesizable Verilog code to implement alternate reading and writing of memory group data after context switching.
[0012] Optionally, the basic CGRA hardware architecture in step 1 includes: size, topological connection relationship between tiles, size of data memory and configuration memory within tiles, complexity of functional units, and number of registers
[0013] Optionally, the description of the CGRA hardware architecture in step 2 specifically includes:
[0014] Use the basic component library provided by the modeling tool to build the CGRA hardware architecture.
[0015] Optionally, the two-level design of the data memory in each tile in step 3 specifically includes a first-level memory and a second-level memory;
[0016] The first-level memory design specifically includes: dividing the data memory into memory group A and memory group B, writing the first round of data to memory group A, and when the embedded CGRA is running, reading the first round of data from memory group A to perform calculations, while writing the second round of data to memory group B. When all the first round of data in memory group A has been read out, it directly switches to memory group B to read the second round of data, and writes the third round of data to memory group A, repeating the above steps to perform alternating read and write operations;
[0017] The second-level memory design specifically includes: dividing memory group A and memory group B into sub-memory group A1, sub-memory group A2, sub-memory group B1 and sub-memory group B2 respectively, so that the alternating read and write operations are only for sub-memory group A1 and sub-memory group B1. When the dynamic resource allocation performs context switching, the first round of data is read from sub-memory group A1 and the second round of data is written to sub-memory group B1 at the same time. At the same time, the remaining data of the first round of data that can be processed under the new resource allocation decision is temporarily written to sub-memory group A2. After the writing is completed, the read operation on sub-memory group A1 will be switched to sub-memory group A2, and the written data will be processed under the new resource allocation decision. After the processing is completed, an alternating operation is performed to read the second round of data from sub-memory group B1 and write the third round of data to sub-memory group A1 at the same time.
[0018] Optionally, the components added in step 3 include a Mux, eight Demuxes, an AXI-Lite receiver, and an AXI-Stream receiver.
[0019] Optionally, all working states of the finite state machine in step 4 include: a normal working state of reading sub-memory group A and writing sub-memory group B, a context switching state, and a working state after context switching, and a normal working state of reading sub-memory group B and writing sub-memory group A, a context switching state, and a working state after context switching;
[0020] The input data of the finite state machine includes: memory access signals from the tile functional unit, AXI-Lite receiver data and AXI-Stream receiver data;
[0021] The memory access signals from the Tile functional unit include write enable wrEn_FU, write address wrAddr_FU, write data wrData_FU and read address rdAddr_FU;
[0022] The AXI-Lite receiver data includes the XI-Lite bus configuration registers Reg_DRA and Reg_ADDR. The size of Reg_DRA is 1 bit and is used to notify the finite state machine for dynamic resource allocation. The size of Reg_ADDR is 16 bits and is used to specify the maximum address of data written to sub-memory group A2 and sub-memory group B2 and processed under the new resource allocation decision.
[0023] AXI-Stream receiver data includes write enable wrEn_MM, write address wrAddr_MM, and write data wrData_MM. The data written to sub-memory group A2 and sub-memory group B2 is received through the AXI-Stream bus. After decoding by the AXI-Stream receiver, write enable wrEn_MM, write address wrAddr_MM, and write data wrData_MM are obtained.
[0024] Optionally, in step 4, sequential logic design for state transfer and state output and combinational logic design for state switching are performed and described in PyMTL3 language, specifically including:
[0025] Step 4.1: When the finite state machine is in normal working state, it reads data from sub-memory group A1 and writes data into sub-memory group B1. When dynamic resource allocation is triggered, it enters the context switching state.
[0026] Step 4.2: When the finite state machine is in the context switching state, read the data of sub-memory group A1 and write it into sub-memory group B1 and sub-memory group A2. When the written data reaches the preset value, enter the state after the context switching;
[0027] Step 4.3: When the finite state machine is in the state after context switching, the data written to the sub-memory group A2 is read. After the reading is completed, the reverse transmission state is entered;
[0028] Step 4.4: Repeat steps 4.1 through 4.4 to read data from sub-memory group B and write it to sub-memory group A. This process obtains the normal working state, context switching state, and working state after context switching of reading from sub-memory group B and writing to sub-memory group A.
[0029] Step 4.5: Infinitely loop steps 4.1 to 4.4 to obtain the output of all states of the finite state machine.
[0030] Optionally, the outputs of the finite state machine in all states in step 4.5 include an 8-bit mode signal, memory access signals for group A and group B, write enable wrEn_A / B, write address wrAddr_A / B, write data wrData_A / B, read address rdAddr_A / B, read data rdData_A / B, mode[15:0] and mode[17:16], wherein mode[15:0] is used to control eight Demux respectively, and direct the correct input signal to the correct sub-memory group according to the current state of the finite state machine, and mode[17:16] is used to control Mux.
[0031] The beneficial effects of the present invention are:
[0032] 1. With the support of the context switching-oriented embedded CGRA secondary memory design designed by the present invention, context switching will not cause task stagnation, thereby improving task throughput.
[0033] 2. Adopting the context-switching embedded CGRA secondary memory designed by this invention can increase throughput by 15.8x, 23.2x, and 17.9x for GEMVER, GESUMMV, and SYMM respectively. Furthermore, within five input data points after triggering resource allocation, an average resource allocation decision-making target throughput exceeding 98.2% can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A block diagram of the embedded CGRA multi-level memory design for context switching provided by the present invention;
[0035] Figure 2 The state transition diagram of the finite state machine provided by the present invention. DETAILED DESCRIPTION
[0036] Specific implementation method 1: In this implementation method, CGRA is a coarse-grained reconfigurable architecture.
[0037] Combine Figure 1-2 This embodiment describes the embedded CGRA multi-level memory design method for context switching described in this embodiment, which includes the following steps:
[0038] S1: Determine the basic CGRA hardware architecture based on actual multi-tasking requirements;
[0039] S101: Determine the basic CGRA hardware architecture, including determining the size of the CGRA processing unit array, the topological connection relationship between tiles, the size of the data memory and configuration memory within the tile, the complexity of the functional unit, and the number of registers.
[0040] S102: According to Figure 1The secondary memory architecture shown in the figure modifies the basic component library Mem provided by OpenCGRA and adds components such as AXI-Lite receiver, AXI-Stream receiver, Demux, and Mux through the PyMTL3 language;
[0041] S103: At the first level, the data memory is divided into two sub-memories: memory group A and memory group B;
[0042] Due to the power consumption and area limitations of chips in embedded scenarios, the data memory size of each tile is generally 1KB, which is difficult to store all the data required for task calculation, especially for tasks with streaming input such as satellite-borne target detection, access control, systems, urban pollution monitoring, etc. Therefore, Figure 1 As shown, this embodiment divides the 1KB data memory into memory group A and memory group B, each of which is 512Bytes. Before CGRA runs, the first round of data is first written to memory group A. When CGRA runs, the first round of data is first read from memory group A to perform calculations, and the second round of data is written to memory group B. When the first round of data in memory group A is read out, it can be directly switched to memory group B to read the second round of data, and the third round of data can be written to memory group A, alternating in sequence. By repeating the above alternating read and write operations, CGRA is able to continuously process streaming input data;
[0043] S104: At the second level, memory group A and memory group B are further divided into two groups of sub-memory, group 1 and group 2. The four groups of sub-memory, sub-memory group A1, sub-memory group A2, sub-memory group B1 and sub-memory group B2, have independent read and write signal lines, including write enable wrEn, write address wrAddr, write data wrData, read address rdAddr, and read data rdData, and perform read and write operations under the control of a finite state machine to hide the context switching delay.
[0044] When reading and writing data in memory group A and memory group B alternately, if dynamic resource allocation occurs, context switching will force task execution to stagnate. The task can only continue to execute after the corresponding data under the newly allocated resources is reloaded into the memory group A and memory group B of the correct tiles. To solve this problem, Figure 1As shown, this embodiment further divides group A and group B into memory group 1 and memory group 2, and makes the alternating read and write operations only for sub-memory group A1 and sub-memory group B1. When the context is switched during dynamic resource allocation, for example, the first round of data is being read from sub-memory group A1 and the second round of data is being written to sub-memory group B1 at the same time, the original operation will continue, and the remaining data of the first round of data that can be processed under the new resource allocation decision will be temporarily written to sub-memory group A2. After the writing is completed, the read operation for sub-memory group A1 will be switched to sub-memory group A2, and the written data will be processed under the new resource allocation decision. After the processing is completed, an alternating operation is performed to read the second round of data from sub-memory group B1 and write the third round of data to sub-memory group A1 at the same time. With the support of the above-mentioned secondary memory, context switching will not cause stagnation to the task, thereby improving the task throughput.
[0045] S2: According to Figure 2 The state transition diagram of the finite state machine shown in the figure is used to design the sequential logic of state transition and state output, as well as the combinational logic of state switching, and is described in PyMTL3 language.
[0046] S201: The input of the finite state machine includes: (1) memory access signals from the independent routing network (FU), including write enable wrEn_FU, write address wrAddr_FU, write data wrData_FU, and read address rdAddr_FU; (2) registers Reg_DRA and Reg_ADDR configured by the AXI-Lite bus, Reg_DRA is 1 bit in size and is used to notify the finite state machine to perform dynamic resource allocation, and Reg_ADDR is 16 bits in size and is used to specify the maximum address of data written to sub-memory group A2 or sub-memory group B2 that can be processed under the new resource allocation decision; (3) data written to sub-memory group A2 or sub-memory group B2 is received through the AXI-Stream bus, and the write enable wrEn_MM, write address wrAddr_MM, and write data wrData_MM are obtained after decoding by the receiver. The output of the finite state machine includes an 18-bit mode signal, memory access signals for group A and group B, write enable wrEn_A / B, write address wrAddr_A / B, write data wrData_A / B, read address rdAddr_A / B, and read data rdData_A / B. Mode[15:0] is used to control the eight Demux registers, directing the correct input signals to the correct sub-memory group according to the current state of the finite state machine. Mode[17:16] is used to control the Mux registers, selecting the output signals of the correct sub-memory group as the read result of the data memory according to the current state of the finite state machine.
[0047] S202: The state of the finite state machine is as follows: Figure 2 As shown, reset forces the state to IDLE. The definition and jump logic of all states are as follows:
[0048] RA1_WB1: Normal working state, reading sub-memory bank A1 and writing sub-memory bank B1. When dynamic resource allocation is triggered (Reg_DRA = 1'd1), it enters RA1_WA2_WB1 state, otherwise it remains in this state.
[0049] RA1_WA2_WB1: Context switch state, continue to read sub-memory group A1 and write to sub-memory group B1, but also write to sub-memory group A2. Write address wrAddr_A2 is generated by self-incrementing the counter. When the target write data amount is reached (wrAddr_A2 = Reg_DRA), enter the RA2_WB1 state, otherwise maintain this state unchanged;
[0050] Entering RA2_WB1: Working state after context switching: Reading the data written to sub-memory A2. When the reading is completed, entering RB1_WA1 state, otherwise maintaining the current state unchanged;
[0051] RB1_WA1, RB1_WB2_WA1, RB2_WA1: The meaning is similar to the above status, but reads from memory bank B and writes to memory bank A.
[0052] Through the above operations, the present invention can perform read and write operations under the control of a finite state machine, hide context switching delays, and avoid task stagnation during context switching, thereby improving task throughput.
[0053] Example
[0054] To verify the technical effects of the present invention, this embodiment adopts the following experimental settings:
[0055] (1) CGRA architecture: The size of the CGRA processing unit array is 16 times 16, with a total of 256 tiles. The data memory size of each tile is 1KB, the configuration memory size is 512Bytes, and the topology of the interconnection between tiles is "orthogonal + diagonal".
[0056] (2) Data-intensive tasks: Three computing tasks in the field of matrix multiplication were selected from the open source CGRA task set Polybench, namely multi-matrix-vector multiplication (GEMVER), sum matrix-vector multiplication (GESUMMV), and symmetric matrix multiplication (SYMM).
[0057] (3) Resource Allocation Decision: Before resource allocation, the three tasks were evenly divided into 16x16 CGRAs, each executing on 85 tiles. The resource allocation decision was to have one task absorb half of the resources from the other two tasks to improve its own throughput. That is, GEMVER, GESUMMV, and SYMM will rotate and execute on 170 tiles. The C source code for the three tasks will be compiled using MLIR and LLVM and pre-mapped under the above resource allocation decision for query during task execution.
[0058] Experimental results: Under the above experimental settings, GEMVER, GESUMMV, and SYMM achieved throughput improvements of 15.8x, 23.2x, and 17.9x, respectively, after adopting secondary memory. Furthermore, within five input data points after triggering resource allocation, they achieved an average resource allocation decision throughput exceeding 98.2%.
[0059] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A context-switching-oriented embedded CGRA secondary memory design method, characterized in that: The steps of the context switching-oriented embedded CGRA secondary memory design method include: Step 1: Determine the basic CGRA hardware architecture based on actual multi-tasking requirements; Step 2: Using the open source CGRA modeling tool OpenCGRA, based on the selected basic CGRA hardware architecture, describe the CGRA hardware architecture using the PyMTL3 language; Step 3: Modify the basic component library Mem provided by OpenCGRA, add components using the PyMTL3 language, complete the secondary design of the data memory in each tile, and obtain the embedded CGRA; The two-level design of data memory in each tile in step 3 specifically includes first-level memory and second-level memory; The first-level memory design specifically includes: dividing the data memory into memory group A and memory group B, writing the first round of data to memory group A, and when the embedded CGRA is running, reading the first round of data from memory group A to perform calculations, while writing the second round of data to memory group B. When all the first round of data in memory group A has been read out, directly switching to memory group B to read the second round of data, and writing the third round of data to memory group A, repeating the above steps to perform alternating read and write operations; The second-level memory design specifically includes: dividing memory group A and memory group B into sub-memory group A1, sub-memory group A2, sub-memory group B1, and sub-memory group B2, respectively, so that alternating read and write operations are only performed on sub-memory group A1 and sub-memory group B1. When dynamic resource allocation performs context switching, the first round of data is read from sub-memory group A1 and the second round of data is simultaneously written to sub-memory group B1. At the same time, the remaining data of the first round of data that can be processed under the new resource allocation decision is temporarily written to sub-memory group A2. After the writing is completed, the read operation on sub-memory group A1 will be switched to sub-memory group A2, and the written data will be processed under the new resource allocation decision. After the processing is completed, an alternating operation is performed to read the second round of data from sub-memory group B1 and write the third round of data to sub-memory group A1. Step 4: Design the sequential logic for state transfer and state output, and the combinational logic for state switching for all working states of the finite state machine in the CGRA in the secondary memory, and describe them in PyMTL3. Step 5: Use OpenCGRA to generate synthesizable Verilog code to implement alternate reading and writing of memory group data after context switching.
2. The context switching-oriented embedded CGRA secondary memory design method according to claim 1, characterized in that: The basic CGRA hardware architecture in step 1 includes: size, topological connection relationship between tiles, size of data memory and configuration memory within tiles, complexity of functional units, and number of registers.
3. The context switching-oriented embedded CGRA secondary memory design method according to claim 1, characterized in that: The description of the CGRA hardware architecture in step 2 specifically includes: Use the basic component library provided by the modeling tool to build the CGRA hardware architecture.
4. The context switching-oriented embedded CGRA secondary memory design method according to claim 1, characterized in that: The components added in step 3 include a Mux, eight Demuxes, an AXI-Lite receiver, and an AXI-Stream receiver.
5. The context switching-oriented embedded CGRA secondary memory design method according to claim 1, characterized in that: All working states of the finite state machine in step 4 include: a normal working state of reading sub-memory group A and writing sub-memory group B, a context switching state, and a working state after context switching; and a normal working state of reading sub-memory group B and writing sub-memory group A, a context switching state, and a working state after context switching; The input data of the finite state machine includes: memory access signals from the tile functional unit, AXI-Lite receiver data and AXI-Stream receiver data; The memory access signal from the Tile functional unit includes a write enable wrEn_FU, a write address wrAddr_FU, a write data wrData_FU and a read address rdAddr_FU; The AXI-Lite receiver data includes registers Reg_DRA and Reg_ADDR configured by the AXI-Lite bus. The Reg_DRA is 1 bit in size and is used to notify the finite state machine to perform dynamic resource allocation. The Reg_ADDR is 16 bits in size and is used to specify the maximum address of data written to sub-memory group A2 and sub-memory group B2 and processed under the new resource allocation decision. The AXI-Stream receiver data includes write enable wrEn_MM, write address wrAddr_MM, and write data wrData_MM. The data written to sub-memory group A2 and sub-memory group B2 are received through the AXI-Stream bus. After decoding, the AXI-Stream receiver obtains write enable wrEn_MM, write address wrAddr_MM, and write data wrData_MM.
6. The context switching-oriented embedded CGRA secondary memory design method according to claim 1, characterized in that: In step 4, the sequential logic design for state transfer and state output and the combinational logic design for state switching are performed and described in PyMTL3 language. Specifically, the following are included: Step 4.1: When the finite state machine is in normal working state, it reads data from sub-memory group A1 and writes data into sub-memory group B1. When dynamic resource allocation is triggered, it enters the context switching state. Step 4.2: When the finite state machine is in the context switching state, read the data of sub-memory group A1 and write it into sub-memory group B1 and sub-memory group A2. When the written data reaches the preset value, enter the state after the context switching; Step 4.3: When the finite state machine is in the state after context switching, the data written to the sub-memory group A2 is read. After the reading is completed, the reverse transmission state is entered; Step 4.4: Repeat steps 4.1 to 4.3 to read data from sub-memory group B and write it to sub-memory group A. Obtain the normal working status, context switching status, and working status after context switching of reading sub-memory group B and writing to sub-memory group A. Step 4.5: Infinitely loop steps 4.1 to 4.4 to obtain the output of the finite state machine in all states.
7. The context switching-oriented embedded CGRA secondary memory design method according to claim 6, characterized in that: The outputs of the finite state machine in all states in step 4.5 include an 8-bit mode signal, memory access signals for group A and group B, write enable wrEn_A / B, write address wrAddr_A / B, write data wrData_A / B, read address rdAddr_A / B, read data rdData_A / B, mode[15:0] and mode[17:16]. The mode[15:0] is used to control eight Demux respectively, and the correct input signal is directed to the correct sub-memory group according to the current state of the finite state machine. The mode[17:16] is used to control the Mux.
Citation Information
Patent Citations
Hierarchical multi-RPU and multi-PEA reconfigurable processor
CN112486908A
Rapid and flexible slice-type dynamic reconstruction configuration system of CGRA
CN117201297A