Data-Driven Coarse-Grained Reconfigurable Array Based Near-Memory Computing System
By designing a near-memory computing system based on data-driven coarse-grained reconfigurable array, the problems of poor universality of existing architectures and low utilization of processing units are solved, and high-energy-efficient near-memory computing and high bandwidth access are achieved.
Patent Information
- Application Number
- CN202210053673.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-01-18
AI Technical Summary
The existing near-memory computing architecture has poor universality, does not support indirect memory access, high compiler requirements, and low processing unit utilization.
A near-memory computing system based on data-driven coarse-grained reconstructable array is designed. The system is divided into an off-chip master control layer, a logic layer of a three-dimensional accelerator and a storage layer. A three-dimensional stacking structure is formed through silicon through-silicon connection to realize direct access and indirect memory access, and the utilization rate of processing units is improved through dynamic execution structure and token buffer.
It realizes high-energy-efficient near-memory computing, fully utilizes the high bandwidth advantages of near-memory computing, improves program universality and processing unit utilization, and reduces the dependence on the compiler.
Smart Images

Figure CN114398308B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of near-memory computing architectures with high energy efficiency ratios. In particular, the present invention relates to a near-memory computing system based on a data-driven coarse-grained reconfigurable array. Background Art
[0002] With the explosive development of Internet services, a large amount of data is generated in the daily use of Internet users. The analysis of massive data has brought huge pressure to computing systems. The performance of traditional computing architectures gradually fails to meet the computing performance requirements of a large number of data-intensive applications. The bottleneck of data-intensive applications lies in the memory wall. Traditional computing architectures transfer data from memory to on-chip memory through a bus, perform a large amount of data operations in the processor, and then write the results back to memory through the bus. As the data scale continues to increase, in order to accelerate the data processing speed, the scale of computing resources increases accordingly. However, the memory bandwidth cannot increase with the scale of the computing system, becoming a bottleneck restricting the computing architecture of modern data centers. Data transfer causes huge time and power consumption overheads in the data analysis algorithm process. Recent research shows that in widely used mobile applications, 62% of the energy consumption is used for data movement. The multi-level cache structure in traditional architectures temporarily stores the read data in a faster cache, which can reduce the number of data transfers via the bus. However, through research, the memory access patterns of many applications result in a large amount of data in the cache not being reused, which instead brings additional latency and power consumption overheads.
[0003] In recent years, the integration technology of semiconductor systems has been further improved, and memory and logic can be tightly integrated. On the premise of this technology, due to the increasing demand for memory systems in new data-intensive applications, the concept of Processing In Memory (PIM) has been proposed again. The main idea of PIM is to perform a large amount of calculations inside the memory chip, avoiding the overhead of data transfer. The implementation method is to directly use the physical characteristics of the storage medium itself for data calculation or integrate the calculation logic of data into the data storage chip. The concept of PIM has been proposed for nearly 50 years, but it has not been widely adopted and studied in the past, mainly due to the following reasons: (1) The semiconductor manufacturing technology in the past could not tightly integrate the storage part and the logic part; (2) The applications in the past were not data-intensive, and the characteristics of PIM had little improvement on the performance of these applications. Now, many data-intensive applications have become the mainstream of applications. As one of the possible technologies to overcome the memory wall, PIM has once again attracted wide attention. Currently, PIM is mainly divided into two categories. The first category is called Processing Using Memory (PUM). This method makes minimal changes to the memory chip to perform simple and powerful general operations, that is, the inherent characteristics of the chip itself or after minor changes to make it have the ability of efficient calculation. The second category is called Processing Near Memory (PNM). This method integrates the calculation logic into the memory controller of traditional DRAM or into the logic layer of new 3D-DRAM.
[0004] The coarse-grained reconfigurable architecture (CGRA) is a special computing architecture different from traditional general-purpose processors and application-specific integrated circuits (ASICs). The former ensures the programmability of the architecture, but is limited by the limited performance of simple general-purpose processors at the same time. The latter improves the execution efficiency of the architecture, but the architecture has a single usage scenario, cannot be configured and reconfigured according to requirements, and cannot effectively amortize high R & D costs. The field programmable gate array (FPGA) is a reconfigurable computing architecture, whose characteristics are between those of general-purpose processors and application-specific integrated circuits. Its reconfiguration unit is reconfigured in the smallest unit of bits (bit). Therefore, the configuration information of FPGA is huge, and the configuration time overhead is long (averaging from ten to dozens of milliseconds), and it can only achieve static reconfiguration and cannot achieve dynamic reconfiguration at runtime. The configuration information of CGRA takes the processing element (PE) as the smallest configuration unit, which greatly reduces the data volume of the reconfigurable configuration information, and the overhead of circuit reconfiguration is also greatly reduced compared with FPGA. Therefore, the circuit structure can be dynamically changed during the operation process, which makes the coarse-grained reconfigurable architecture more flexible than FPGA when executing tasks.
[0005] Near-memory computing architectures are new computing architectures proposed to reduce the overhead of data movement in traditional computing architectures. By integrating computing logic into DRAM memory circuits, data can be computed while being read, avoiding the overhead of data movement on the bus between the processor and memory. In the prior art, there are near-memory computing architectures developed for graph algorithms, where a sequentially configured information processor with a simple structure is configured under each channel in the bottom logic layer of 3D-DRAM. Each processor is only responsible for computing the data in its corresponding channel. However, the computing power of this computing logic is insufficient and the energy efficiency ratio is low, unable to fully utilize the bandwidth advantage of in-memory computing. PIM-Enabled Instructions (PEI) deploy computing tasks to the PIM for operation in terms of configuration information granularity, with the same architecture granularity as that based on traditional processors. The programming model only needs minor modifications to use 3D-memory-based PIM for acceleration. However, its drawback is that its computing logic fragmentarily processes some configuration information of the central processor, and the execution lacks coherence, unable to maximize the utilization of the in-memory computing architecture bandwidth. GRIM-Filter is an in-memory accelerator for accelerating the genomic seed screening algorithm. This architecture loads the genomic seed screening algorithm onto the logic layer computing engine of the 3D memory; NATSA is a near-memory computing accelerator for time series analysis. NATSA implements the matrix profile algorithm on the accelerator, which is a latest algorithm for time series analysis entirely through PNM. The drawbacks of these two computing architectures are that their application fields are too single and lack generality. The present invention combines CGRA as the computing logic to achieve a relatively high level of both performance and generality of the architecture.
[0006] The research on coarse-grained reconfigurable architecture processors at home and abroad mainly focuses on optimizing for algorithm characteristics and reducing reconfiguration overhead, with less consideration of the impact of the memory system on computing performance and power consumption during the computing process. Y. Park, J. J. K. Park, and S. Mahlke studied the improvement of the energy efficiency of the overall computing architecture by the heterogeneous array structure in 2012 in terms of the heterogeneity, complexity, and integration method of the processing elements (PEs). Some researchers also studied the interconnection structure between PEs in the reconfigurable array, exploring the impact of different interconnection methods on the programmability, performance, power consumption, and area of the architecture. Z. Kwok and S. J. E. Wilton and Bouwens et al. explored the appropriate ratio of shared memory, global registers, local registers, and array scale size inside the array to optimize the combination of performance, power consumption, and area. It is similar to the structure of the present invention, but the focus of their research is on how to minimize the change of the existing DRAM circuit structure to combine CGRA with the 3D memory.
[0007] Most existing near-memory computing architectures use traditional sequential processors as the computing logic in the logic layer, unable to take advantage of the huge bandwidth of the storage layer in the near-memory computing architecture. A few architectures use dedicated accelerators as the near-memory computing logic layer, with performance improvement but limited generality of the architecture. Extremely few architectures use reconfigurable arrays as the computing logic, but they have problems such as the memory access range being limited to on-chip shared memory, not supporting indirect memory access, high requirements for compilers, and low utilization of processing units.
[0008] Therefore, those skilled in the art are committed to developing a near-memory computing system and construction method based on a data-driven coarse-grained reconfigurable array. Summary of the Invention
[0009] In view of the above defects of the prior art, the technical problems to be solved by the present invention are the problems of poor generality of the existing architecture, not supporting indirect memory access, high requirements for compilers, and low utilization of processing units.
[0010] A near-memory computing system based on a data-driven coarse-grained reconfigurable array, the computing system being a heterogeneous acceleration system, the system being divided into three levels, namely an off-chip master control layer, a logic layer of a three-dimensional accelerator, and a storage layer;
[0011] The off-chip master control layer consists of a main processor and a processor main memory. The main processor transports the data to be calculated from the processor main memory to the storage layer of the near-memory computing architecture through a bus, transports the configuration information to the configuration information registers of each reconfigurable array in the logic layer through a bus, sends the configuration task parameters to the configuration information scheduler of each reconfigurable array through a bus, and issues a start calculation signal through the bus after the transportation is completed, and the reconfigurable array starts to perform the calculation task;
[0012] The logic layer uses 16 coarse-grained reconfigurable arrays as the computing logic, and the arrays are connected to each memory controller through an internal bus to achieve access to different memory channels;
[0013] The memory controller in the logic layer is connected to the storage blocks in the storage layer through through-silicon vias to form a three-dimensional stacked accelerator to reduce the physical distance of memory access.
[0014] Furthermore, the coarse-grained reconfigurable array includes 64 processing units arranged in 8 rows and 8 columns, a shared memory, a memory access combiner, a global configuration information memory, and an array configuration information scheduler. The processing units are of heterogeneous design and are respectively responsible for data calculation and data access. Data transmission between processing units is completed by data routing between processing units, and the routing forms a Mesh on-chip network. The array configuration information scheduler of the coarse-grained reconfigurable array distributes the configuration information in the global configuration information register through the row bus. The memory access combiner and the shared memory are directly connected to the access units through the column bus to achieve direct access of the reconfigurable array to the memory and access to the on-chip memory. The array has two interaction interfaces. One of them connects the global configuration information register and the array configuration information scheduler to complete the interaction of configuration information, and the other interface is directly connected to the CrossBar bus and is responsible for data interaction between the array and the memory. The array configuration information scheduler dynamically distributes the configuration information in the global configuration information register through the processing unit status signal in the row bus, and distributes the configuration information to each processing unit through the row bus for task processing. The shared data memory transfers the preset data from the memory to the shared data memory through direct memory access, and then the array quickly accesses the data through the column bus. The memory access combiner collects the direct access of the LSUs in the array to the memory, exchanges data with the memory through the memory direct access interface, and distributes the data to each LSU by the internal logic after receiving the memory data. The first row and the last row of the array are each designed with 8 LSUs to correspond to the burst access length of the memory, and the highest bandwidth of the memory is fully utilized during the direct access to the memory. The 48 processing units in the middle 6 rows are arithmetic logic units and are responsible for data calculation according to the configuration information.
[0015] Furthermore, the processing units are divided into 2 types of structures, namely access units and arithmetic logic units;
[0016] The access unit includes a token buffer, an address generator, and a store / read configuration information queue, and has 7 external interfaces, namely the processing unit routing input and output interfaces, the array configuration information scheduler interface, the memory access interface, the memory reply interface, the shared memory access interface, and the shared memory reply interface, and the interface data width is 4 Bytes. The arithmetic logic unit includes a token buffer, an execution circuit, a data emission circuit, and 3 interfaces, namely the processing unit routing input and output interfaces and the configuration information input interface, and the interface data width is 4 Bytes.
[0017] Further, the token buffer is used to store the configuration information of the LSU, record the status of the current operand of the configuration information, send the operand of the corresponding operation to the address generator when all the operands of the configuration information are in the ready state, and change the corresponding operand in the buffer according to the operand increment information in the configuration information; the address generator calculates the input operand to generate a memory access address; the LSU selects the corresponding interface for memory access according to the memory access operation type. If the memory successfully receives the configuration information, the memory access configuration information is stored in the store / read configuration information queue; the store / read configuration information queue records the completion status of all the sent memory access configuration information. When the configuration information at the front of the queue is in the completion state, the store / read configuration information queue sends a data sending request to the processing unit routing interface according to the sending information part of the configuration information, and sends the memory access result to the destination processing unit; the operation configuration information of the LSU is divided into read operation and write operation. The operand 1 and operand 2 of the read operation are input to the address generator to generate a memory access address; the operand 1 of the write operation is the data to be stored, and the operand 2 and operand 3 are input to the address generator to generate a memory access address; there are two sources of the configuration information operand. One is the immediate number automatically managed by the token buffer, and the other is the output result of other processing units, which is determined by the operand-related field in the configuration information; the token buffer stores the configuration information of the ALU, records the operation of the configuration information and the status of the current operand, sends the operand of the corresponding operation to the execution circuit for calculation when all the operands of the configuration information are in the ready state. The execution circuit is designed in a pipelined manner, and one configuration information is emitted per cycle. When continuously performing the execution calculation configuration information, the calculation delay is hidden; the calculation result of the data emission circuit is taken out from the execution circuit, and the result is sent to the processing unit routing through the processing unit routing interface according to the result transmission configuration information in the configuration information. The source of the configuration information operand is the same as that of the access unit operand.
[0018] Further, the token buffer is provided with 4 configuration information buffer bits. When one or more pieces of configuration information cannot be executed due to unready operands, the processing unit preferentially processes other ready instructions. The 4 configuration information buffer bits are connected to the array configuration information scheduler of the processing unit through an interface. When a completion signal is sent from the configuration information completion signal interface, the array configuration information scheduler sets the configuration information for the processing unit through the configuration information register port. The configuration information buffer bits record the operation instruction OP of the configuration information, operand-related information, including operand matching information, operand ready status, operand arguments, the current iteration cycle IterID of the configuration information execution, the ID of the configuration information in the entire program, and the target address for transmitting the execution result. The token buffer determines whether the data transmitted through the processing unit routing data interface matches the currently cached configuration information according to the operand matching information. If they match, the operand is received. The token buffer introduces an arbiter for execution arbitration when multiple configuration information are ready. The arbiter determines the instruction to be executed in the next cycle according to the execution unit feedback signal interface and the ready status of each instruction in the instruction status table at this time, and outputs the corresponding instruction and operand through the instruction output port, preferentially executing the configuration information with a smaller iteration cycle.
[0019] Further, the shared memory adopts two sets of asymmetric ports, including the interface between the shared memory and the 3D memory and the interface between the shared memory and the reconfigurable array. The shared memory and the reconfigurable array are connected through a crossbar CrossBar. The shared memory adopts a multi-bank design, which is divided into 8 banks in total. Each bank is configured with a separate access port. The shared memory access can read in at most 8 32-bit operands. When a memory access conflict occurs, the arbiter will delay the conflicting configuration information by one cycle for execution.
[0020] Furthermore, the processing unit routing adopts a 16-channel design. Each channel corresponds to the corresponding token buffer ID and operand ID of the target processing unit. The channel selection is determined by the function DC_ID * 4 + OP_ID - 1. For example, if the output of the local processing unit is the operand 1 of the configuration information 0 cached in the token buffer of the target processing unit, then the 0th channel is used for transmission. The transmission direction of each channel is divided into 5 directions: east, south, west, north, and local. Corresponding input and output queues and interfaces, transmission controllers, and a 5x5 CrossBar are set for each direction. The processing unit sends a transmission request to the local interface of the processing unit routing through the corresponding interface. If the input queue in the local direction is not full, a receive signal is returned, and the request is added to the local input queue. Every cycle, the routing algorithm processor checks the transmission destination addresses of the first configuration information in all input queues, and decides the destination output queue through which the information is transmitted through the CrossBar according to the routing algorithm. When conflicts occur because the destination output queues of different input queues are the same output queue, the routing algorithm processor acts as an arbiter and delays one of the requests by one cycle for transmission. Every cycle, all output queues of the routing attempt to send the first packet in the output queue through the interfaces in the corresponding directions.
[0021] Furthermore, the structure of the global configuration information register adopts two sets of ports with an asymmetric design. When exchanging data with the main processor and memory, a 64-bit bus interface is used. After the request signal from the main processor arrives, a response signal is sent, indicating that the configuration information transfer is successfully executed through the DMA interface. Since the iteration-related configuration information, output configuration information, and input configuration information in the configuration information are dynamically specified by the configuration information scheduler during operation, for the 16 arrays in the entire architecture, the global configuration information register stores 1 copy of the configuration information, providing 66 bits of basic configuration information for each processing unit. When the global configuration information register interacts with the configuration information scheduler, a configuration information register output interface with a width of 528 bits is used. Each 528 bits of the internal unit of the global configuration information register is regarded as a unit, and the address of each unit is set with a corresponding id. When retrieving and transmitting, it is configured to be transmitted to the reconfigurable array PEA through the port according to the corresponding id.
[0022] Furthermore, the configuration information includes iteration cycle offset, immediate number self-increment, and branch instructions;
[0023] The iteration cycle offset enables the correct execution of configuration information with dependencies between iteration cycles;
[0024] The immediate number self-increment, in conjunction with the upper limit of the iteration count, enables the processing unit to automatically adjust the immediate number part of the instruction according to the configuration information before reaching the specified iteration count;
[0025] The branch instruction determines whether the operation is a branch based on the input operand 4 requirement bit in the configuration information. If the operand 4 is True, the operation is executed; if the operand 4 is False, the operation is not executed.
[0026] Implementation method of a near-memory computing system based on a data-driven coarse-grained reconfigurable array, including:
[0027] Step 1: Abstract the interface behavior, and implement the basic behaviors of the interface, including sending requests, receiving replies, and binding receiving ports, through C++ functions.
[0028] Step 2: Transplant models such as CPU, memory, and cross buses in the open-source platform, and add the port model implemented in Step 1 to the model.
[0029] Step 3: Use the vector data structure to implement the configuration information cache bit, and define arbitration functions, configuration information matching functions, and configuration information refresh functions.
[0030] Step 4: Utilize the interface defined in Step 1 and the token buffer implemented in Step 3 to abstract the 3-stage pipeline behavior of the processing unit into 3 functions, and implement the tick() function to sequentially call the 3 pipeline functions to achieve cycle simulation of 2 types of processing units.
[0031] Step 5: Use C++ data structures to implement the storage structures of shared memory, processing unit routing, and global configuration information registers; define the cycle behaviors of these components in the tick() function; add the interface defined in Step 1.
[0032] Step 6: Form an interconnection structure through the interface for all the components implemented in Steps 4 - 5 to form a reconfigurable array.
[0033] Step 7: According to the appendix Figure 1 Integrate all the modules in Steps 1 - 6 into the final near-memory reconfigurable array architecture.
[0034] Step 8: Write the configuration information, send the configuration information to 16 global configuration information registers through the bus, and run the program.
[0035] Technical effects
[0036] 1. Compared with the existing near-memory computing architectures, the present invention can give full play to the high-bandwidth advantage of near-memory computing and improve the energy efficiency ratio of the entire architecture.
[0037] 2. The present invention has better program generality compared with other reconfigurable architectures, is programmer-friendly, has a higher utilization rate of processing units, and has less dependence on compilers.
[0038] 3. The present invention utilizes a memory access unit to enable direct data exchange between the reconfigurable array and the memory unit during the computing process, expanding the memory access range and achieving indirect memory access.
[0039] 4. The present invention uses a processing unit with a dynamic execution structure to reduce the dependence of the reconfigurable architecture on the compiler and increase the generality of the program.
[0040] 5. The present invention uses a token buffer to implement the simultaneous execution of instructions in different iteration cycles, improving the utilization rate of the processing units in the reconfigurable array. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic diagram of the near-memory architecture of the present invention;
[0042] Figure 2 is a schematic diagram of the reconfigurable array architecture;
[0043] Figure 3 is a schematic diagram of the memory access unit and the arithmetic logic unit;
[0044] Figure 4 is a schematic diagram of the token buffer structure;
[0045] Figure 5 is a schematic diagram of the shared memory structure;
[0046] Figure 6 is a schematic diagram of the multi-channel processing unit routing structure;
[0047] Figure 7 is a schematic diagram of the configuration information register structure;
[0048] Figure 8 is a schematic diagram of the configuration information structure. DETAILED DESCRIPTION OF THE INVENTION
[0049] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification, making its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0050] In the drawings, components with the same structure are denoted by the same reference numerals, and components with similar structures or functions are denoted by similar reference numerals. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of some components in the drawings is appropriately exaggerated.
[0051] The overall architecture of the present invention is as shown in Figure 1As shown, the overall architecture is a near-memory processing architecture based on a dynamic CGRA. It is an overall heterogeneous acceleration system, which can be divided into three layers: the off-chip master control layer, the logic layer of the three-dimensional accelerator, and the storage layer. The off-chip master control layer consists of a main processor and the main memory of the processor. The main processor transfers the data to be calculated from the main memory of the processor to the storage layer of the near-memory computing architecture through the bus, transfers the configuration information to the configuration information registers of each reconfigurable array in the logic layer through the bus, sends the configuration task parameters to the configuration information scheduler of each reconfigurable array through the bus, and issues a start calculation signal through the bus after the transfer is completed. The reconfigurable array then starts the calculation task. The logic layer consists of 16 coarse-grained reconfigurable arrays as the calculation logic. The arrays are connected to each memory controller through an internal bus to achieve access to different memory channels. The memory controller in the logic layer is connected to the storage blocks in the storage layer using Through Silicon Via (TSV) technology to form a three-dimensional stacked accelerator, reducing the physical distance of memory access and providing high bandwidth and higher resource efficiency.
[0052] The coarse-grained reconfigurable array (Coarse-Grained Reconfigurable Architecture, CGRA) consists of 64 processing elements (Process Element, PE) arranged in 8 rows and 8 columns. The processing elements are of heterogeneous design, respectively completing data calculation and data access. The data transfer between the processing elements is responsible for by the data routing between PEs. The routing forms a Mesh on-chip network, as Figure 2As shown. The array configuration information scheduler of the array distributes the configuration information in the global configuration information register through the row bus. The memory access merger and the shared memory are directly connected to the access unit (Load / Store Unit, LSU) through the column bus to achieve the direct access of the reconfigurable array to the memory and the access to the on-chip memory. As shown in the figure, the array has two interaction interfaces. One of them connects the global configuration information register to the array configuration information scheduler to complete the interaction of the configuration information. The other interface is directly connected to the CrossBar bus and is responsible for the data interaction between the array and the memory. The array configuration information scheduler dynamically distributes the configuration information in the global configuration information register through the processing unit status signal in the row bus, and distributes the configuration information to each processing unit through the row bus for task processing. The shared data memory transports the preset data from the memory to the shared data memory through Direct Memory Access (DMA). Subsequently, the array quickly accesses the data through the column bus. The memory access merger collects the direct access of the LSU in the array to the memory, exchanges data with the memory through the memory direct access interface, and distributes the data to each LSU by the internal logic after receiving the memory data. The first row and the last row of the array are each designed with 8 LSUs to correspond to the burst access length of the memory, and to fully utilize the highest bandwidth of the memory during the direct access to the memory. The middle 6 rows, a total of 48 processing elements (PEs), are arithmetic logic units and are responsible for data calculation according to the configuration information.
[0053] The processing element (PE) is divided into 2 structures, namely the access unit and the arithmetic logic unit (Algorithm Logic Unit, ALU).
[0054] The access unit is as Figure 3As shown in the figure, it consists of a token buffer, an address generator, and a storage / reading configuration information queue. There are 7 external interfaces, namely the processing unit routing input and output interfaces, the array configuration information scheduler interface, the memory access interface, the memory reply interface, the shared memory access interface, and the shared memory reply interface. The interface data bit width is 4 Bytes. Among them, the token buffer stores the configuration information of the LSU, records the status of the current operand of the configuration information, and sends the operand corresponding to the operation to the address generator when all the operands of the configuration information are in the ready state, and changes the corresponding operand in the buffer according to the operand increment information in the configuration information. The address generator calculates the input operand to generate a memory access address. The LSU selects the corresponding interface for memory access according to the memory access operation type. If the memory successfully receives the configuration information, it stores the memory access configuration information in the storage / reading configuration information queue. The storage / reading configuration information queue records the completion status of all the issued memory access configuration information. When the configuration information at the front of the queue is in the completed state, the storage / reading configuration information queue sends a data sending request to the processing unit routing interface according to the sending information part of the configuration information, and sends the memory access result to the destination processing unit. The operation configuration information of the LSU is divided into read operation and write operation. For the read operation, operand 1 and operand 2 are input to the address generator to generate a memory access address; for the write operation, operand 1 is the data to be stored, and operand 2 and operand 3 are input to the address generator to generate a memory access address. There are 2 sources of configuration information operands. One is the immediate number automatically managed by the token buffer, and the other is the output result of other processing units, which is determined by the relevant fields of the operands in the configuration information.
[0055] The arithmetic logic unit is as Figure 3 As shown in the figure, it consists of a token buffer, an execution circuit, and a data emission circuit. There are 3 interfaces, namely the processing unit routing input and output interfaces and the configuration information input interface. The interface data bit width is 4 Bytes. The token buffer stores the configuration information of the ALU, records the execution operation of the configuration information and the status of the current operand, and sends the operand corresponding to the operation to the execution circuit for calculation when all the operands of the configuration information are in the ready state. The execution circuit is designed in a pipelined manner, and one configuration information can be emitted per cycle. When continuously performing execution calculation configuration information, the calculation delay can be hidden. The data emission circuit takes the calculation result from the execution circuit and sends the result to the processing unit routing through the processing unit routing interface according to the result transmission configuration information in the configuration information. The source of the configuration information operand is the same as that of the access unit operand.
[0056] The physical structure of the token buffer is as Figure 4As shown, to improve the utilization rate of the processing unit, 4 configuration information cache bits are set. When several pieces of configuration information cannot be executed because the operands are not ready, the processing unit can preferentially process other ready instructions. The 4 configuration information cache bits are connected to the array configuration information scheduler of the processing unit through the configuration information register port 401. When the configuration information completion signal interface 402 issues a completion signal, the array configuration information scheduler sets the configuration information for the processing unit through the configuration information register port 401. The configuration information cache bits record the operation instruction OP of the configuration information, operand-related information (operand matching information, operand ready status, operand arguments), the current iteration cycle IterID of the configuration information execution, the ID of the configuration information in the entire program, the target address for transmitting the execution result, etc. The token buffer determines whether the data transmitted by the processing unit routing data interface 403 matches the currently cached configuration information according to the operand matching information. If they match, the operand is received. The token buffer introduces an arbiter to perform execution arbitration when multiple configuration information is ready. The arbiter determines the instruction to be executed in the next cycle through the arbitration algorithm according to the execution unit feedback signal interface 404 and the ready status of each instruction in the instruction status table at this time, and outputs the corresponding instruction and operand through the instruction output port 405. The algorithm currently adopted by the present invention is to preferentially execute the configuration information with a smaller iteration cycle.
[0057] The design of the shared memory is similar to that of the global configuration information register, and two sets of asymmetric ports are adopted, as Figure 5 shown. The interface 501 between the shared memory and the three-dimensional memory has a width of 32 bits; the interface 502 between the shared memory and the reconfigurable array is connected to the reconfigurable array through a crossbar CrossBar. To improve the access speed of the shared memory, a multi-bank design is adopted, which is divided into 8 banks in total. Each bank is configured with a separate access port. In this way, each shared memory access can read in up to 8 32-bit operands at most, and can be further expanded. When a memory access conflict occurs, the arbiter will delay the execution of the conflicting configuration information by one cycle.
[0058] The processing unit routing is as Figure 6As shown in the figure, in order to completely avoid the routing transmission deadlock problem caused by the load balancing problem, the routing adopts a 16-channel design. Each channel corresponds to the corresponding token buffer ID and operand ID of the target processing unit. The channel selection is determined by the function DC_ID * 4 + OP_ID - 1. For example, if the output of the local processing unit is the operand 1 of the configuration information 0 cached in the token buffer of the target processing unit, then the 0th channel is used for transmission. Each operand of each configuration information is transmitted using a separate channel, ensuring the execution correctness of the dynamically reconfigurable array. The transmission direction of each channel is divided into 5 directions: east, south, west, north, and local. Corresponding input and output queues and interfaces are set for each direction. In addition, there is also a transmission controller and a 5x5 CrossBar. The processing unit sends a transmission request to the local interface of the processing unit routing through the corresponding interface. If the local input queue in the routing local direction is not full, a receive signal is returned, and the request is added to the local input queue. Every cycle, the routing algorithm processor checks the transmission destination address of the first configuration information in all input queues, and decides the destination output queue through which the information is transmitted through the CrossBar according to the routing algorithm. When conflicts occur where the destination output queues of different input queues are the same output queue, the routing algorithm processor acts as an arbiter and delays one of the requests by one cycle for transmission. Every cycle, all output queues of the routing attempt to send the first packet in the output queue through the interfaces in the corresponding directions.
[0059] The structure of the global configuration information register is as Figure 7 shown. Two sets of ports with an asymmetric design are adopted. When exchanging data with the main processor and memory, a 64-bit bus interface is used to ensure good compatibility of the system. After the request signal 701 of the main processor arrives, a response signal 702 is sent, indicating that the configuration information transfer is successfully executed through the DMA interface 703. Since the iteration-related configuration information, output configuration information, and input configuration information in the configuration information are dynamically specified by the configuration information scheduler during operation, for the 16 arrays in the entire architecture, the global configuration information register only needs to store 1 copy of the configuration information, and only needs to provide 66 bits of basic configuration information for each processing unit. When the global configuration information register interacts with the configuration information scheduler, a configuration information register output interface 704 with a width of 528 bits is used to meet the speed during reconstruction. Each internal unit of the global configuration information register takes 528 bits as a unit, and the corresponding id is set for the address of each unit. During retrieval and transmission, it is configured according to the corresponding id and transmitted to the reconfigurable array PEA through the 704 port.
[0060] The configuration information format of the present invention considers multiple aspects to ensure the generality and correctness of the program:
[0061] 1) Design of the upper limit of iteration times: In the previous execution process of CGRA, configuration information needed to be reread from the configuration information register in each machine cycle, resulting in a large amount of configuration information reading overhead and configuration information register overhead. However, the same configuration information was often reused during this calculation process. Designing an upper limit of iteration times for the configuration information can reduce the reading of invalid configuration information and, at the same time, reduce the configuration information register overhead.
[0062] 2) Iterative cycle offset: Introducing the concept of iterative cycles increases the execution efficiency of the array, but it causes operand matching to only accept results of the same iterative cycle. Therefore, an iterative cycle offset is introduced in the configuration information so that configuration information with dependencies between iterative cycles can be correctly executed.
[0063] 3) Immediate number auto-increment: In the previous configuration information of CGRA, most of the content was the same, and the difference was only in the slight change of the immediate number, which brought the overhead of rereading the configuration information and increased the storage overhead of the configuration information register. The present invention sets an immediate number auto-increment in the configuration information. Combined with the upper limit of iteration times, the processing unit automatically adjusts the immediate number part of the instruction according to the configuration information before reaching the specified iteration times, avoiding the invalid reading of the configuration information register and reducing the storage pressure of the configuration information register.
[0064] 4) Implementing branch instructions: To implement branch instructions, a fourth input operand is set for the configuration information. Whether this operation is a branch is determined by the demand bit of input operand 4 in the configuration information. If operand 4 is True, the operation is executed; if operand 4 is False, the operation is not executed.
[0065] On the above basis, the specific content of the 130 - bit configuration word is as Figure 8 shown:
[0066] 1) 801: Bits 129 - 114, a total of 16 bits, represent the upper limit of iteration times of the configuration information;
[0067] 2) 802: Bits 113 - 84, a total of 30 bits, represent 3 output directions reserved for the calculation / memory access results of the configuration information. Each direction is 10 bits. Among them, the high 6 bits represent the destination PE_ID, the middle 2 bits represent the token buffer channel ID of the target PE, and the last 2 bits represent the operand ID of the target PE.
[0068] 3) 803: Bits 83 - 82, a total of 2 bits, represent the number of output target PEs.
[0069] 4) 804: A total of 16 bits, representing the configuration information of logical judgment operand 4. When the operand is required, the high 8 bits represent the PE_ID, and the low 8 bits represent the possible offset between the execution cycle of the configuration information and the iteration cycle of the data source configuration information. This offset is generated by the cross-cycle dependence of the instruction.
[0070] 5) 805: Operand 4 represents a logical judgment operand, with a total of 1 bit, representing the requirement status of logical judgment operand 4. 0 indicates that the operand is not required, and 1 indicates that the source of this operand is other processing units.
[0071] 6) 806 / 808 / 810: A total of 16 bits, representing the configuration information of operand 3 / 2 / 1. When the source of the operand is other PEs, the high 8 bits represent the PE_ID, and the low 8 bits represent the possible offset between the execution cycle of the configuration information and the iteration cycle of the data source configuration information. This offset is generated by the cross-cycle dependence of the program. When the source of the operand is an immediate number, the high 8 bits represent the initial value of the immediate number, and the low 8 bits represent the self-increment of the immediate number.
[0072] 7) 807 / 809 / 811: A total of 1 bit, representing the source status of operand 3 / 2 / 1. 0 indicates that the source of this operand is an immediate number, and 1 indicates that the source of this operand is other processing units.
[0073] 8) 812: Bits 13 to 6, a total of 8 bits, representing the ID of this configuration information in the global register, used to match the dependencies between configuration information.
[0074] 9) 813: Bits 5 to 0, a total of 6 bits, representing the type of storage configuration information / arithmetic logic operation executed by the processing unit.
[0075] The present invention is finally implemented by modeling the hardware through C++. C++ simulates the behavior patterns of components and interconnects the components through interfaces. Finally, the components cooperate to perform simulations. The specific steps are as follows:
[0076] Step 1: Abstract the interface behavior and implement the basic behaviors of the interface (sending requests, receiving replies, binding receiving ports) through C++ functions.
[0077] Step 2: Transplant models such as CPUs, memories, and cross buses in the open-source platform, and add the port model implemented in Step 1 to the models.
[0078] Step 3: According to the appendix Figure 4 , use the vector data structure to implement the configuration information cache bits, and define arbitration functions, configuration information matching functions, and configuration information refreshing functions.
[0079] Step 4: According to the appendix Figure 3Scheme, add the interface defined in step 1 and the token buffer implemented in step 3, abstract the 3-stage pipelining behavior of the processing unit into 3 functions, implement the tick() function to call the 3 pipeline functions in sequence, and implement the cycle simulation of 2 types of processing units.
[0080] Step 5: According to Appendix Figure 5 , Appendix Figure 6 and Appendix Figure 7 Use C++ data structures to implement the storage structures of shared memory, processing unit routing, and global configuration information registers; define the cycle behavior of these components in the tick() function; add the interface defined in step 1.
[0081] Step 6: According to Appendix Figure 2 Form an interconnection structure for all the components implemented in steps 4 - 5 through the interface to form a reconfigurable array.
[0082] Step 7: According to Appendix Figure 1 Integrate all the modules in steps 1 - 6 into the final near-memory reconfigurable array architecture.
[0083] Step 8: Write the configuration information according to the configuration information format in Appendix Figure 8 Send the configuration information to 16 global configuration information registers through the bus and run the program.
[0084] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A near-memory computing system based on a data-driven coarse-grained reconfigurable array, characterized in that, the computing system is a heterogeneous acceleration system, which is divided into three levels, namely an off-chip master control layer, a logic layer of a three-dimensional accelerator, and a storage layer; the off-chip master control layer consists of a main processor and a processor main memory. The main processor transports the data to be calculated from the processor main memory to the storage layer of the near-memory computing architecture through a bus, transports the configuration information to the configuration information registers of each reconfigurable array in the logic layer through a bus, sends the configuration task parameters to the configuration information scheduler of each reconfigurable array through a bus, and issues a start calculation signal through the bus after the transportation is completed, and the reconfigurable array starts to perform the calculation task; the logic layer uses 16 coarse-grained reconfigurable arrays as computing logics. The arrays are connected to each memory controller through an internal bus to achieve access to different memory channels; the memory controller in the logic layer is connected to the storage blocks in the storage layer using through-silicon vias to form a three-dimensional stacked accelerator to reduce the physical distance of memory access; the coarse-grained reconfigurable array includes 64 processing units arranged in 8 rows and 8 columns, a shared memory, a memory access merger, a global configuration information memory, and an array configuration information scheduler. The processing units are of heterogeneous design and respectively complete data calculation and data access. The data transmission between the processing units is completed by the data routing between the processing units, and the routing forms a Mesh on-chip network; the array configuration information scheduler of the coarse-grained reconfigurable array distributes the configuration information in the global configuration information register through the row bus. The memory access merger and the shared memory are directly connected to the access units through the column bus to achieve the direct access of the reconfigurable array to the memory and the access to the on-chip memory; the array has two interaction interfaces, one of which connects the global configuration information register and the array configuration information scheduler to complete the interaction of the configuration information, and the other interface is directly connected to the CrossBar bus to be responsible for the data interaction between the array and the memory; the array configuration information scheduler dynamically distributes the configuration information in the global configuration information register through the processing unit status signal in the row bus, and distributes the configuration information to each processing unit through the row bus for task processing; the shared data memory transports the preset data from the memory to the shared data memory through direct memory access, and then the array quickly accesses the data through the column bus; the memory access merger collects the direct access of the LSU in the array to the memory, exchanges data with the memory through the memory direct access interface, and distributes the data to each LSU by the internal logic after receiving the memory data; the first row and the last row of the array are each designed with 8 LSUs to correspond to the burst access length of the memory, and the highest bandwidth of the memory is fully utilized during the direct access to the memory; the 48 processing units in the middle 6 rows are arithmetic logic units, which are responsible for data calculation according to the configuration information; the processing units are divided into 2 structures, namely access units and arithmetic logic units; The access unit includes a token buffer, an address generator, and a storage / read configuration information queue. It has 7 external interfaces, namely the processing unit routing input and output interfaces, the array configuration information scheduler interface, the memory access interface, the memory reply interface, the shared memory access interface, and the shared memory reply interface. The interface data width is 4 Bytes; The arithmetic logic unit includes a token buffer, an execution circuit, a data emission circuit, and 3 interfaces, namely the processing unit routing input and output interfaces and the configuration information input interface. The interface data width is 4 Bytes; The token buffer is used to store the configuration information of the LSU, record the status of the current operand of the configuration information, send the operand corresponding to the operation to the address generator when all the operands of the configuration information are in the ready state, and change the corresponding operand in the buffer according to the operand increment information in the configuration information; The address generator calculates the input operand to generate a memory access address; The LSU selects the corresponding interface for memory access according to the memory access operation type. If the memory successfully receives the access request, it stores the memory access configuration information in the storage / read configuration information queue; the storage / read configuration information queue records the completion status of all the sent memory access configuration information. When the configuration information at the front of the queue is in the completed state, the storage / read configuration information queue sends a data transmission request to the processing unit routing interface according to the transmission information part of the configuration information, and sends the memory access result to the destination processing unit; the operation configuration information of the LSU is divided into read operation and write operation. For the read operation, operand 1 and operand 2 are input to the address generator to generate a memory access address; for the write operation, operand 1 is the data to be stored, and operand 2 and operand 3 are input to the address generator to generate a memory access address; there are 2 sources of the configuration information operand. One is the immediate number automatically managed by the token buffer, and the other is the output result of other processing units, which is determined by the operand-related fields in the configuration information; The token buffer stores the configuration information of the ALU, records the operation of the configuration information and the status of the current operand, and sends the operand corresponding to the operation to the execution circuit for calculation when all the operands of the configuration information are in the ready state. The execution circuit is designed with pipelining, and one configuration information is emitted per cycle. When continuously performing execution calculation configuration information, the calculation delay is hidden; The data emission circuit takes the calculation result from the execution circuit and sends the result to the processing unit routing through the processing unit routing interface according to the result transmission configuration information in the configuration information. The source of the configuration information operand is the same as that of the access unit operand.
2. The near-memory computing system based on a data-driven coarse-grained reconfigurable array according to claim 1, characterized in that The token buffer is set with 4 configuration information buffer bits. When one or more pieces of configuration information cannot be executed because the operands are not ready, the processing unit preferentially processes other ready instructions. The 4 configuration information buffer bits are connected to the array configuration information scheduler of the processing unit through an interface. When the completion signal interface issues a completion signal for the configuration information, the array configuration information scheduler sets the configuration information for the processing unit through the configuration information register port. The configuration information buffer bits record the operation instruction OP of the configuration information, operand-related information, including operand matching information, operand ready status, operand arguments, the current iteration cycle IterID of the configuration information execution, the ID of the configuration information in the entire program, and the target address for transmitting the execution result. The token buffer determines whether the data transmitted through the processing unit routing data interface matches the currently cached configuration information according to the operand matching information. If it matches, the operand is received. The token buffer introduces an arbiter to perform execution arbitration when multiple configuration information is ready. The arbiter determines the instruction to be executed in the next cycle according to the execution unit feedback signal interface and the ready status of each instruction in the instruction status table at this time, and outputs the corresponding instruction and operand through the instruction output port, preferentially executing the configuration information with a smaller iteration cycle.
3. The near-memory computing system based on a data-driven coarse-grained reconfigurable array according to claim 2, characterized in that the shared memory adopts two sets of asymmetric ports, including the interface between the shared memory and the 3D memory and the interface between the shared memory and the reconfigurable array. The shared memory and the reconfigurable array are connected through a crossbar CrossBar. The shared memory adopts a multi-bank design, which is divided into 8 banks in total. Each bank is configured with a separate access port. The shared memory access can read in at most 8 32-bit operands. When a memory access conflict occurs, the arbiter will delay the conflicting configuration information by one cycle for execution.
4. The near-memory computing system based on a data-driven coarse-grained reconfigurable array according to claim 3, characterized in that the processing unit routing adopts a 16-channel design. Each channel corresponds to the corresponding token buffer ID and operand ID of the target processing unit. The channel selection is determined according to the function DC_ID * 4 + OP_ID - 1. The transmission direction of each channel is divided into 5 directions: east, south, west, north, and local. Each direction is provided with a corresponding input / output queue and interface, a transmission controller, and a 5x5 CrossBar; The processing unit sends a transmission request to the local interface routed by the corresponding interface to the processing unit. If the local input queue for routing in the local direction is not full, a receive signal is returned, and the request is added to the local input queue. Every cycle, the routing algorithm processor checks the transmission destination addresses of the first configuration information in all input queues, and decides the destination output queue through which the information is transmitted via the CrossBar according to the routing algorithm. When conflicts occur where the destination output queues of different input queues are the same output queue, the routing algorithm processor acts as an arbiter and delays the transmission of one of the requests by one cycle. Every cycle, all output queues of the routing attempt to send the first packet in the output queue through the interfaces in the corresponding directions.
5. The near-memory computing system based on a data-driven coarse-grained reconfigurable array according to claim 4, wherein the structure of the global configuration information register adopts two sets of ports with an asymmetric design. When exchanging data with the main processor and the memory, a 64-bit bus interface is adopted. After the request signal from the main processor arrives, a response signal is sent, indicating that the transmission successfully executes the configuration information transfer through the DMA interface. Since the iteration-related configuration information, output configuration information, and input configuration information in the configuration information are dynamically specified by the configuration information scheduler during operation, for the 16 arrays in the entire architecture, the global configuration information register stores 1 copy of the configuration information, providing 66 bits of basic configuration information for each processing unit. When the global configuration information register interacts with the configuration information scheduler, a configuration information register output interface with a width of 528 bits is adopted. Each internal unit of the global configuration information register takes 528 bits as a unit, and the address of each unit is set with a corresponding id. When retrieving and transmitting, it is configured to be transmitted to the reconfigurable array PEA through the port according to the corresponding id.
6. The near-memory computing system based on a data-driven coarse-grained reconfigurable array according to claim 5, wherein the configuration information includes an iteration cycle offset, an immediate number self-increment, and a branch instruction; the iteration cycle offset enables the correct execution of configuration information with dependencies between iteration cycles; the immediate number self-increment, in cooperation with the upper limit of the number of iterations, enables the processing unit to automatically adjust the immediate number part of the instruction according to the configuration information before reaching the specified number of iterations; the branch instruction determines whether the operation is a branch according to the demand bit of input operand 4 in the configuration information. If operand 4 is True, the operation is executed; if operand 4 is False, the operation is not executed.
7. A method for implementing a near-memory computing system based on a data-driven coarse-grained reconfigurable array according to claim 6, wherein, it includes: Step 1: Abstract the interface behavior, and implement the basic behaviors of the interface, including sending requests, receiving replies, and binding receive ports, through C++ functions; Step 2: Transplant models such as the CPU, memory, and cross bus in the open-source platform, and add the port model implemented in Step 1 to the model; Step 3: Implement the configuration information cache bit using the vector data structure, and define the arbitration function, the configuration information matching function, and the configuration information refreshing function; Step 4: Utilize the interface defined in Step 1 and the token buffer implemented in Step 3 to abstract the 3-stage pipelined behavior of the processing unit into 3 functions, and implement the tick() function to sequentially call the 3 pipelined functions to achieve the cycle simulation of 2 types of processing units; Step 5: Implement the storage structures of the shared memory, the processing unit routing, and the global configuration information register using the C++ data structure; define the cycle behavior of these components in the tick() function; add the interface defined in Step 1; Step 6: Form an interconnection structure through the interface for all the components implemented in Steps 4 - 5 to form a reconfigurable array; Step 7: Integrate all the modules in Steps 1 - 6 into the final near-memory reconfigurable array architecture according to Figure 1; Step 8: Write the configuration information, send the configuration information to the 16 global configuration information registers through the bus, and run the program.
Citation Information
Patent Citations
In-memory calculation method based on coarse-grained reconfigurable array
CN112463719A
Coarse-grained reconfigurable architecture system for large-scale MIMO signal detection
CN113055060A