A programmable computing and storage integrated acceleration array based on multi-stage pipeline
By designing a grid-like programmable array module and a multi-stage pipelined processing mode, a high-efficiency in-memory computing acceleration array was achieved, which improved computing speed and energy consumption control, and solved the problems of storage bandwidth bottleneck and low algorithm complexity.
Patent Information
- Application Number
- CN202210756356.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing in-memory computing designs cannot allocate computing resources reasonably according to specific algorithms, resulting in low algorithm complexity and high power consumption, with storage bandwidth becoming a computing bottleneck.
It adopts a grid-like programmable array module, which is interconnected with each other and with each unit through interconnected resources. It allows for programming, design of multi-level pipeline processing mode, supports multiple programmable array modules, and has a high degree of matching between algorithms and hardware in-memory computing acceleration array.
It improves the space utilization of computing and storage units, supports multi-level pipeline operation, has high computing speed and energy consumption control capabilities, and solves the problems of slow data transfer and storage bandwidth bottleneck.
Smart Images

Figure CN115098434B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer data processing technology, and specifically relates to a programmable in-memory computing acceleration array based on a multi-stage pipeline. Background Technology
[0002] With the development of cloud computing, artificial intelligence, and data preprocessing applications in recent years, slow data transfer and high energy consumption have become key bottlenecks in computing centers dealing with the massive data deluge. Retrieving data from external storage units often takes hundreds or thousands of times longer than computation time, with 60%-90% of the energy consumed in the process being wasted, resulting in very low energy efficiency. Storage bandwidth has become a major obstacle to data computing applications. For example, in accelerating neural network computations and radar data front-end preprocessing, the biggest challenge is the frequent movement of data between computing and storage units. Multi-stage pipelined programmable in-memory computing (IMC) acceleration arrays can solve this problem. Currently, in-memory computing designs do not allow software programmability. Once the hardware design is complete, the storage-computing relationship is fixed, making it impossible to rationally configure storage resources based on specific algorithms or control energy consumption based on these resources. Therefore, current designs result in limited algorithm complexity supported by storage resources and high power consumption for the same computing power. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multi-stage pipelined programmable in-memory computing acceleration array composed of grid-like programmable array modules, in-memory computing units within the modules and interconnected through interconnected resources, allowing programming, possessing high algorithm and hardware compatibility, and having high utilization of computing and storage unit space.
[0004] The objective of this invention is achieved through the following technical solution: a programmable in-memory computing acceleration array based on a multi-stage pipeline, comprising multiple programmable array modules and a power clock management unit, wherein the programmable array modules and the power clock management unit are connected via a management bus; the multiple programmable array modules are divided into a grid by connection resource lines, and each programmable array module is connected to the others by interconnection resources;
[0005] The interconnection resources consist of many metal wire segments. The interconnection resources are divided into single-length lines and double-length lines. Two spatially adjacent programmable array modules are connected by single-length lines, and non-adjacent programmable array modules are connected by double-length lines. Single-length lines or double-length lines of different grids are interconnected by matrix switches.
[0006] Furthermore, the programmable array module internally comprises input and output interfaces, multiple computing units, multiple storage units, internal wiring resources, a power control unit, and a delay unit. The input and output interfaces realize the input and output of the system clock, the input and output of module-level pipeline operation signals, and the input and output of data. The input and output interfaces, computing units, and storage units are connected through internal wiring resources. The power control unit is used to receive control signals sent by the power clock management unit and control the power on and off of the programmable array module. The delay unit is used to divide and delay the input system clock signal.
[0007] A computing unit can access one or more storage units, while a storage unit can only be accessed by one computing unit.
[0008] Within the programmable array module, each in-memory computing unit controls the timing of computation through a system clock and pipeline operation signals. In all in-memory computing units, the computation cycle is defined as m system clock cycles. A computation clock is generated through a delay unit. The rising edge of the computation clock samples the computation input and pipeline operation input signals, while the falling edge outputs the computation output and pipeline operation signals. The computation output signals of all in-memory computing units can be connected to the computation output signals of other in-memory computing units, and the pipeline operation signals of all in-memory computing units can be connected to the pipeline operation signals of other in-memory computing units. Through the pipeline operation signals and the system clock signal, the in-memory computing unit can perform computation in each computation clock cycle and notify the next-level in-memory computing unit to continue computation in the next computation cycle.
[0009] The data width of the storage unit should be greater than or equal to the data input width plus the width of the multiplication coefficient plus 1.
[0010] Furthermore, the power control unit is connected to the power clock management unit via a bus, which consists of a clock line and a data line, with the clock and data lines being at a high level when idle.
[0011] The bus communication protocol uses an address plus data format for transmission. When the start signal is a low level on the data line, a falling edge is generated on the clock line. When the confirmation signal is sent by the power clock management unit as an address signal or a data signal, the power control unit pulls the data line low on the rising edge of the clock. When the stop signal is a low level on the data line, a rising edge is generated on the clock line.
[0012] The address signal is 32 bits, with the high 16 bits representing the x-coordinate of the programmable block array grid and the low 16 bits representing the y-coordinate of the grid; the data bits are 8 bits, where 0 indicates that the power supply to the programmable array module is off, and 1 indicates that the power supply to the programmable array module is on; communication between the power clock management unit and the power control unit is only allowed, and communication requests between the power control unit and the power clock management unit are not allowed.
[0013] The power input enable signal of the programmable array module is controlled by the power control unit, which is used to control the power supply of each programmable array module.
[0014] Furthermore, the matrix switch achieves selective connection of horizontal and vertical metal lines through gating transistors, with a gating transistor set at the intersection of each horizontal and vertical metal line.
[0015] The beneficial effects of this invention are:
[0016] 1. Composed of a grid-like programmable array module, facilitating engineering design and implementation;
[0017] 2. The computing units are interconnected through wires, allowing for programming and possessing a high degree of algorithm and hardware compatibility; it also boasts high space utilization of computing and storage units.
[0018] 3. Supports multi-level pipeline operation, with high time utilization of computing and storage units, and each computing unit has high computing speed.
[0019] 4. A single programmable array module consists of multiple programmable memory units, and a single programmable array module has a high computing speed;
[0020] 5. This design contains multiple programmable array modules, and the programmable blocks are independent of each other. Therefore, this in-memory computing acceleration array has high computing speed while also being able to control energy consumption, resulting in a high energy consumption and computing power conversion rate.
[0021] 6. Each computing unit is equipped with one or more data storage units, which can solve the problems of slow data transfer and storage bandwidth bottlenecks. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the programmable in-memory computing acceleration array of the present invention;
[0023] Figure 2 This is a diagram of the programmable block array organization structure of the present invention;
[0024] Figure 3 This is a diagram showing the interconnection relationship between programmable blocks and interconnect resources;
[0025] Figure 4 This is a diagram showing the internal structure of the programmable array module of the present invention;
[0026] Figure 5 This is a pipeline structure diagram of a single in-memory computing unit of the present invention;
[0027] Figure 6 This is the data register cascading method of the present invention;
[0028] Figure 7 This is a timing diagram of the bus signals;
[0029] Figure 8 This is a diagram of a programmable matrix switch structure. Detailed Implementation
[0030] This invention reduces engineering design complexity by implementing a grid-like programmable array module. Interconnectivity resources are designed between modules and units, allowing software programming so that users can operate the actual hardware according to custom algorithms, fully utilizing hardware resources to complete algorithm calculations. A multi-stage pipelined processing mode is designed, ensuring that computing units are always in computational mode and storage units are always in data writing or reading mode, avoiding idle waiting states and improving the computational speed of individual in-memory units. Multiple programmable in-memory units are implemented within a single programmable array module, further increasing the computational speed of the individual module. Multiple programmable blocks are designed to form a programmable array, improving the overall design's computational speed. The programmable blocks are independent of each other, with identical input / output interfaces and power switches, enabling power consumption control.
[0031] This invention improves overall computing power while controlling energy consumption. The programmable portion of the design is programmed by the user, who can control the power supply of the programmable modules and define the interconnection operations of the in-memory computing units according to the specific software algorithm, maximizing the integration of software and hardware to achieve both improved computing power and energy consumption control. The technical solution of this invention is further described below with reference to the accompanying drawings.
[0032] like Figure 1 As shown, the present invention provides a programmable in-memory computing acceleration array based on a multi-stage pipeline, comprising multiple programmable array modules (hereinafter referred to as programmable blocks) and a power clock management unit. The programmable array modules and the power clock management unit are connected via a management bus. The multiple programmable array modules are divided into a grid by connection resource lines, with grid coordinates of (xa, yb). Each programmable array module is connected to the others by interconnection resources.
[0033] The interconnect resources consist of many metal wire segments. These interconnect resources are divided into single-length wires and double-length wires. Two spatially adjacent programmable array modules are connected via a single-length wire, while non-adjacent programmable array modules are connected via double-length wires, such as... Figure 2 As shown in the diagram, solid lines represent double-long lines, and dashed lines represent single-long lines. Single-long or double-long lines in different grids are interconnected via matrix switches, as shown below. Figure 3 As shown. A single long line is a single signal line; a double long line is a pair of differential lines, i.e., two signal lines. Double long lines are used for two programmable array modules that are not spatially adjacent, which can increase the anti-interference capability of signal transmission.
[0034] The programmable array module internally consists of input and output interfaces, multiple computing units, multiple storage units, internal wiring resources, a power control unit, and a delay unit. The input and output interfaces realize the input and output of the system clock, the input and output of module-level pipeline operation signals, and the input and output of data. The input and output interfaces, computing units, and storage units are connected through internal wiring resources. The power control unit is used to receive control signals sent by the power clock management unit and control the power on and off of the programmable array module. The delay unit is used to divide and delay the input system clock signal.
[0035] The combination of computing and storage units within the programmable module allows users to define the programming configuration, and this programming functionality is implemented via programming switches, such as... Figure 4 As shown. There are no strict requirements on the number of computation units and storage units; they do not need to be equal. There are two ways to combine computation units and storage units: one computation unit can access one storage unit, and one computation unit can access multiple storage units, but one storage unit can only be accessed by one computation unit. In the first case, for example, if the data input width is 16 bits, after processing, the maximum width might be 32 bits. In this case, the data width needs to be increased for storage, i.e., storage units are cascaded. The second case addresses scenarios where the computational data equals the stored data. In this scenario, overflow data from multiplication calculations is not processed; only the input and output data widths are retained.
[0036] Based on two combinations of computing units and storage units, a separate data storage unit is allocated to each computing unit, which contains a data register. The computing unit includes an addition coefficient unit, a multiplication coefficient unit, and a multiply-accumulate unit. Initialization input signals are input to the addition coefficient unit and the multiplication coefficient unit, respectively, and the outputs of the two coefficient units are input to the multiply-accumulate unit. The system clock signal, after frequency division and delay, is input to both the multiply-accumulate unit and the data register. The pipeline operation input signal is also input to the multiply-accumulate unit, which in turn inputs the pipeline operation output signal. The calculation input data is placed in the data register. The multiply-accumulate unit accesses the data register to obtain the input data, processes it, and then places it back into the data register. The data register outputs the calculation result. The specific process is as follows: Figure 5 As shown in the diagram. In some specific application scenarios, the multiplication and addition coefficients used by the computing units in each in-memory unit are different. When the system starts, the user needs to set the addition and multiplication coefficients for each in-memory unit. These multiplication and addition coefficients are provided by the user, i.e., the initialization input in the diagram. The first-level pipeline operation signal is input by the user or operator, and the last-level signal is received and processed by the user or operator.
[0037] These data storage units and paired computing units use independent data transmission lines, ensuring that only one computing unit accesses the data register at a time, eliminating access conflicts and access waits. All data write and read operations are executed immediately. Compared to traditional in-memory computing architectures, there is no time-sharing multiplexing of data registers, thus eliminating the speed bottleneck of storage bandwidth. During system initialization, the multiplication and addition coefficients of all computing units are initialized, and after initialization, the multiplication and addition coefficients remain unchanged.
[0038] Within the programmable array module, each in-memory computing unit controls the timing of computation through a system clock and pipelined operation signals. In all in-memory computing units, the computation cycle is defined as m system clock cycles. A delay unit generates the computation clock, and the rising edge of the clock samples the computation input and pipelined operation input signals, while the falling edge outputs the computation output and pipelined operation signals. The computation output signals of all in-memory computing units can be connected to the computation output signals of other in-memory computing units, and the pipelined operation signals of all in-memory computing units can be connected to the pipelined operation signals of other in-memory computing units. Through the pipelined operation signals and the system clock signal, the in-memory computing unit can perform computation in each computation clock cycle and notify the next-level in-memory computing unit to continue computation in the next computation cycle. The computation cycle is defined as the number of system clock cycles, determined by the number of clock cycles required for the multiply-accumulate unit overhead; theoretically, the computation cycle can be defined as one system clock cycle.
[0039] The data width of the storage unit should be greater than or equal to the data input width plus the width of the multiplication coefficient plus 1. The storage mode of the storage unit is defined according to requirements and can be either little-endian or big-endian. Here is an example of a cascading method for little-endian storage: the most significant bit (msb) of data register 1 is connected to the least significant bit (lsb) of data register 2, and so on. The data register cascading method is as follows: Figure 6 As shown.
[0040] The power clock management unit controls the system's power consumption. Depending on the specific algorithm, power consumption can be controlled by designing different system clock frequencies. The system clock drives data storage and multiply-accumulate operations. Because the system clock does not pass through the switching matrix, there is no time delay caused by the switching matrix. Therefore, a delay module is designed at the system clock input of each memory unit. This module delays the clock signal to facilitate clock and data synchronization.
[0041] The power control unit is connected to the power clock management unit via a bus. The bus consists of a clock line and a data line. The clock and data lines are at a high level when idle.
[0042] The bus communication protocol uses an address plus data format for transmission. When the start signal is a low level on the data line, a falling edge is generated on the clock line. When the confirmation signal is sent by the power clock management unit as an address signal or a data signal, the power control unit pulls the data line low on the rising edge of the clock. When the stop signal is a low level on the data line, a rising edge is generated on the clock line.
[0043] The address signal is 32 bits, with the high 16 bits representing the x-coordinate of the programmable block array's grid coordinates and the low 16 bits representing the y-coordinate. The data bits are 8 bits; a value of 0 indicates the programmable array module is powered off, and a value of 1 indicates the programmable array module is powered on. Figure 7 As shown. Only the power clock management unit is allowed to initiate communication with the power control unit; the power control unit is not allowed to initiate communication requests with the power clock management unit.
[0044] The power input enable signal of the programmable array module is controlled by the power control unit, which controls the power supply of each programmable array module. Array modules that are not in use are not powered, thus achieving the purpose of energy consumption control.
[0045] The matrix switch achieves selective connection of horizontal and vertical metal lines through gating transistors. A gating transistor is placed at the intersection of each horizontal and vertical metal line, such as... Figure 8As shown. The gating transistor can employ existing mature technologies, such as EEPROM transistors using floating gate technology or avalanche injection MOSFETs. When designing the gating transistor for this matrix switch, a power-down memory function needs to be implemented, i.e., the power-down retention of user-programmable functions. Single-long lines and double-long lines implement the user's programmable interconnect function through the matrix switch. Each time a signal passes through a programmable matrix switch, a delay is added. Therefore, a delay module is set in the system clock inside the programmable module to offset the impact of physical hardware delays. Furthermore, users are required to use adjacent programmable blocks as much as possible. The delay between interconnect resources and programmable blocks is controlled through strict programming prohibition principles: the in-memory units of adjacent pipelines must be in the same or adjacent programmable blocks.
[0046] The programmable in-memory computing acceleration array of this invention consists of a grid-like array of programmable blocks, reducing the difficulty of engineering implementation and facilitating design customization. The computing units are interconnected via wiring, allowing for programming and enabling users to customize storage and computing units. This design achieves high algorithm-hardware compatibility and efficient in-memory unit space utilization. Support for multi-stage pipelined operations further enhances in-memory unit space utilization, improving computational speed from the perspective of individual computing units. Each array of programmable blocks comprises multiple programmable in-memory units, resulting in high computational speed for each block. The design includes multiple independent arrays of programmable blocks, allowing for high computational speed while maintaining energy efficiency, achieving a high energy-to-computing power conversion rate. Utilizing the in-memory computing design, each computing unit is equipped with one or more data storage units, solving the problems of slow data transfer and storage bandwidth bottlenecks.
[0047] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A programmable in-memory computing acceleration array based on a multi-stage pipeline, characterized in that, It includes multiple programmable array modules and a power clock management unit, which are connected via a management bus; the multiple programmable array modules are divided into a grid by connection resource lines, and each programmable array module is connected to the others by interconnection resources. The interconnection resources consist of many metal wire segments. The interconnection resources are divided into single-length lines and double-length lines. Two spatially adjacent programmable array modules are connected by single-length lines, and non-adjacent programmable array modules are connected by double-length lines. Single-length lines or double-length lines of different grids are interconnected by matrix switches. The programmable array module internally consists of input and output interfaces, multiple computing units, multiple storage units, internal wiring resources, a power control unit, and a delay unit. The input and output interfaces realize the input and output of the system clock, the input and output of module-level pipeline operation signals, and the input and output of data. The input and output interfaces, computing units, and storage units are connected through internal wiring resources. The power control unit is used to receive control signals sent by the power clock management unit and control the power on and off of the programmable array module. The delay unit is used to divide and delay the input system clock signal. A computing unit can access one or more storage units, and a storage unit can only be accessed by one computing unit. Each computing unit is allocated a separate storage unit, which contains a data register; The calculation unit includes an addition coefficient unit, a multiplication coefficient unit, and a multiply-accumulate unit. Initialization input signals are input to the addition coefficient unit and the multiplication coefficient unit, respectively, and the outputs of the two coefficient units are input to the multiply-accumulate unit. The system clock signal, after frequency division and delay, is input to the multiply-accumulate unit and the data register, respectively. The pipeline operation input signal is also input to the multiply-accumulate unit, which in turn inputs the pipeline operation output signal. The calculation input data is placed in the data register. The multiply-accumulate unit accesses the data register to obtain the input data, processes it, and then places it back into the data register. The data register outputs the calculation result. Within the programmable array module, each in-memory computing unit controls the timing of computation through a system clock and pipeline operation signals. In all in-memory computing units, the computation cycle is defined as m system clock cycles. A delay unit generates the computation clock, and the rising edge of the clock samples the computation input and pipeline operation input signals, while the falling edge outputs the computation output and pipeline operation signals. The computation output signals of all in-memory computing units can be connected to the computation output signals of other in-memory computing units, and the pipeline operation signals of all in-memory computing units can be connected to the pipeline operation signals of other in-memory computing units. Through the pipeline operation signals and the system clock signal, the in-memory computing unit can perform computation in each computation clock cycle and notify the next-level in-memory computing unit to continue computation in the next computation cycle. The power control unit is connected to the power clock management unit via a bus. The bus consists of a clock line and a data line. The clock and data lines are at a high level when idle. The bus communication protocol uses an address plus data format for transmission. When the start signal is a low level on the data line, a falling edge is generated on the clock line. When the confirmation signal is sent by the power clock management unit as an address signal or a data signal, the power control unit pulls the data line low on the rising edge of the clock. When the stop signal is a low level on the data line, a rising edge is generated on the clock line. The address signal is 32 bits, with the high 16 bits representing the x-coordinate of the programmable block array grid and the low 16 bits representing the y-coordinate of the grid; the data bits are 8 bits, where 0 indicates that the power supply to the programmable array module is off, and 1 indicates that the power supply to the programmable array module is on; communication between the power clock management unit and the power control unit is only allowed, and communication requests between the power control unit and the power clock management unit are not allowed. The power input enable signal of the programmable array module is controlled by the power control unit, which is used to control the power supply of each programmable array module.
2. The programmable in-memory computing acceleration array based on a multi-stage pipeline according to claim 1, characterized in that, The data width of the storage unit should be greater than or equal to the data input width plus the width of the multiplication coefficient plus 1.
3. The programmable in-memory computing acceleration array based on a multi-stage pipeline according to claim 1, characterized in that, The matrix switch achieves selective connection of horizontal and vertical metal lines through gating transistors, with a gating transistor set at each intersection of the horizontal and vertical metal lines.
Citation Information
Patent Citations
Memory device for performing in-memory processing
CN114388012A
Field programmable gate array with distributed RAM and increased cell utilization
CN1194702A
Calculation device and control method therefor
JP2014052918A