Internal data caching and carrying circuit of storage and calculation integrated chip
Through the internal data cache handling circuit of the memory and computing integrated chip based on flash array, the unbalanced Xbar hybrid dimension mapping method is adopted to solve the communication bottlenecks and cache difficulties caused by data splicing in the memory and computing integrated chip, and achieve more efficient data processing.
Patent Information
- Application Number
- CN202510839943.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the integrated storage and computing chip, data splicing increases local communication pressure, forming a communication bottleneck, and due to changes in neural network structure, it becomes difficult to determine the number of caches, and traditional methods cannot effectively solve the data splicing problem.
The internal data cache and handling circuit of integrated memory and computing chip based on flash array is adopted. Through the unbalanced Xbar hybrid dimension mapping method, the input channel is sliced to achieve decoupling of the input channel and the output channel to avoid data splicing.
Reduce the amount of data per Tile, reduce local communication pressure and hotspots, optimize cache requirements, and improve data processing efficiency.
Smart Images

Figure CN120337833A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital integrated circuits, and particularly relates to an internal data cache transfer circuit of a memory-computation integrated chip, and specifically relates to an internal data cache transfer circuit of a memory-computation integrated chip based on a flash array for multiple neural network models. Background Art
[0002] Memory-computation integrated chips are an important technical path in the post-Moore era, and their development depends on the common breakthroughs in storage technology, circuit design, and algorithm-architecture optimization. Data stitching is a cancer that harms the memory-computation integrated architecture. On the one hand, data stitching will increase the local communication pressure, easily cause local communication hotspots and form communication bottlenecks. On the other hand, when multiplexing data is stitched, in order to avoid message-dependent deadlocks, the same number of receiving buffers (caches) need to be set at the local ports of the routing units where the stitching points are located. This will not only increase the hardware overhead, but also as the neural network structure changes, the position and fan-in of data stitching will also change, making it very difficult to determine the number of buffers in a homogeneous design. Under the traditional Xbar (crossbar) mapping method and communication planning method, data stitching is inevitable because the calculation results of Tiles responsible for different output channel ranges in the same layer need to be stitched.
[0003] Therefore, an internal data cache transfer circuit of a memory-computation integrated chip based on a flash array for multiple neural network models is proposed to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide an internal data cache transfer circuit of a memory-computation integrated chip based on a flash array, and a hybrid-dimensional mapping method for a non-uniform Xbar is proposed. By slicing the input channels, the decoupling of the input channels and output channels on the same Xbar is realized, thus avoiding data stitching.
[0005] To solve the above technical problems, the present invention provides an internal data cache transfer circuit of a memory-computation integrated chip based on a flash array, which is used for data processing of data streams of different layers of a neural network before flowing into the flash memory-computation array, and includes: A dynamic buffer unit, which is responsible for receiving the cast data stream from the on-chip network, caching the data into a multi-bank dynamic buffer, and when the data in the buffer can perform a convolution once, the data in the buffer will be input into the flash memory-computation array; during the convolution process, the feature image pixels required for the convolution process will be continuously released to reduce the buffer requirement by dynamically managing the buffer. The arbitration and cross - switch unit consists of an arbiter and two cross - switches. One cross - switch is responsible for transmitting the addresses in the address mapping unit to the dynamic buffer unit, and the strobe signal of this cross - switch directly comes from the output of the arbiter. The other cross - switch is responsible for transmitting the data in the dynamic buffer unit to the address mapping unit, and the strobe signal of this cross - switch comes from the result of delaying the output of the arbiter by one clock cycle. The address mapping unit slices the input channels to decouple the input channels and output channels on the same cross - switch. It receives data from the arbiter through the cross - switch, realizes the mapping between the feature image pixels and the addresses in the cache, and outputs this data to the data buffer.
[0006] Preferably, the width of the dynamic buffer unit is 16 channels * quantization bit width. During data storage, when the number of channels of the feature map is less than 16 channels or not an integer multiple of 16, the feature map should be supplemented to make the number of channels of the feature map reach an integer multiple of 16 to ensure storage alignment.
[0007] Preferably, the dynamic buffer unit also includes a preset buffer management strategy, which specifically includes: in the merged address lines, the lower bit bits are used to select the corresponding bank, and the higher bit bits are used to represent the corresponding row address; the order of data entering the buffer is in the order of channel - row - column.
[0008] Preferably, the address mapping unit also includes a preset address mapping strategy, which specifically includes: first, judge the relationship between the number of input channels and the word line WL in the cross - switch. When the number of input channels is less than the word line WL in the cross - switch, a per - channel mapping mechanism is adopted; when the number of input channels is greater than the number of columns of the cross - switch, the input channels are divided into multiple segments equal to the number of columns of the cross - switch, and the weights corresponding to each segment adopt the per - channel mapping mechanism respectively.
[0009] Preferably, the arbiter is a 3*3 matrix arbiter, and the 3*3 matrix arbiter specifically includes: AND gates AND1~AND9, OR gates OR1~OR3, selectors MUX1~MUX6, request signals Request0~Request2, disable signals Disable0~Disable2, and grant signals Grant0~Grant2; wherein the request signal Request0 is connected to the first input terminal of AND gate AND1, the first input terminal of AND gate AND2, and the first input terminal of AND gate AND3; the request signal Request1 is connected to the first input terminal of AND gate AND4, the first input terminal of AND gate AND5, and the first input terminal of AND gate AND6; the request signal Request2 is connected to the first input terminal of AND gate AND7, the first input terminal of AND gate AND8, and the first input terminal of AND gate AND9; wherein the output of AND gate AND1 outputs the grant signal Grant0, and at the same time is connected to the second input terminal of selector MUX1, the second input terminal of selector MUX2, the first input terminal of selector MUX3, and the first input terminal of selector MUX5; the second input terminal of AND gate AND1 is connected to the disable signal Disable0, the output terminal of OR gate OR1 outputs the disable signal Disable0, and the two input terminals of OR gate OR1 are respectively connected to the output terminals of AND gate AND4 and AND gate AND7; the second input terminal of AND gate AND7 is connected to the output terminal of selector MUX5, the second input terminal of selector MUX5 is connected to the second input terminal of selector MUX6, the first input terminal of selector MUX4, and the first input terminal of selector MUX2; the second input terminal of selector MUX3 is connected to the output terminal of AND gate AND5, the first input terminal of selector MUX1, the first input terminal of selector MUX6, and the second input terminal of selector MUX4, and the output terminal of AND gate AND5 outputs the grant signal Grant1; the two input terminals of OR gate OR2 are respectively connected to the output terminals of AND gate AND8 and AND gate AND2, the output terminal of OR gate OR2 outputs the disable signal Disable1, and is connected to the second input terminal of AND gate AND5; the second input terminal of AND gate AND8 is connected to the output terminal of selector MUX6; the second input terminal of AND gate AND2 is connected to the output terminal of selector MUX1; the second input terminal of AND gate AND6 is connected to the output terminal of selector MUX4; the two input terminals of OR gate OR3 are respectively connected to the output terminals of AND gate AND6 and AND gate AND3, the output terminal of OR gate OR3 outputs the disable signal Disable2, and is connected to the second input terminal of AND gate AND9, and the output terminal of AND gate AND9 outputs the grant signal Grant2; the second input terminal of AND gate AND3 is connected to the output terminal of selector MUX2.
[0010] Preferably, in the 3×3 matrix arbiter, if the value at the Wij position of the matrix is 1, it represents that the priority of i is greater than that of j; if it is 0, it represents that the priority of i is less than that of j. After each arbitration is successful, the value of the matrix Wij is updated. The row where the successfully arbitrated request is located is set to 0, and the column where it is located is set to 1. The arbitration function is realized by maintaining the value of the matrix Wij.
[0011] Preferably, the priorities of the arbitration requests are: Request signal Request0 > Request signal Request1 > Request signal Request2.
[0012] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses an internal data cache transfer circuit based on a flash array for multiple neural network models. This circuit is for the data processing process before the data stream of different layers of the neural network flows into the flash memory computing array, including operations such as data caching and data transfer. The data is sliced through the input channel and transferred to the flash memory computing array according to the number of data channels and the number of pixel sizes of the current neural network layer. By proposing a hybrid dimension mapping method (i.e., address mapping strategy) for non-uniform Xbar, when performing the hybrid dimension mapping mechanism, first judge the relationship between the number of input channels and the Xbar Word Line (WL). When the number of input channels is less than the Xbar WL, a per-channel mapping mechanism is adopted; when the number of input channels is greater than the Xbar column number, the input channels are divided into multiple segments equal to the Xbar column number, and per-channel mapping is respectively adopted for the weights corresponding to each segment. By slicing the input channels, the decoupling of the input channels and output channels on the same Xbar is achieved, thus avoiding data stitching. By avoiding data stitching, the amount of data received by each Tile is reduced by half. Therefore, avoiding data stitching is beneficial to reducing the local communication pressure and local hot spots. Description of the Drawings
[0013] Figure 1 is the circuit block diagram of the data cache and transfer provided by the present invention.
[0014] Figure 2 is the schematic diagram of the input data storage format provided by the present invention.
[0015] Figure 3 is the schematic diagram of the buffer management strategy provided by the present invention.
[0016] Figure 4 is the circuit diagram of the 3×3 matrix arbiter provided by the present invention.
[0017] Figure 5 is the schematic diagram of the pixel point coordinates (x, y, c) provided by the present invention. Detailed Embodiment
[0018] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0019] As Figure 1 shown, the embodiment of the present invention specifically provides an internal data cache transfer circuit for a memory-computation integrated chip based on a flash array, which completes the data processing process of the data streams of different layers of the neural network before flowing into the flash memory-computation array, including operations such as data caching and data transfer. The data is sliced through the input channel and transferred to the flash memory-computation array according to the number of data channels and the number of pixel sizes of the current neural network layer, and mainly consists of three parts: a dynamic buffer unit, an arbitration and cross-switch unit, and an address mapping unit.
[0020] The dynamic buffer unit is responsible for receiving the cast data stream from the on-chip network, caching the data into the multi-bank dynamic buffer, and when the data in the buffer is sufficient to perform a convolution, the data in the buffer will be input into the flash memory-computation array. During this process, the feature image pixels that need to be released due to the completion of the convolution will be continuously released, and the cache requirement will be reduced by dynamically managing the buffer. The input data storage format is as Figure 2 shown. Since 16 channels are bound during data transmission, in order to ensure symmetry, the width of the buffer here is also set to 16 channels * quantization bit width. In the example Figure 2 , each color represents 16 channels. In the actual storage process, the feature maps with less than 16 channels and non-integer multiples of 16 channels should be supplemented to 16 channels for alignment in storage. Here, a 3*3 size, stride-1 convolution window is used as an example to illustrate the buffer management strategy: Figure 3 In it, the pixel points marked with "X" are the pixels that have been released when performing the black convolution window, the pixel points marked with "0" and "√" are the pixels that must be maintained in the buffer when performing the current convolution, the pixel points marked with "." are the pixels that will be released after performing this convolution, and the pixel points marked with "-" are the pixels that may exist in the buffer when performing the current convolution. The multi-bank buffer can be regarded as a large buffer. In the combined address line, the lower bits are used to select the corresponding bank (memory bank), and the higher bits indicate the corresponding row address. Combining the above buffer management strategy, it can be seen that the storage form change of a frame of feature map in the buffer over time is as Figure 3 . The gray part is the position where the valid data is stored in the buffer. In this way, the capacity requirement of the buffer will be much smaller than the buffer required to cache the entire frame of image.
[0021] An arbitration and crossbar unit, which consists of an arbiter and two crossbars. Here, the arbiter circuit uses a matrix arbiter. Figure 4 Taking a 3*3 matrix arbiter as an example to illustrate the function of the arbiter. If the value at the Wij position is 1, it means that the priority of i is greater than that of j; if it is 0, it means that the priority of i is less than that of j. After each successful arbitration, the value of the W matrix is updated. The row where the successfully arbitrated request is located is set to 0, and the column is set to 1. The arbitration function is realized by maintaining the value of the W matrix. The arbitration request is the one-hot code generated by decoding the lower 2 bits of all read pointers. The generated arbitration success signal will be used as the strobe signal of the crossbar. The internal circuit structure of the crossbar consists of multiple multiplexers (mux). In this circuit, there are two crossbars. One is responsible for transmitting addresses, and the other is responsible for transmitting data. Due to the register out characteristic of the RAM, the address and data are in adjacent two cycles. Therefore, the strobe signal sel of the address crossbar directly comes from the output of the arbitration, while the strobe signal sel of the data crossbar comes from the result after delaying the output of the arbiter by one clock cycle. The reason why this circuit structure can be adopted here is that the address mapping strategy design in the subsequent address mapping unit can ensure that the data winning the arbitration can be transmitted to the specified position without blocking.
[0022] The arbiter is a 3*3 matrix arbiter, and the 3*3 matrix arbiter specifically includes: AND gates AND1~AND9, OR gates OR1~OR3, selectors MUX1~MUX6, request signals Request0~Request2, disable signals Disable0~Disable2, and grant signals Grant0~Grant2. Among them, the request signal Request0 is connected to the first input end of AND gate AND1, the first input end of AND gate AND2, and the first input end of AND gate AND3; the request signal Request1 is connected to the first input end of AND gate AND4, the first input end of AND gate AND5, and the first input end of AND gate AND6; the request signal Request2 is connected to the first input end of AND gate AND7, the first input end of AND gate AND8, and the first input end of AND gate AND9. Among them, the output of AND gate AND1 outputs the grant signal Grant0, and at the same time is connected to the second input end of selector MUX1, the second input end of selector MUX2, the first input end of selector MUX3, and the first input end of selector MUX5; the second input end of AND gate AND1 is connected to the disable signal Disable0, the output end of OR gate OR1 outputs the disable signal Disable0, and the two input ends of OR gate OR1 are respectively connected to the output ends of AND gate AND4 and AND gate AND7; the second input end of AND gate AND7 is connected to the output end of selector MUX5, the second input end of selector MUX5 is connected to the second input end of selector MUX6, the first input end of selector MUX4, and the first input end of selector MUX2; the second input end of selector MUX3 is connected to the output end of AND gate AND5, the first input end of selector MUX1, the first input end of selector MUX6, and the second input end of selector MUX4, and the output end of AND gate AND5 outputs the grant signal Grant1; the two input ends of OR gate OR2 are respectively connected to the output ends of AND gate AND8 and AND gate AND2, the output end of OR gate OR2 outputs the disable signal Disable1, and is connected to the second input end of AND gate AND5; the second input end of AND gate AND8 is connected to the output end of selector MUX6; the second input end of AND gate AND2 is connected to the output end of selector MUX1; the second input end of AND gate AND6 is connected to the output end of selector MUX4; the two input ends of OR gate OR3 are respectively connected to the output ends of AND gate AND6 and AND gate AND3, the output end of OR gate OR3 outputs the disable signal Disable2, and is connected to the second input end of AND gate AND9, and the output end of AND gate AND9 outputs the grant signal Grant2; the second input end of AND gate AND3 is connected to the output end of selector MUX2.
[0023] The working principle of this matrix arbitration circuit is as follows: When the three input request signals Request0~Request2 simultaneously request access to the address or data, one requester is selected to obtain the access right, that is, corresponding to Grant0~Grant2, to avoid conflicts, and the priority judgment is realized by using a combinational logic circuit.
[0024] Priority order: Request0 > Request1 > Request2 (that is, the 0th requester has the highest priority).
[0025] Logical expression: Grant0 = Request0 (As long as Request0 is high, it is immediately authorized, ignoring other requests).
[0026] Grant1 =!Request0 & Request1 (When Request0 is invalid, Request1 can be authorized).
[0027] Grant2 =!Request0 &!Request1 & Request2 (Only when both Request0 and Request1 are invalid, Request2 is authorized).
[0028] The address mapping unit mainly realizes the mapping between the feature image pixels and the addresses in the cache. As Figure 5 shown, taking an image with H rows, W columns, and C channels (16 channels) as an example to illustrate how to find the position of a certain pixel point (x, y, c) in the cache address when the coordinates of a certain pixel point are known. The data enters the cache in the order of channel - row - column. Therefore, y - 1 rows of data have been input before (x, y, 0), and x - 1 channels of data have been input in the yth row. For (x, y, c), there are C - 1 more inputs of the same position but different channels. Since the buffer has frame periodicity, the physical address of a pixel in the buffer is:
[0029] According to the formula, it is known that buffer_size, W, and C are all fixed, while x and y will change as the sliding window moves. Since the coordinates of the lower right corner of the convolution window have been calculated in the previous buffer management strategy, and the size of the convolution window is fixed. For flash, after mapping a network, the position of the input corresponding to a specific BL line in the convolution window in the flash is also fixed. Therefore, we can calculate the addresses required by the flash input BL one by one. Moreover, many consecutive pixels in the convolution window also have consecutive characteristics in the cache. Using these characteristics, the configuration information can be compressed.
[0030] The configuration information only needs to include X_pos, Y_pos, and start_ch, which are the relative positions of the BL input in the convolutional window, and ch_num, which is the number of consecutive pixels after this pixel that can be sequentially transmitted on adjacent BLs. To increase the data parallelism, an address mapping unit equivalent to the number of banks is used here to fetch data in parallel from multiple banks, maximizing the advantages of multiple banks.
[0031] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure shall fall within the scope of protection of the claims.
Claims
1. An in-memory computing chip internal data cache transfer circuit, which is used for data processing of data streams of different layers of a neural network before flowing into a flash in-memory computing array, and is characterized in that, Including: A dynamic buffer unit, which is responsible for receiving the cast data stream from the on-chip network, caching the data into a multi-bank dynamic buffer, and when the data in the buffer can perform a convolution once, the data in the buffer will be input to the flash memory and computing array; during the convolution process, the feature image pixels required for the convolution process will be continuously released to reduce the buffer requirement by dynamically managing the buffer. An arbitration and crossbar unit, which consists of an arbiter and two crossbars; one of the crossbars is responsible for transmitting the address in the address mapping unit to the dynamic buffer unit, and the strobe signal of this crossbar directly comes from the output of the arbiter; the other crossbar is responsible for transmitting the data in the dynamic buffer unit to the address mapping unit, and the strobe signal of this crossbar comes from the result after delaying the output of the arbiter by one clock cycle. An address mapping unit, which decouples the input channel and the output channel on the same crossbar by slicing the input channel; receives the data from the arbiter through the crossbar, realizes the mapping between the feature image pixels and the addresses in the buffer, and outputs the data to the data buffer.
2. The internal data cache transfer circuit of the in-memory computing chip according to claim 1, wherein The width of the dynamic buffer unit is 16 channels * quantization bit width, and when the data is stored, when the number of channels of the feature map is less than 16 channels or not an integer multiple of 16, the feature map should be supplemented so that the number of channels of the feature map reaches an integer multiple of 16 to ensure storage alignment.
3. The internal data cache transfer circuit of a memory - in - computing chip according to claim 2, wherein The dynamic buffer unit also includes a preset buffer management strategy, specifically including: in the merged address line, the low bit is used to select the corresponding bank, and the high bit is used to represent the corresponding row address; the data enters the buffer in the order of channel - row - column.
4. The internal data cache transfer circuit of a memory - in - computing chip according to claim 1, characterized in that, The address mapping unit also includes a preset address mapping strategy, specifically including: first, judge the relationship between the number of input channels and the word line WL in the crossbar. When the number of input channels is less than the word line WL in the crossbar, a per-channel mapping mechanism is adopted; when the number of input channels is greater than the number of columns of the crossbar, the input channels are divided into multiple segments equal to the number of columns of the crossbar, and the weights corresponding to each segment are respectively mapped using the per-channel mapping mechanism.
5. The internal data cache transfer circuit of an in-memory computing chip according to claim 1, characterized in that, The arbiter is a 3*3 matrix arbiter, and the 3*3 matrix arbiter specifically includes: AND gates AND1~AND9, OR gates OR1~OR3, selectors MUX1~MUX6, request signals Request0~Request2, disable signals Disable0~Disable2, and grant signals Grant0~Grant2; among them, the request signal Request0 is connected to the first input terminal of AND gate AND1, the first input terminal of AND gate AND2, and the first input terminal of AND gate AND3; the request signal Request1 is connected to the first input terminal of AND gate AND4, the first input terminal of AND gate AND5, and the first input terminal of AND gate AND6; the request signal Request2 is connected to the first input terminal of AND gate AND7, the first input terminal of AND gate AND8, and the first input terminal of AND gate AND9; among them, the output of AND gate AND1 outputs the grant signal Grant0, and at the same time is connected to the second input terminal of selector MUX1, the second input terminal of selector MUX2, the first input terminal of selector MUX3, and the first input terminal of selector MUX5; the second input terminal of AND gate AND1 is connected to the disable signal Disable0, the output terminal of OR gate OR1 outputs the disable signal Disable0, and the two input terminals of OR gate OR1 are respectively connected to the output terminals of AND gate AND4 and AND gate AND7; the second input terminal of AND gate AND7 is connected to the output terminal of selector MUX5, the second input terminal of selector MUX5 is connected to the second input terminal of selector MUX6, the first input terminal of selector MUX4, and the first input terminal of selector MUX2; the second input terminal of selector MUX3 is connected to the output terminal of AND gate AND5, the first input terminal of selector MUX1, the first input terminal of selector MUX6, and the second input terminal of selector MUX4, and the output terminal of AND gate AND5 outputs the grant signal Grant1; the two input terminals of OR gate OR2 are respectively connected to the output terminals of AND gate AND8 and AND gate AND2, the output terminal of OR gate OR2 outputs the disable signal Disable1, and is connected to the second input terminal of AND gate AND5; the second input terminal of AND gate AND8 is connected to the output terminal of selector MUX6; the second input terminal of AND gate AND2 is connected to the output terminal of selector MUX1; the second input terminal of AND gate AND6 is connected to the output terminal of selector MUX4; the two input terminals of OR gate OR3 are respectively connected to the output terminals of AND gate AND6 and AND gate AND3, the output terminal of OR gate OR3 outputs the disable signal Disable2, and is connected to the second input terminal of AND gate AND9, and the output terminal of AND gate AND9 outputs the grant signal Grant2; the second input terminal of AND gate AND3 is connected to the output terminal of selector MUX2.
6. The internal data cache transfer circuit of a memory - in - computing chip according to claim 5, characterized in that, In the 3*3 matrix arbiter, the value at the matrix Wij position, if it is 1, represents that the priority of i is greater than j; if it is 0, it represents that the priority of i is less than j; after each arbitration is successful, the value of the matrix Wij is updated, the row where the successfully arbitrated request is located is set to 0, and the column is set to 1, and the arbitration function is realized by maintaining the value of the matrix Wij.
7. The internal data cache transfer circuit of a memory - in - computing chip according to claim 6, characterized in that, The priorities of the requests for arbitration are as follows: Request signal Request0 > Request signal Request1 > Request signal Request2.
Citation Information
Patent Citations
One-cycle router on chip based on quick path technology
CN102185751A
High-order router row buffer optimization structure
CN108111438A
Multicast implementation method and device based on Tile exchange architecture
CN117714392A
Data processing system and method for testing a data processor having a cache memory
US5586279A