A data cache transfer circuit inside a storage and computing integrated chip
Through the internal data cache and handling circuit of the integrated memory and computing chip based on flash array, the unbalanced Xbar hybrid dimension mapping method is adopted to solve the communication bottlenecks and cache difficulties caused by data splicing, and achieve more efficient data processing.
Patent Information
- Application Number
- CN202510839943.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the integrated storage and computing chip, data splicing increases local communication pressure, forms a communication bottleneck, and it is difficult to determine the number of caches in isomorphic design, which is difficult to effectively solve in the existing technology.
The internal data cache and handling circuit of integrated memory and computing chip based on flash array is adopted. Through the unbalanced Xbar hybrid dimension mapping method, the input channel is sliced to achieve decoupling of the input channel and the output channel to avoid data splicing.
Reduce the amount of data per Tile, reduce local communication pressure and hotspots, optimize cache requirements, and improve data processing efficiency.
Smart Images

Figure CN120337833B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital integrated circuits, and in particular relates to a data cache transport circuit inside a storage and computing integrated chip, and specifically relates to a data cache transport circuit inside a storage and computing integrated chip based on a flash array for multiple neural network models. Background Art
[0002] Integrated storage and computing chips are a key technological path in the post-Moore era. Their development relies on breakthroughs in storage technology, circuit design, and algorithm-architecture optimization. Data splicing is a major threat to integrated storage and computing architectures. On the one hand, data splicing increases local communication pressure, easily creating hotspots and bottlenecks. On the other hand, when splicing multiple data paths, to avoid message dependency deadlock, an equal number of receive buffers must be set up on the local ports of the routing units where the splicing points are located. This not only increases hardware overhead, but also changes the location and fan-in of data splicing as the neural network structure evolves, making it very difficult to determine the number of buffers in homogeneous designs. Under traditional Xbar (crossbar) mapping and communication planning methods, data splicing is inevitable because the computation results of tiles in the same layer that are responsible for different output channel ranges need to be spliced.
[0003] Therefore, a data cache transfer circuit inside a flash array-based storage and computing integrated chip for various neural network models is proposed to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide a data cache transfer circuit within a storage and computing integrated chip based on a flash array. A mixed-dimensional mapping method for non-balanced Xbar is proposed. By slicing the input channels, the input and output channels on the same Xbar are decoupled, thereby avoiding data splicing.
[0005] To solve the above technical problems, the present invention provides a data cache transfer circuit inside a flash array-based storage and computing integrated chip, which is used to process data streams from different layers of a neural network before they flow into the flash storage and computing array, including:
[0006] The dynamic cache unit is responsible for receiving the cast data stream from the on-chip network and caching the data in the multi-bank dynamic cache. When the data in the cache can be used for a convolution, the data in the cache will be input into the flash storage array. During the convolution process, the feature map pixels required to complete the convolution process will be continuously released to reduce the cache demand by dynamically managing the cache.
[0007] The arbitration and crossbar switch unit consists of an arbitrator and two crossbar switches. One crossbar switch is responsible for transmitting the address in the address mapping unit to the dynamic cache unit, and the selection signal of this crossbar switch comes directly from the output of the arbitrator. The other crossbar switch is responsible for transmitting the data in the dynamic cache unit to the address mapping unit, and the selection signal of this crossbar switch comes from the result of a beat of the arbitrator output.
[0008] The address mapping unit slices the input channels to decouple the input channels and output channels on the same cross switch. It receives data from the arbitrator through the cross switch, maps the feature map pixels to the addresses in the cache, and outputs the data to the data cache.
[0009] Preferably, the width of the dynamic cache unit is 16 channels * quantization bit width, and during the data storage process, when the number of channels of the feature map is less than 16 channels or is not an integer multiple of 16, the feature map should be supplemented so that the number of channels of the feature map reaches an integer multiple of 16 to ensure storage alignment.
[0010] Preferably, the dynamic cache unit also includes a preset cache management strategy, specifically including: in the merged address line, the lower bit is used to select the corresponding bank, and the higher bit is used to represent the corresponding row address; the order in which data enters the cache is in the order of channel-row-column.
[0011] Preferably, the address mapping unit further includes a preset address mapping strategy, specifically including: first determining the relationship between the number of input channels and the word lines WL in the crossbar switch; when the number of input channels is less than the word lines WL in the crossbar switch, adopting a channel-by-channel mapping mechanism; when the number of input channels is greater than the number of crossbar switch columns, dividing the input channels into multiple segments equal to the number of crossbar switch columns, and applying a channel-by-channel mapping mechanism to the weights corresponding to each segment.
[0012] Preferably, the arbiter is a 3*3 matrix arbiter, and the 3*3 matrix arbiter specifically includes: AND gates AND1~AND9, OR gates OR1~OR3, selectors MUX1~MUX6, request signals Request0~Request2, disable signals Disable0~Disable2 and grant signals Grant0~Grant2; wherein the request signal Request0 is connected to input terminal 1 of AND gate AND1, input terminal 1 of AND gate AND2 and input terminal 1 of AND gate AND3; the request signal Request1 is connected to input terminal 1 of AND gate AND4, input terminal 2 of AND gate AND5. The request signal Request2 is connected to the input terminal 1 of the AND gate AND7, the input terminal 1 of the AND gate AND8 and the input terminal 1 of the AND gate AND9; the output of the AND gate AND1 outputs the authorization signal Grant0, and is simultaneously connected to the input terminal 2 of the selector MUX1, the input terminal 2 of the selector MUX2, the input terminal 1 of the selector MUX3 and the input terminal 1 of the selector MUX5; the input terminal 2 of the AND gate AND1 is connected to the disable signal Disable0, the output terminal of the OR gate OR1 outputs the disable signal Disable0, and the two input terminals of the OR gate OR1 are respectively connected to the AND gate AND4 and the AND gate The output end of AND7; the second input end of AND gate AND7 is connected to the output end of selector MUX5, the second input end of selector MUX5 is connected to the second input end of selector MUX6, the first input end of selector MUX4 and the first input end of selector MUX2; the second input end of selector MUX3 is connected to the output end of AND gate AND5, the first input end of selector MUX1, the first input end of selector MUX6 and the second input end of selector MUX4, and the output end of AND gate AND5 outputs the authorization signal Grant1; the two input ends of OR gate OR2 are respectively connected to the output ends of AND gate AND8 and AND gate AND2, and the output end of OR gate OR2 outputs the disable signal Signal Disable1 is connected to the second input terminal of AND gate AND5; the second input terminal of AND gate AND8 is connected to the output terminal of selector MUX6; the second input terminal of AND gate AND2 is connected to the output terminal of selector MUX1; the second input terminal of AND gate AND6 is connected to the output terminal of selector MUX4; the two input terminals of OR gate OR3 are respectively connected to the output terminals of AND gate AND6 and AND gate AND3, the output terminal of OR gate OR3 outputs disable signal Disable2, and is connected to the second input terminal of AND gate AND9, and the output terminal of AND gate AND9 outputs authorization signal Grant2; the second input terminal of AND gate AND3 is connected to the output terminal of selector MUX2.
[0013] Preferably, in the 3*3 matrix arbiter, the value at the matrix Wij position, if it is 1, represents that the priority of i is greater than j; if it is 0, represents that the priority of i is less than j; after each successful arbitration, the value of the matrix Wij is updated, and the row where the request for successful arbitration is located is set to 0, and the column where it is located is set to 1, and the arbitration function is realized by maintaining the value of the matrix Wij.
[0014] Preferably, the priority of the arbitration requests is: request signal Request0 > request signal Request1 > request signal Request2.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] This invention discloses a data cache and transport circuit within a flash array-based integrated memory and computing chip for various neural network models. This circuit processes data streams from different neural network layers before they flow into the flash memory and computing array, including data caching and data transport. Data is sliced through input channels and transported to the flash memory and computing array based on the number of data channels and pixel size of the current neural network layer. A hybrid dimensional mapping method (i.e., address mapping strategy) for uneven Xbars is proposed. This hybrid dimensional mapping mechanism first determines the relationship between the number of input channels and Xbar word lines (WLs). When the number of input channels is less than the Xbar WL, a channel-by-channel mapping mechanism is employed. When the number of input channels is greater than the number of Xbar columns, the input channels are divided into multiple segments equal to the number of Xbar columns, and the weights corresponding to each segment are mapped channel-by-channel. Slicing the input channels decouples the input and output channels on the same Xbar, thus avoiding data splicing. Avoiding data splicing reduces the amount of data received by each tile by half, thus reducing local communication pressure and hotspots. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a circuit block diagram of data caching and transport provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the input data storage format provided by the present invention.
[0019] Figure 3 This is a schematic diagram of the cache management strategy provided by the present invention.
[0020] Figure 4 This is a circuit diagram of a 3*3 matrix arbiter provided by the present invention.
[0021] Figure 5 It is a schematic diagram of pixel coordinates (x, y, c) provided by the present invention. DETAILED DESCRIPTION
[0022] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used to facilitate and clearly illustrate the embodiments of the present invention.
[0023] like Figure 1 As shown, an embodiment of the present invention specifically provides a data cache and transfer circuit within a flash array-based integrated storage and computing chip. This circuit completes the data processing process for data streams at different layers of a neural network before they flow into the flash storage and computing array, including data caching and data transfer operations. Data is sliced through input channels and transferred to the flash storage and computing array based on the number of data channels and pixel size of the current neural network layer. It mainly consists of three parts: a dynamic cache unit, an arbitration and crossbar switch unit, and an address mapping unit.
[0024] The dynamic buffer unit is responsible for receiving the cast data stream from the on-chip network and caching the data in the multi-bank dynamic buffer. When the data in the buffer is sufficient to complete a convolution, the data in the buffer will be input to the flash storage array. In this process, the feature map pixels that need to be released due to the completion of the convolution are continuously released, and the buffer requirements are reduced by dynamically managing the buffer. The input data storage format is as follows: Figure 2 As shown in the example, since the data is bound to 16 channels during transmission, in order to ensure symmetry, the buffer width is also set to 16 channels * quantization bit width. Figure 2 In the example, each color represents 16 channels. In actual storage, feature maps with less than 16 channels or channels that are not an integer multiple of 16 should be supplemented to 16 channels to align them in storage. Here, a 3*3 convolution window with a stride of 1 is used as an example to illustrate the buffer management strategy: Figure 3 The pixels marked with "X" are the pixels that have been released when the black convolution window is executed, the pixels marked with "0" and "√" are the pixels that must be kept in the cache when the current convolution is executed, the pixels marked with "." are the pixels that will be released after the current convolution is executed, and the pixels marked with "-" are the pixels that may exist in the cache when the current convolution is executed. The multi-bank buffer can be regarded as a large cache. The lower bit in the merged address line is used to select the corresponding bank (storage body), and the higher bit indicates the corresponding row address. Combined with the above cache management strategy, it can be seen that the storage format of a frame feature map in the cache changes with time as shown below. Figure 3The gray area is the location where valid data is stored in the buffer. In this way, the buffer capacity requirement will be much smaller than the buffer required to cache the entire frame image.
[0025] The arbitration and crossbar unit (crossbar) consists of an arbitrator and two crossbar switches. The arbitrator circuit here uses a matrix arbitrator. Figure 4 A 3*3 matrix arbiter is used as an example to illustrate its functionality. A value of 1 at position Wij indicates that i has a higher priority than j, while a value of 0 indicates that i has a lower priority than j. After each successful arbitration, the W matrix is updated, with the row containing the successfully arbitrated request set to 0 and the column containing the successfully arbitrated request set to 1. This maintains the W matrix to implement arbitration. The arbitration request is a one-hot code generated by decoding the lower two bits of all read pointers. The resulting arbitration success signal serves as the crossbar select signal. The crossbar circuit structure consists of multiple selector muxes. In this circuit, there are two crossbars: one for address transmission and the other for data transmission. Due to the register-out nature of RAM, the address and data are transmitted in adjacent cycles. Therefore, the address crossbar select signal (sel) is derived directly from the arbitrator output, while the data crossbar select signal (sel) is derived from the arbitrator output after a tap. This circuit structure is suitable because the address mapping strategy in the subsequent address mapping unit ensures that the winning data is transmitted to the designated location without blocking.
[0026] The arbiter is a 3*3 matrix arbiter, which specifically includes: AND gates AND1~AND9, OR gates OR1~OR3, selectors MUX1~MUX6, request signals Request0~Request2, disable signals Disable0~Disable2 and grant signals Grant0~Grant2; wherein the request signal Request0 is connected to input terminal 1 of AND gate AND1, input terminal 1 of AND gate AND2 and input terminal 1 of AND gate AND3; the request signal Request1 is connected to input terminal 1 of AND gate AND4, input terminal 1 of AND gate AND5 and input terminal 1 of AND gate AND6. Input terminal 1 of AND gate AND6; request signal Request2 is connected to input terminal 1 of AND gate AND7, input terminal 1 of AND gate AND8, and input terminal 1 of AND gate AND9; the output of AND gate AND1 outputs authorization signal Grant0, and is simultaneously connected to input terminal 2 of selector MUX1, input terminal 2 of selector MUX2, input terminal 1 of selector MUX3, and input terminal 1 of selector MUX5; input terminal 2 of AND gate AND1 is connected to disable signal Disable0, the output terminal of OR gate OR1 outputs disable signal Disable0, and the two input terminals of OR gate OR1 are respectively connected to AND gate AND4 and AND gate AN The output end of D7; the second input end of the AND gate AND7 is connected to the output end of the selector MUX5, the second input end of the selector MUX5 is connected to the second input end of the selector MUX6, the first input end of the selector MUX4 and the first input end of the selector MUX2; the second input end of the selector MUX3 is connected to the output end of the AND gate AND5, the first input end of the selector MUX1, the first input end of the selector MUX6 and the second input end of the selector MUX4, and the output end of the AND gate AND5 outputs the authorization signal Grant1; the two input ends of the OR gate OR2 are respectively connected to the output ends of the AND gate AND8 and the AND gate AND2, and the output end of the OR gate OR2 outputs the disable signal Signal Disable1 is connected to the second input terminal of AND gate AND5; the second input terminal of AND gate AND8 is connected to the output terminal of selector MUX6; the second input terminal of AND gate AND2 is connected to the output terminal of selector MUX1; the second input terminal of AND gate AND6 is connected to the output terminal of selector MUX4; the two input terminals of OR gate OR3 are respectively connected to the output terminals of AND gate AND6 and AND gate AND3, the output terminal of OR gate OR3 outputs the disable signal Disable2, and is connected to the second input terminal of AND gate AND9, and the output terminal of AND gate AND9 outputs the authorization signal Grant2; the second input terminal of AND gate AND3 is connected to the output terminal of selector MUX2.
[0027] The working principle of this matrix arbitration circuit is: when the three input request signals Request0~Request2 request access to the address or data at the same time, one requester is selected to obtain the access permission, that is, the corresponding Grant0~Grant2, to avoid conflicts and realize priority judgment by using combinational logic circuits.
[0028] Priority order: Request0>Request1>Request2 (i.e., requester 0 has the highest priority).
[0029] Logical expressions:
[0030] Grant0 = Request0 (as long as Request0 is high, grant immediately and ignore other requests).
[0031] Grant1 = !Request0&Request1 (Request1 can only be authorized when Request0 is invalid).
[0032] Grant2 = !Request0&!Request1&Request2 (Request2 is granted only when both Request0 and Request1 are invalid).
[0033] The address mapping unit mainly realizes the mapping between feature map pixels and addresses in the cache. Figure 5 As shown in the figure, an image with H rows, W columns, and C channels (16 channels) is used as an example to illustrate how to find the location of a pixel in the cache address when the coordinates of that pixel (x, y, c) are known. Data enters the cache in the order of channel-row-column. Therefore, before (x, y, 0), y-1 rows of data have already been input, and x-1 channels of data have already been input in row y. For (x, y, c), there are C-1 additional inputs at the same location but different channels. Because the cache has frame periodicity, the physical address of a pixel in the cache is:
[0034]
[0035] According to the formula, buffer_size, W, and C are fixed, while x and y change as the sliding window moves. Since the previous buffer management strategy already calculates the coordinates of the convolution window's lower right corner, and the convolution window size is fixed, the position of the input corresponding to a specific BL line in the flash memory within the convolution window is also fixed after mapping a network. Therefore, we can calculate the address of the required BL input to the flash memory one by one. Since many consecutive pixels in the convolution window are also consecutive in the cache, these characteristics can be used to compress configuration information.
[0036] The configuration information only needs to include X_pos, Y_pos, and start_ch, which represent the relative position of the BL input within the convolution window. ch_num specifies the number of consecutive pixels following this pixel that can be sequentially transferred to the adjacent BL. To increase data parallelism, address mapping units equal to the number of banks are used to fetch data from multiple banks in parallel, maximizing the advantages of multiple banks.
[0037] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention. Any changes and modifications made by ordinary technicians in the field of the present invention based on the above disclosure shall fall within the scope of protection of the claims.
Claims
1. A data cache transport circuit inside a memory-computing integrated chip, used to process data streams from different layers of a neural network before they flow into a flash memory-computing array, characterized in that: include: The dynamic cache unit is responsible for receiving the cast data stream from the on-chip network and caching the data in the multi-bank dynamic cache. When the data in the cache can be used for a convolution, the data in the cache will be input into the flash storage array. During the convolution process, the feature map pixels required to complete the convolution process will be continuously released to reduce the cache demand by dynamically managing the cache. The arbitration and crossbar switch unit consists of an arbitrator and two crossbar switches. One crossbar switch is responsible for transmitting the address in the address mapping unit to the dynamic cache unit, and the selection signal of this crossbar switch comes directly from the output of the arbitrator. The other crossbar switch is responsible for transmitting the data in the dynamic cache unit to the address mapping unit, and the selection signal of this crossbar switch comes from the result of a beat of the arbitrator output. The address mapping unit slices the input channels to decouple the input and output channels on the same crossbar switch. The crossbar switch receives data from the arbitrator, maps the feature map pixels to the addresses in the cache, and outputs the data to the data buffer. The address mapping unit also includes a preset address mapping strategy, specifically including: first determining the relationship between the number of input channels and the word lines WL in the crossbar switch; when the number of input channels is less than the word lines WL in the crossbar switch, adopting a channel-by-channel mapping mechanism; when the number of input channels is greater than the number of crossbar switch columns, dividing the input channels into multiple segments equal to the number of crossbar switch columns, and applying a channel-by-channel mapping mechanism to the weights corresponding to each segment.
2. The data cache transfer circuit inside a memory-computing integrated chip according to claim 1, characterized in that: The width of the dynamic cache unit is 16 channels * quantization bit width, and during the data storage process, when the number of channels of the feature map is less than 16 channels or is not an integer multiple of 16, the feature map should be supplemented so that the number of channels of the feature map reaches an integer multiple of 16 to ensure storage alignment.
3. The internal data cache transfer circuit of a memory-computing integrated chip according to claim 2, characterized in that: The dynamic cache unit also includes a preset cache management strategy, specifically including: in the merged address line, the low bit is used to select the corresponding bank, and the high bit is used to indicate the corresponding row address; the order in which data enters the cache is in the order of channel-row-column.
4. The data cache transfer circuit inside a memory-computing integrated chip according to claim 1, characterized in that: The arbiter is a 3*3 matrix arbiter, which specifically includes: AND gates AND1~AND9, OR gates OR1~OR3, selectors MUX1~MUX6, request signals Request0~Request2, disable signals Disable0~Disable2 and grant signals Grant0~Grant2; wherein the request signal Request0 is connected to input terminal 1 of AND gate AND1, input terminal 1 of AND gate AND2 and input terminal 1 of AND gate AND3; the request signal Request1 is connected to input terminal 1 of AND gate AND4, input terminal 1 of AND gate AND5 and input terminal 1 of AND gate AND6. Input terminal 1 of AND gate AND6; request signal Request2 is connected to input terminal 1 of AND gate AND7, input terminal 1 of AND gate AND8, and input terminal 1 of AND gate AND9; the output of AND gate AND1 outputs authorization signal Grant0, and is simultaneously connected to input terminal 2 of selector MUX1, input terminal 2 of selector MUX2, input terminal 1 of selector MUX3, and input terminal 1 of selector MUX5; input terminal 2 of AND gate AND1 is connected to disable signal Disable0, the output terminal of OR gate OR1 outputs disable signal Disable0, and the two input terminals of OR gate OR1 are respectively connected to AND gate AND4 and AND gate AN The output end of D7; the second input end of the AND gate AND7 is connected to the output end of the selector MUX5, the second input end of the selector MUX5 is connected to the second input end of the selector MUX6, the first input end of the selector MUX4 and the first input end of the selector MUX2; the second input end of the selector MUX3 is connected to the output end of the AND gate AND5, the first input end of the selector MUX1, the first input end of the selector MUX6 and the second input end of the selector MUX4, and the output end of the AND gate AND5 outputs the authorization signal Grant1; the two input ends of the OR gate OR2 are respectively connected to the output ends of the AND gate AND8 and the AND gate AND2, and the output end of the OR gate OR2 outputs the disable signal Signal Disable1 is connected to the second input terminal of AND gate AND5; the second input terminal of AND gate AND8 is connected to the output terminal of selector MUX6; the second input terminal of AND gate AND2 is connected to the output terminal of selector MUX1; the second input terminal of AND gate AND6 is connected to the output terminal of selector MUX4; the two input terminals of OR gate OR3 are respectively connected to the output terminals of AND gate AND6 and AND gate AND3, the output terminal of OR gate OR3 outputs the disable signal Disable2, and is connected to the second input terminal of AND gate AND9, and the output terminal of AND gate AND9 outputs the authorization signal Grant2; the second input terminal of AND gate AND3 is connected to the output terminal of selector MUX2.
5. The data cache transfer circuit inside a memory-computing integrated chip according to claim 4, characterized in that: In the 3*3 matrix arbiter, the value at the matrix Wij position, if it is 1, it means that the priority of i is greater than j; if it is 0, it means that the priority of i is less than j; after each successful arbitration, the value of the matrix Wij is updated, and the row where the request for successful arbitration is located is set to 0, and the column where it is located is set to 1. The arbitration function is realized by maintaining the value of the matrix Wij.
6. The internal data cache transfer circuit of a memory-computing integrated chip according to claim 5, characterized in that: The priority of the arbitration requests is: request signal Request0 > request signal Request1 > request signal Request2.
Citation Information
Patent Citations
High-order router row buffer optimization structure
CN108111438A
Multicast implementation method and device based on Tile exchange architecture
CN117714392A