Data transmission method, data transmission apparatus and electronic device
Through a new data transmission method, the buffer unit and the specified storage sequence are used to solve the problem of low data transmission efficiency in matrix operations, realizing real-time data re-arrangement during the reading process, thereby improving computing efficiency.
Patent Information
- Application Number
- PCT/CN2024/097168
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-06-04
- Publication Date
- 2025-06-19
AI Technical Summary
In matrix operation, because the reusability of data requires multiple readings and rearrangements during the operation, the data transmission efficiency is inefficient and the calculation time is increased.
A data transmission method is provided, by sequentially reading storage unit data at multiple depths and storing them into a buffer unit, writing data to the second memory in a designated order, ensuring that adjacent sub-data units are stored in the same storage unit.
The data is rearranged at the same time during the reading of data, so that the rearrangement does not take up extra time and improves the efficiency of data reading.
Smart Images

Figure CN2024097168_19062025_PF_FP_ABST
Abstract
Description
Data transmission method, data transmission device and electronic device
[0001] This application claims priority to Chinese patent application No. 202311700987.3 filed on December 12, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as a part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to a data transmission method, a data transmission device, and an electronic device. Background Art
[0003] In matrix operations, due to the reusability of the data of two input matrices (for example, matrix A and matrix B), the data is generally first read from an external memory such as double-data-rate synchronous dynamic random-access memory (DDR) into a local data memory (LSM, or simply referred to as "local memory") during the operation process. Then, the data is read from the local data memory into a general register and sent to the matrix operation unit for operation. After the operation is completed, the result is written back to the external memory DDR.
[0004] Summary of the Invention
[0005] At least one embodiment of the present disclosure provides a data transmission method for transmitting data from a first memory to a second memory, wherein the first memory includes a memory bank with a depth of N, and a first storage unit of the memory bank at each depth has a width of K bits, and the second memory includes a plurality of consecutively arranged second storage units, and the width of the second storage unit is K bits. The method comprises: sequentially reading data stored in the first storage unit at each depth of the N depths, and sequentially storing the data of the N depths into a buffer unit, wherein the data is m bits, and each first storage unit is divided into K / m sub-data units according to m bits from low to high bits, wherein N, K and m are all positive integers; and writing the data in the memory bank from the buffer unit to the second memory in a specified order, wherein the specified order comprises: K / m data at the same position in adjacent K / m first storage units in the memory bank are stored in the same second storage unit of the second memory.
[0006] For example, in a method provided in an embodiment of the present disclosure, data stored in a first storage unit with a low depth among K / m adjacent first storage units in a memory bank is located in a low bit of the second storage unit.
[0007] For example, in a method provided in an embodiment of the present disclosure, the data stored in the first storage unit at each of the N depths is read in sequence, including: reading the data stored in the first storage unit at each of the N depths in sequence according to the depth order, and reading the data in the K / m sub-data units stored in the same first storage unit in parallel.
[0008] For example, in a method provided in an embodiment of the present disclosure, the buffer unit includes a plurality of buffers, each buffer is used to store one of the m-bit data, and the number of the plurality of buffers is greater than or equal to K / m.
[0009] For example, in the method provided in one embodiment of the present disclosure, the buffer unit includes (K / m) 2 A continuously arranged buffer zone is provided, and data in the storage body is written from the buffer unit to the second memory in a specified order, including: after sequentially storing data in first storage units of K / m depths into the buffer unit, starting to read data in K / m sub-data units at the same position in adjacent K / m first storage units in the storage body from the buffer unit.
[0010] For example, in a method provided in an embodiment of the present disclosure, after the data in the first storage units of K / m depths are sequentially stored in the buffer unit, data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body are started to be read from the buffer unit, including: after the data in the first storage units of K / m depths are sequentially stored in the buffer unit, the data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body are read in parallel; and the data in the K / m sub-data units at the same position in the adjacent K / m first storage units are written into the same second storage unit of the second memory.
[0011] For example, in a method provided in an embodiment of the present disclosure, multiple sub-data units in the storage body are continuously indexed in order of depth from small to large and position from small to large. After the data in the first storage units of K / m depths are sequentially stored in the buffer unit, the data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body are read in parallel, including: from the first cycle to the K / m-th cycle, in each cycle, the data in the sub-data units with index numbers (i-1)×K / m, ...K / m×i-1 are read from the storage body and stored in the buffer unit, where i represents the cycle count value; and starting from the K / m+1-th cycle, for each cycle, while reading out the data in the sub-data units with index numbers (i-1)×K / m, ...K / m×i-1 and storing them in the buffer unit, the K / m data at the same position in the adjacent K / m first storage units in the storage body are read in parallel from the buffer unit.
[0012] At least one embodiment of the present disclosure provides a data transmission device for transmitting data from a first memory to a second memory, wherein the first memory includes a memory bank with a depth of N, and a first storage unit of the memory bank at each depth has a width of K bits, and the second memory includes a plurality of consecutively arranged second storage units, and a width of the second storage unit has K bits, and the device includes: a buffer unit; a reading unit configured to sequentially read data stored in the first storage unit at each depth of N depths, and sequentially store the data of the N depths into the buffer unit, wherein the data is m bits, and each first storage unit is divided into K / m sub-data units according to m bits from low to high bits; and a writing unit configured to write the data in the memory bank from the buffer unit to the second memory in a specified order, wherein the specified order includes: K / m data at the same position in adjacent K / m first storage units in the memory bank are stored in the same second storage unit of the second memory, wherein N, K and m are all positive integers.
[0013] For example, in the device provided in one embodiment of the present disclosure, the buffer unit includes: a plurality of registers, each register is configured to store the data read from the first storage unit, the reading unit includes: a plurality of first multiplexers, the plurality of first multiplexers are configured to read the data stored in the first storage unit at each depth of the N depths of each of the storage bodies, and sequentially store the data stored in the first storage unit at each depth into K / m registers in the plurality of registers, the writing unit includes: a plurality of second multiplexers, the plurality of second multiplexers are configured to select K / m data located at the same position in adjacent K / m first storage units in the storage body from the plurality of registers, and write them into the same second storage unit of the second memory.
[0014] For example, in the device provided in one embodiment of the present disclosure, the first memory includes P storage bodies, the number of the buffer units is P, each buffer unit includes a preset number of registers, and the P storage bodies and the P buffer units correspond one to one, where P is a positive integer.
[0015] For example, in the device provided in one embodiment of the present disclosure, the preset number is a positive integer greater than or equal to K / m.
[0016] At least one embodiment of the present disclosure provides an electronic device, comprising the data transmission device provided in any embodiment of the present disclosure, a first memory, and a second memory, wherein the first memory is coupled to the data transmission device, and the data transmission device is coupled to the second memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0018] FIG1A is a schematic diagram showing a dot product of rows of matrix A and columns of matrix B in a matrix multiplication operation;
[0019] FIG1B shows a schematic diagram of a matrix operation data flow;
[0020] FIG1C shows a schematic diagram of a matrix data arrangement in an external memory and a local memory in a column direction;
[0021] FIG1D is a schematic diagram showing a matrix data in a storage bank in a local memory arranged in a column direction;
[0022] FIG2A shows a flow chart of a data transmission method provided by at least one embodiment of the present disclosure;
[0023] FIG2B shows a schematic system architecture diagram for executing a data transmission method provided by at least one embodiment of the present disclosure;
[0024] 3A to 3G are schematic diagrams of a method for rearranging data in a 16-bit data format, provided by at least one embodiment of the present disclosure;
[0025] 4A to 4M are schematic diagrams of a method for rearranging data in an 8-bit format, provided by at least one embodiment of the present disclosure;
[0026] FIG5 shows a schematic block diagram of a data transmission device provided by at least one embodiment of the present disclosure;
[0027] FIG6A shows a schematic structural diagram of a data transmission device in FIG5 provided by at least one embodiment of the present disclosure;
[0028] FIG6B shows a schematic structural diagram of another data transmission device provided by at least one embodiment of the present disclosure; and
[0029] FIG7 shows a schematic diagram of an electronic device provided by at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0031] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0032] FIG1A is a schematic diagram showing a dot product of rows of a matrix A and columns of a matrix B in a matrix multiplication operation.
[0033] The data required for matrix operations is generally arranged linearly along the rows or columns of the matrix in the external memory. As shown in Figure 1A, in the matrix multiplication operation C = C + A * B, it is necessary to perform a dot product between each row of matrix A and each column of matrix B, and then update the corresponding elements in matrix C. Therefore, if the data in matrix A is arranged linearly along the columns in the external memory, the data in matrix A needs to be rearranged so that the dot product of each row of matrix A and each column of matrix B is achieved.
[0034] FIG1B shows a schematic diagram of a matrix operation data flow.
[0035] As shown in Figure 1B, for example, the external memory DDR stores matrix A, which is read from the external memory DDR into the local data memory, and then read from the local data memory into the read and data reordering unit. The read and data reordering unit reorganizes and reorders the matrix, for example, reorganizing the data organized in the column direction in the row direction. The reorganized and reordered data is sent to the vector general register (also called "vector register") array, and then sent to the matrix operation unit for operation. After the operation is completed, the result is written back to the external memory DDR. The local data memory is usually a static random access memory (SRAM) array composed of multiple memory banks, which can support a larger read and write data bit width.
[0036] FIG1C is a schematic diagram showing that matrix data in an external memory and a local memory are arranged in a column direction; FIG1D is a schematic diagram showing that matrix data in a storage bank in a local memory are arranged in a column direction.
[0037] As shown in Figures 1C and 1D, the local data memory generally includes multiple banks, such as bank 0, bank 1, bank 2, and bank 3, and the multiple banks are independent of each other. Each bank includes multiple storage cells at different depths, and the bit width of each storage cell can be 32 bits, 16 bits, or 8 bits, etc. In the embodiment of the present disclosure, the bit width of each storage bank is 32 bits for illustration. For example, if the storage cell bit width is 32 bits, each storage cell can store two 16-bit data or matrix elements. In the embodiment of the present disclosure, the depth of the storage cell increases from right to left, and the depth of the rightmost storage cell in Figure 1D is 0. For example, the lower 16 bits P0 of the memory cell with a depth of 0 in memory bank 0 store element 0 in matrix A, and the upper 16 bits P1 store element 1 in matrix A; the lower 16 bits P2 of the memory cell with a depth of 1 in memory bank 0 store element 8 in matrix A, and the upper 16 bits P3 store element 9 in matrix A; the lower 16 bits P4 of the memory cell with a depth of 2 in memory bank 0 store element 16 in matrix A, and the upper 16 bits P5 store element 17 in matrix A; the lower 16 bits P6 of the memory cell with a depth of 3 in memory bank 0 store element 24 in matrix A, and the upper 16 bits P7 store element 25 in matrix A. The other memory banks of the local memory, such as memory bank 1, memory bank 2, and memory bank 3, are similar to memory bank 0 and will not be described in detail. Referring to Figures 1C and 1D, for example, element 2 in matrix A is stored in the lower 16 bits of the memory cell with a depth of 0 in memory bank 1, and the storage method of other elements is similar.
[0038] Similarly, the data format of the elements in matrix A can also be 8 bits. In this case, each storage cell can store 4 elements. For example, element 0 in matrix A is stored in the lower 8 bits P0 of the storage cell at depth 0 in memory bank 0. Element 1 is stored in bits P1 (8th to 15th bits) of the storage cell at depth 0 in memory bank 0. Element 2 is stored in bits P2 (16th to 23rd bits) of the storage cell at depth 0 in memory bank 0. Element 3 is stored in bits P3 (24th to 31st bits) of the storage cell at depth 0 in memory bank 0. For example, the lower 8 bits P4 of the storage cell at depth 1 in memory bank 0 store element 16 of matrix A. Similarly, bits P5 to P7 of the storage cell store elements 17 to 19 of matrix A, respectively. Similarly, bits P8 to P11 of the storage cell at depth 2 in memory bank 0 store the four elements of matrix A, respectively. Bits P12 to P15 of the storage cell at depth 3 in memory bank 0 store the four elements of matrix A, respectively.
[0039] As shown in Figures 1C and 1D, when the data of matrix A is stored in the external memory DDR in a column-oriented manner, after being read from the external memory DDR into the local data memory, the local data memory is still stored in a column-oriented manner. For example, when the data format of each element in matrix A is 16 bits, elements 0 to 15 occupy the storage space of two columns of the local memory; when the data format of each element in matrix A is 8 bits, elements 0 to 15 occupy the storage space of one column of the local memory. Similarly, when the data format of each element in matrix A is 4 bits, elements 0 to 15 occupy the storage space of 1 / 2 columns of the local memory. This requires reorganizing the data so that it can be read out in a row-oriented format and then sent to the operation unit to facilitate matrix multiplication operations.
[0040] As shown in Figures 1B to 1D, matrix operations require that the matrix data be read out from the local memory and stored in the reading and data rearrangement unit first, and then the reading and data rearrangement unit rearranges the matrix data. Rearranging the matrix data takes additional time, and the matrix data reading efficiency is low.
[0041] It should be noted that although the embodiments of the present disclosure use matrices as an example to illustrate that data rearrangement takes additional time, this does not mean that the present invention is only applied to rearrangement of matrix data. The embodiments of the present disclosure can be applied to any data that needs to be rearranged, not limited to matrix data.
[0042] At least one embodiment of the present disclosure provides a data transmission method, a data transmission device, and an electronic device. The data transmission method is used to transmit data from a first memory to a second memory, wherein the first memory includes a memory bank of N depths, wherein a first storage unit at each depth of the memory bank has a width of K bits, and the second memory includes a plurality of consecutively arranged second storage units, wherein the second storage units have a width of K bits. The method comprises: sequentially reading data stored in the first storage units at each of the N depths, and sequentially storing the data at the N depths into a buffer unit, wherein the data is m bits, and each first storage unit is divided into K / m sub-data units from the lowest bit to the highest bit according to the m-bit structure; and writing the data in the memory bank from the buffer unit to the second memory in a specified order, wherein the specified order includes: K / m data at the same position in adjacent K / m first storage units in the memory bank are stored in the same second storage unit of the second memory, where N, K, and m are all positive integers. The data transmission method can simultaneously rearrange the data during the data reading process, so that the data rearrangement step does not take up additional time, thereby improving data reading efficiency.
[0043] FIG2A shows a flowchart of a data transmission method provided by at least one embodiment of the present disclosure.
[0044] As shown in FIG2A , the method may include steps S10 to S20. The data transmission method is used to transmit data from a first memory to a second memory, wherein the first memory includes a memory bank having a depth of N, and a first memory cell at each depth of the memory bank has a width of K bits, and the second memory includes a plurality of consecutively arranged second memory cells, and the width of the second memory cells is K bits.
[0045] Step S10: Read the data stored in the first storage unit at each depth of N depths in turn, and store the data of N depths in the buffer unit in turn. The data is m bits, and each first storage unit is divided into K / m sub-data units according to m bits from low to high.
[0046] Step S20: Write the data in the storage body from the buffer unit to the second memory in a specified order, wherein the specified order includes K / m data at the same position in adjacent K / m first storage units in the storage body being stored in the same second storage unit of the second memory.
[0047] FIG2B shows a schematic system architecture diagram for executing a data transmission method provided by at least one embodiment of the present disclosure.
[0048] As shown in FIG2B , the system includes a memory 201, a buffer unit 202, and a memory 203. Memory 201 is an example of a first memory, and memory 203 is an example of a second memory. Memory 201 is, for example, the local data memory shown in FIG1C and FIG1D , and memory 203 is, for example, the vector general register shown in FIG1B .
[0049] Memory 201 includes at least one memory bank 211 (an example of a memory bank) having a depth of N. FIG2B shows only one memory bank 211 as an example. In practice, as shown in FIG1D and FIG1C , the local data memory may include multiple memory banks. Each depth is considered a memory cell, and N first memory cells are arranged, for example, along a row direction. For example, memory bank 211 includes memory cell B1, memory cell B2, memory cell B3, memory cell B4, ..., memory cell BN (hereinafter referred to as memory cells B1-BN), where memory cells B1-BN are examples of first memory cells, and memory cells B1-BN are arranged sequentially.
[0050] Each memory cell at the N depths has a width of K bits, that is, a memory cell has a bit width of K bits and can store K bits of data. For example, memory cells B1 to BN each include K bits.
[0051] In some embodiments, for example, if the data format stored in memory 201 is mbit, each storage unit is divided into K / m sub-data units from low to high bits according to m bits, and each sub-data unit is used to store m-bit data. For example, if the data in matrix A is stored in memory 201, and each element in matrix A is in mbit data format, then for example, storage unit B1 is divided into K / m sub-data units from low to high bits, and each sub-data unit is used to store one element in matrix A. Bits 0 to (m-1) are the first sub-data unit, used to store one element in matrix A, and bits m to (2m-1) are the second sub-data unit, used to store another element in matrix A, and so on.
[0052] For example, in the example of K=32 and m=16, each storage unit is divided into two sub-data units. For example, storage unit B1 is divided into sub-data units B10 and B11, where sub-data unit B10 stores element 0 in matrix A, and sub-data unit B11 stores element 1 in matrix A; storage unit B2 is divided into sub-data units B20 and B21, where sub-data unit B20 stores element 8 in matrix A, and sub-data unit B21 stores element 9 in matrix A; storage unit B3 is divided into sub-data units B30 and B31, storage unit B4 is divided into sub-data units B40 and B41, and storage unit BN is divided into sub-data units BN0 and BN1.
[0053] Similarly, the second memory 203 includes a plurality of consecutive second storage units, such as storage unit C1, storage unit C2, storage unit C3, storage unit C4, ..., storage unit Cx, each of which is configured to store K bits of data. If the data format stored in the second memory 203 is m bits, then one storage unit can also store K / m data. The size of the second memory 203 can be the same as that of the first memory 201 (e.g., x = N), or it can be different.
[0054] The buffer unit 202 may include multiple buffers, and the number of buffers may be determined according to the data format of the matrix and whether parallel reading is performed. In some embodiments of the present disclosure, the multiple buffer units correspond to the multiple memory banks in a one-to-one manner.
[0055] In some embodiments of the present disclosure, each buffer is used to store m bits of data, and the number of buffers corresponding to each memory bank is greater than or equal to K / m.
[0056] For example, based on the condition that a single memory bank is 32 bits wide, each memory unit is 32 bits, and each includes two 16-bit data formats. In order to realize the matrix row torque array or the matrix column to matrix row conversion, it is necessary to operate on at least two 16-bits of each column. Each buffer unit can include (32 / 16) = 2 buffers. If the first memory includes M (M is a positive integer) memory banks, then a total of 2*M buffers are required. Accordingly, for an 8-bit data format, each buffer unit can include (32 / 8) = 4 buffers. Accordingly, for a 4-bit data format, each buffer unit can include (32 / 4) = 8 buffers.
[0057] In other embodiments of the present disclosure, in order to achieve parallel processing, the number of buffers can be increased, and the number of buffers in each buffer unit is greater than K / m. For example, each buffer unit includes (K / m) 2 A continuous array of buffers.
[0058] For example, for a 16-bit data format, each buffer unit 202 may include 4 buffers, each buffer storing 16 bits; for an 8-bit data format, each buffer unit 202 may include 16 buffers, each buffer storing 8 bits; for a 4-bit data format, each buffer unit 202 may include 64 buffers, each buffer storing 4 bits.
[0059] For step S10 in FIG2A , for example, the data stored in N storage cells are read sequentially in ascending order of the serial numbers of the storage cells. In some embodiments of the present disclosure, the data stored in the first storage cell at each of the N depths are read sequentially in depth order, and the data in the K / m sub-data units stored in the same first storage cell are read in parallel. That is, the data in the K / m sub-data units stored in the first storage cell at the same depth are read in parallel. For example, the data in storage cells B1 to BN are read sequentially, and the two data in each storage cell are read out at the same time. For example, after the data in sub-data unit B10 and sub-data unit B11 are read out at the same time, the data in sub-data unit B20 and sub-data unit B21 are read out at the same time, and so on. Reading data at the same depth in parallel can improve reading efficiency.
[0060] For example, each cycle reads 32 bits of data and stores the 32 bits of data in the buffer unit. For example, each memory bank corresponds to four buffers. In the first cycle, the data of sub-data units B10 and B11 in memory unit B1 are read in parallel, and the data of sub-data units B10 and B11 are stored in buffers 212 and 222 of buffer unit 202, respectively. In the second cycle, the data of sub-data units B20 and B21 in memory unit B2 are read in parallel, and the data of sub-data units B20 and B21 are stored in buffers 232 and 242 of buffer unit 202, respectively.
[0061] In step S20 in FIG2A , K / m data are read from the buffer unit in a specified order and written into, for example, a vector general register. The specified order may be such that K / m data at the same position in K / m adjacent first storage cells in the memory bank are stored in the same second storage cell of the vector general register. That is, K / m data are read from the buffer unit in each cycle, and these K / m data are located in K / m adjacent first storage cells, and the positions of these K / m data in these K / m adjacent first storage cells are the same. For example, these K / m data are all from the sth position to the tth position in the first storage cell. Both s and t are positive integers.
[0062] For example, if K bit = 32 bits and m bit = 16 bits, then two data are read from the buffer unit in each cycle, and these two data are located in two adjacent first storage units and have the same position in the first storage units. For example, in one cycle, two data located in storage units B1 and B2 are read from the buffer unit, and these two data are located from the sth bit to the tth bit of storage units B1 and B2, respectively. For example, these two data are respectively in sub-data unit B10 (bits 0 to 15) and sub-data unit B20 (bits 0 to 15), or respectively in sub-data units B11 (bits 16 to 32) and B21 (bits 16 to 32).
[0063] For example, the data stored in the sub-storage unit B10 is a, and the data stored in the sub-storage unit B11 is b. The data a and b in the buffers 212 and 232 are read from the buffer unit 202 and stored in the second storage unit C1.
[0064] For example, the data stored in the first storage cell with the lowest depth among the K / m adjacent first storage cells in a memory bank is located in the lower bits of the second storage cell. A memory typically stores data in ascending order of depth, storing the data stored in the first storage cell with the lowest depth in the lower bits of the second storage cell, thereby maintaining the order of the K / m data.
[0065] For example, if the depth of sub-data unit B10 is lower than that of sub-data unit B20, data a in sub-data unit B10 is stored in the lower bits of storage unit C1, and data b in sub-data unit B20 is stored in the higher bits of storage unit C1.
[0066] In some embodiments of the present disclosure, the buffer unit includes (K / m) 2 Buffer units include (K / m) 2 A consecutive array of buffers can be read and written in parallel and can also save the number of buffers.
[0067] In some embodiments of the present disclosure, for example, after the data in the first storage units of K / m depth are sequentially stored in the buffer unit, the data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body are read from the buffer unit.
[0068] For example, in the example of K / m = 2, after the data in the first storage cells of two depths are sequentially stored in the buffer unit, the data in the two sub-data units located at the same position in the two adjacent first storage cells in the memory bank are read from the buffer unit. For example, as shown in Figure 2B, in the first cycle, the two data a and c in storage cell B1 are stored in buffer 212 and buffer 222 respectively. In the second cycle, the two data b and d in storage cell B2 are stored in buffer 232 and buffer 242 respectively. Then, after the second cycle, data a and data b are read from buffer 202.
[0069] In some embodiments of the present disclosure, after storing K / m depths of data in the buffer unit in sequence, data in K / m sub-data units at the same position in adjacent K / m first storage units in the storage body are read in parallel; and data in K / m sub-data units at the same position in adjacent K / m first storage units are written into the same second storage unit of the second memory.
[0070] For example, multiple sub-data units in a memory bank are sequentially indexed in order of depth and position, and after sequentially storing data in first memory units at a depth of K / m into a buffer unit, data in K / m sub-data units at the same position in adjacent K / m first memory units in the memory bank are read in parallel, including: from a first cycle to a K / mth cycle, data indexed (i-1)×K / m, ...K / m×i-1 are read from the memory bank in each cycle and stored in the buffer unit, where i represents a cycle count value; and starting from the K / m+1th cycle, for each cycle, while data indexed (i-1)×K / m, ...K / m×i-1 are read and stored in the buffer unit, K / m data at the same position in adjacent K / m first memory units in the memory bank are read from the buffer unit in parallel. This method is described below with reference to the embodiments of FIG. 3A to FIG. 3G .
[0071] Figures 3A to 3G are schematic diagrams of a method for rearranging 16-bit data according to at least one embodiment of the present disclosure. Figures 3A to 3G only show schematic diagrams of one memory bank, and similar operations are performed on each memory bank in the local data memory.
[0072] As shown in Figure 3A, multiple first storage units of a storage body in the local data memory are consecutively indexed in order of depth from small to large and position from small to large, for example, the index numbers are index 0, index 1, index 2, index 3, index 4, index 5, index 6 and index 7 respectively.
[0073] When the data format is mbit=16bit and Kbit=32bit, K / m=2.
[0074] In the first cycle (ie, i=1), data of index 0 and index 1 are read out from each memory bank and sent to the buffer.
[0075] In the second cycle (ie, i=2), the data of index 2 and index 3 are read out from each memory bank and sent to the buffer.
[0076] In the third cycle (i.e., i=3), the data of index 4 and index 5 are read from each memory bank and sent to the buffer. Simultaneously, the data of index 0 and index 2 are read from the buffer and sent to the vector general register array.
[0077] In the fourth cycle (i.e., i=4), the data of index 6 and index 7 are read from each memory bank and sent to the buffer. Simultaneously, the data of index 1 and index 3 are read from the buffer and sent to the vector general register array.
[0078] In the fifth cycle (i.e., i=5), the data of index 8 and index 9 are read from each memory bank and sent to the buffer. Simultaneously, the data of index 4 and index 6 are read from the buffer and sent to the vector general register array.
[0079] In the sixth cycle (i.e., i=6), the data at index 10 and index 11 are read from each memory bank and sent to the buffer. Simultaneously, the data at index 5 and index 7 are read from the buffer and sent to the vector general register array.
[0080] In cycle 7 (i.e., i=7), the data at indexes 12 and 13 are read from each memory bank and sent to the buffer. Simultaneously, the data at indexes 8 and 10 are read from the buffer and sent to the vector general register array. Aside from the differences in the indexes used for reading and writing, the hardware control logic remains the same as in cycle 3.
[0081] During the entire execution process, data is read from the local data memory in the "column direction" of the matrix, combined in the "row direction", and then written into the vector memory array.
[0082] Figures 4A to 4M are schematic diagrams of a method for rearranging data in an 8-bit format according to at least one embodiment of the present disclosure. Figures 4A to 4M only show schematic diagrams of one memory bank, and similar operations are performed on each memory bank in the local data memory.
[0083] When the data format is m=8 and K=32, K / m=4.
[0084] In the first cycle (ie, i=1), data of index 0, index 1, index 2, and index 3 are read from each memory bank and sent to the buffer.
[0085] In the second cycle (ie, i=2), data of index 4, index 5, index 6, and index 7 are read from each memory bank and sent to the buffer.
[0086] In the third cycle (ie, i=3), data at index 8, index 9, index 10, and index 11 are read from each memory bank and sent to the buffer.
[0087] In the fourth cycle (ie, i=4), data of index 12, index 13, index 14, and index 15 are read from each memory bank and sent to the buffer.
[0088] In the fifth cycle (i.e., i=5), the data at index 16, index 17, index 18, and index 19 are read from each memory bank and sent to the buffer. Simultaneously, the data at index 0, index 4, index 8, and index 12 are read from the buffer and sent to the vector general register array.
[0089] In the sixth cycle (i.e., i=6), the data at index 20, index 21, index 22, and index 23 are read from each memory bank and sent to the buffer. Simultaneously, the data at index 1, index 5, index 9, and index 13 are read from the buffer and sent to the vector general register array.
[0090] In the 7th cycle (i.e., i=7), the data at index 24, index 25, index 26, and index 27 are read from each memory bank and sent to the buffer. Simultaneously, the data at index 2, index 6, index 10, and index 14 are read from the buffer and sent to the vector general register array.
[0091] In the eighth cycle (i.e., i=8), the data at indexes 28, 29, 30, and 31 are read from each memory bank and sent to the buffer. Simultaneously, the data at indexes 3, 7, 11, and 15 are read from the buffer and sent to the vector general register array.
[0092] In the ninth cycle (i.e., i=9), the data at indexes 32, 33, 34, and 35 are read from each memory bank and sent to the buffer. Simultaneously, the data at indexes 16, 20, 24, and 28 are read from the buffer and sent to the vector general register array.
[0093] In the 10th cycle (i.e., i=10), the data at indexes 36, 37, 38, and 39 are read from each memory bank and sent to the buffer. Simultaneously, the data at indexes 17, 21, 25, and 29 are read from the buffer and sent to the vector general register array.
[0094] In the 11th cycle (i.e., i=11), the data at index 40, index 41, index 42, and index 43 are read from each memory bank and sent to the buffer. Simultaneously, the data at index 18, index 22, index 26, and index 30 are read from the buffer and sent to the vector general register array.
[0095] In the 12th cycle (i.e., i=12), the data at indexes 44, 45, 46, and 47 are read from each memory bank and sent to the buffer. Simultaneously, the data at indexes 19, 23, 27, and 31 are read from the buffer and sent to the vector general register array.
[0096] In cycle 13 (i.e., i = 13), data at indexes 48, 49, 50, and 51 are read from each bank and sent to the buffer. Simultaneously, data at indexes 32, 36, 40, and 44 are read from the buffer and sent to the vector general register array. Aside from the differences in the indexes used for reading and writing, the hardware control logic remains the same as in cycle 5.
[0097] During the entire execution process, data is read from the local data memory in the "column direction" of the matrix, combined in the "row direction", and then written into the vector memory array.
[0098] In other embodiments of the present disclosure, when the data format is m=4 and K=32, K / m=8, and from the first cycle to the eighth cycle, data indexed by (i-1)×K / m, ...K / m×i-1 are read from the memory bank in each cycle and stored in the buffer unit, where i represents the cycle count value; and starting from the ninth cycle, for each cycle, while data indexed by (i-1)×K / m, ...K / m×i-1 are read and stored in the buffer unit, 8 data at the same position in 8 adjacent first storage cells in the memory bank are read from the buffer unit in parallel. For m=4 bits and K=32, the implementations of m=8 and m=16 are similar and will not be repeated here.
[0099] Another aspect of the present disclosure provides a data transmission device for transmitting data from a first memory to a second memory, wherein the first memory includes a memory body with a depth of N, and the width of the first storage unit of the memory body at each depth is K bits; the second memory includes a plurality of consecutively arranged second storage units, and the width of the second storage unit is K bits.
[0100] FIG5 shows a schematic block diagram of a data transmission device 500 provided by at least one embodiment of the present disclosure.
[0101] As shown in FIG. 5 , the data transmission device 500 includes a reading unit 501 , a buffer unit 502 , and a writing unit 503 .
[0102] The reading unit 501 is configured to read the data stored in the first storage unit at each depth of N depths in turn, and store the data of N depths in the buffer unit 502 in turn. The data is m bits, and each first storage unit is divided into K / m sub-data units according to m bits from low to high.
[0103] The write unit 503 is configured to write the data in the storage body from the buffer unit to the second memory in a specified order, and the specified order includes: K / m data at the same position in adjacent K / m first storage units in the storage body are stored in the same second storage unit of the second memory.
[0104] In some embodiments of the present disclosure, the buffer unit 502 includes a plurality of registers, the read unit 501 includes a plurality of first multiplexers, and the write unit 503 includes a plurality of second multiplexers. Each of the plurality of registers is configured to store data read from a first storage unit. The plurality of first multiplexers is configured to read data stored in a first storage unit at each depth of N depths of each storage body, and sequentially store the data stored in the first storage unit at each depth into K / m registers in the plurality of registers. The plurality of second multiplexers is configured to select K / m data at the same position in adjacent K / m first storage units in the storage body from the plurality of registers, and write them into the same second storage unit of the second memory.
[0105] Each of the multiple registers acts as a buffer as described above.
[0106] FIG6A shows a schematic structural diagram of a data transmission device 500 in FIG5 provided by at least one embodiment of the present disclosure.
[0107] 6A , the reading unit 501 includes a plurality of multiplexers, such as a multiplexer 511, a multiplexer 521, a multiplexer 531, a multiplexer 541, a multiplexer 551, and a multiplexer 561. The multiplexers 511, 521, 531, 541, 551, and 561 are all examples of first multiplexers.
[0108] The buffer unit 502 includes a plurality of registers, including, for example, a register 512, a register 522, a register 532, and a register 542. Each register is a buffer.
[0109] The writing unit 503 includes a plurality of multiplexers, such as multiplexers 513, 523, 533, 543, 553, and 563. Multiple multiplexers such as multiplexers 513, 523, 533, 543, 553, and 563 are examples of second multiplexers.
[0110] In FIG6A , each square represents a register, and each trapezoid represents a multiplexer.
[0111] The data transmission device 500 in FIG6A illustrates a hardware structure for rearranging 16-bit data corresponding to a single memory bank. As described above, a single memory bank can correspond to four buffers, or registers, each storing a 16-bit piece of data. This allows the buffer unit to concurrently write and read 16-bit data.
[0112] For read unit 501, for example, multiplexers 511 and 521 are both one-to-many multiplexers, i.e., one input terminal and multiple output terminals. Multiplexers 531, 541, 551, and 561 can be many-to-one multiplexers, i.e., multiple input terminals and one output terminal.
[0113] For example, the input end of the multiplexer 511 is coupled to the sub-data units located in the first row in the storage body (for example, the sub-data units of index 0, index 2, index 4 and index 6), and the output end of the multiplexer 511 is coupled to the multiplexer 531, the multiplexer 541 and the multiplexer 551, thereby receiving the first data of the sub-data units located in the first row in the storage body and outputting the first data to one of the multiplexer 531, the multiplexer 541 and the multiplexer 551.
[0114] The input end of the multiplexer 521 is coupled to the sub-data units located in the second row of the storage body (for example, the sub-data units of index 1, index 3, index 5 and index 7), and the output end of the multiplexer 521 is coupled to the multiplexer 541, the multiplexer 551 and the multiplexer 561, thereby receiving the second data of the sub-data units located in the second row of the storage body and outputting the second data to one of the multiplexer 541, the multiplexer 551 and the multiplexer 561.
[0115] Multiplexer 531, multiplexer 541, multiplexer 551 and multiplexer 561 are coupled to register 512, register 522, register 532 and register 542 in a one-to-one correspondence, so that the data provided by multiplexer 531, multiplexer 541, multiplexer 551 and multiplexer 561 are written into the corresponding registers respectively.
[0116] For the write unit 503 , for example, multiplexers 513 , 523 , 533 , and 543 are one-to-many multiplexers; multiplexers 553 and 563 may be many-to-one multiplexers.
[0117] The input terminals of multiplexers 513, 523, 533, and 543 are coupled to registers 512, 522, 532, and 542 in a one-to-one correspondence. The output terminal of multiplexer 513 is coupled to the input terminal of multiplexer 553. The output terminal of multiplexer 523 is coupled to the input terminals of multiplexers 553 and 563, respectively. The output terminal of multiplexer 533 is coupled to the input terminals of multiplexers 553 and 563, respectively. The output terminal of multiplexer 543 is coupled to the input terminal of multiplexer 563. The output terminal of multiplexer 553 is coupled to the first row of sub-data units of the second memory, and the output terminal of multiplexer 563 is coupled to the second row of sub-data units of the second memory.
[0118] For example, in the first cycle, multiplexer 511 and multiplexer 521 receive data in the sub-data unit of index 0 and the sub-data unit of index 1 from the storage unit of depth 0 in the storage body of the first memory, respectively. For the sake of ease of description and understanding below, it is assumed that the data stored in index x is x, and x is a positive integer. In the first cycle, multiplexer 511 and multiplexer 521 receive data 0 in the sub-data unit of index 0 and data 1 in the sub-data unit of index 1 in the storage unit of depth 0. In addition, multiplexer 511 writes data 0 into register 512 via multiplexer 531; multiplexer 521 writes data 1 into register 522 via multiplexer 541.
[0119] In the second cycle, multiplexers 511 and 521 receive data 2 from the sub-data unit with index 2 and data 3 from the sub-data unit with index 3 in the memory unit with depth 1. Furthermore, multiplexer 511 writes data 2 into register 532 via multiplexer 551; multiplexer 521 writes data 3 into register 542 via multiplexer 561.
[0120] In the third cycle, multiplexer 513 receives data 0 from register 512 and provides data 0 to the lower 16 bits of the memory cell with a depth of 0 in the second memory via multiplexer 553. Multiplexer 533 receives data 2 from register 532 and provides data 2 to the upper 16 bits of the memory cell with a depth of 0 in the second memory via multiplexer 563. Simultaneously, multiplexers 511 and 521 receive data 4 from the sub-data unit with an index of 4 and data 5 from the sub-data unit with an index of 5 in the memory cell with a depth of 2. Multiplexer 511 writes data 4 to register 512 via multiplexer 531, and multiplexer 521 writes data 5 to register 532 via multiplexer 551.
[0121] In the fourth cycle, multiplexer 523 receives data 1 from register 522 and provides data 1 to the lower 16 bits of the storage unit with a depth of 1 in the second memory via multiplexer 553. Multiplexer 543 receives data 3 from register 542 and provides data 3 to the upper 16 bits of the storage unit with a depth of 1 in the second memory via multiplexer 563. At the same time, multiplexers 511 and 521 receive data 6 from the sub-data unit with an index of 6 and data 7 from the sub-data unit with an index of 7 in the storage unit with a depth of 3. Multiplexer 511 writes data 6 to register 522 via multiplexer 541, and multiplexer 521 writes data 7 to register 542 via multiplexer 561.
[0122] In each subsequent cycle, the multiplexer and the register cooperate to perform similar operations as above, so that the data transmission device 500 completes the steps of Figures 3A to 3G, which will not be repeated here.
[0123] FIG6B shows a schematic structural diagram of another data transmission device 600 provided by at least one embodiment of the present disclosure.
[0124] The data transmission device 600 in FIG6B illustrates an example of a hardware structure for rearranging 8-bit data corresponding to a single memory bank. As described above, a single memory bank can correspond to 16 buffers, or 16 registers, each storing an 8-bit piece of data. This allows the buffer unit to concurrently write and read 8-bit data.
[0125] In the example of FIG6B , the read unit 601 includes, for example, four one-to-many multiplexers 611 and sixteen many-to-one multiplexers 621; the buffer unit 602 includes, for example, sixteen registers 612; and the write unit 603 includes, for example, sixteen one-to-many multiplexers 613 and four many-to-one multiplexers 623. Similarly, in FIG6A , each square represents a register, and each trapezoid represents a multiplexer.
[0126] The operations performed by multiple multiplexers 611 and multiple multiplexers 621 in read unit 601 are similar to those of the multiple multiplexers in the read unit in FIG6A ; the operations performed by 16 registers 612 are similar to those of registers 512, 522, 532, and 542 in FIG6A ; and the operations performed by multiple multiplexers 613 and multiple multiplexers 623 in write unit 603 are similar to those of the multiple multiplexers included in write unit 503 in FIG6A . These multiplexers and registers work together to complete the steps described in FIG4A to FIG4M .
[0127] It should be noted that FIG6A and FIG6B are merely schematic hardware structure diagrams and have no limiting effect on the embodiments of the present disclosure. The hardware structure in actual applications may be more or less than the hardware structure shown in FIG6A and FIG6B .
[0128] Figures 6A and 6B only illustrate the embodiments of the present disclosure by taking the example of transmitting data in a storage body in the first memory to the second memory, but in actual applications, the first memory usually includes multiple storage bodies, and the implementation method described in Figure 6A or Figure 6B can be executed on each storage body separately.
[0129] FIG6A and FIG6B illustrate embodiments of the present disclosure using examples of rearrangement during 16-bit data transmission and rearrangement during 8-bit data transmission, respectively, but this does not limit the present disclosure. The embodiments of the present disclosure are also applicable to, for example, 4-bit data transmission and rearrangement, 32-bit data transmission and rearrangement, etc., and those skilled in the art can adaptively modify the embodiments provided in the present disclosure for use in 4-bit data transmission and rearrangement, 32-bit data transmission and rearrangement.
[0130] In some embodiments of the present disclosure, the first memory includes P memory banks, the number of buffer units is P, each buffer unit includes a preset number of registers, the P memory banks correspond to the P buffer units one-to-one, and P is a positive integer. The preset number is a positive integer greater than or equal to K / m. For example, each buffer unit includes K / m registers as K / m buffers, and P memory banks require P buffer units, which means P*(K / m) buffers are required. For example, each buffer unit may include (K / m) 2 registers.
[0131] The hardware structure in the embodiments of the present disclosure can simultaneously rearrange data during the data reading process, so that this step of rearranging data does not take up additional time, thereby improving the efficiency of data reading. In addition, the matrix rearrangement hardware structure of the present invention is applicable to local data memories with a configurable number of memory banks, meeting the needs of local data memories with different bit widths (bandwidths); and the hardware structure in the embodiments of the present disclosure is applicable to matrix operations in various data formats, including but not limited to 16-bit, 8-bit, and 4-bit.
[0132] In an embodiment of the present disclosure, the data transmission device hardware structure includes a hardware structure of P single memory banks, where the number P is consistent with the number of memory banks in the local data memory. The data transmission device input comes from the local data memory, and the data transmission device output is sent to the vector general register array.
[0133] The resources and connections within each single-bank hardware structure vary depending on whether the matrix data format is 16-bit, 8-bit, or 4-bit. If the data is 16-bit (as shown in Figure 6A), the input data is fed into 4*P 16-bit buffers at a specific cycle according to a specific connection relationship and then output from the buffers in a specific order. If the data is 8-bit (as shown in Figure 6B), the input data is fed into 16*P 8-bit buffers at a specific cycle according to a specific connection relationship and then output from the buffers in a specific order. Similarly, if the data is 4-bit, the input data is fed into 64*P 4-bit buffers at a specific cycle according to a specific connection relationship and then output from the buffers in a specific order.
[0134] It should be noted that in the embodiments of the present disclosure, the various units of the data transmission device correspond to the various steps of the aforementioned data transmission method. For the specific functions of the data transmission device, please refer to the relevant description of the data transmission method and will not be repeated here. The data transmission device 500 shown in Figure 5 and the components and structures of the data transmission device shown in Figure 6B are merely exemplary and non-limiting. The data transmission device may also include other components and structures as needed.
[0135] FIG7 shows a schematic block diagram of an electronic device 700 provided by at least one embodiment of the present disclosure.
[0136] For example, as shown in FIG. 7 , the electronic device 700 includes a data transmission device 710 , a first memory 720 , and a second memory 730 .
[0137] The data transmission device 710 is, for example, the data transmission device provided by any embodiment of the present disclosure. For example, the data transmission device 710 may be configured as shown in FIG6A or FIG6B. The first memory 720 is coupled to the data transmission device 710, and the data transmission device 710 is coupled to the second memory 730.
[0138] The first memory 720 is, for example, the local data memory shown in FIG. 1B , and the second memory 730 is, for example, the vector general register array shown in FIG. 1B .
[0139] The data transmission device 710 is used to transmit data from a first memory to a second memory, where the first memory includes a memory body with a depth of N, and the width of the first storage unit at each depth of the memory body is K bits. The second memory includes multiple consecutively arranged second storage units, and the width of the second storage unit is K bits. The method includes: reading the data stored in the first storage unit at each depth of N depths in sequence, and storing the data of N depths in a buffer unit in sequence, where the data is m bits, and each first storage unit is divided into K / m sub-data units from low to high bits according to m bits; and writing the data in the memory body from the buffer unit to the second memory in a specified order, and the specified order includes: K / m data at the same position in adjacent K / m first storage units in the memory body are stored in the same second storage unit of the second memory.
[0140] This electronic device can simultaneously rearrange data during the data reading process, eliminating the need for extra time to rearrange data, thereby improving data reading efficiency. Furthermore, the matrix rearrangement hardware structure of the present invention is applicable to local data memories with a configurable number of memory banks, meeting the needs of local data memories with different bit widths (bandwidths). Furthermore, the hardware structure of the disclosed embodiments is applicable to matrix operations on a variety of data formats, including but not limited to 16-bit, 8-bit, and 4-bit formats.
[0141] For example, in addition to being implemented in hardware, in other embodiments of the present disclosure, the data transmission device 710 may be implemented in hardware, software, firmware, or any feasible combination thereof. For example, the data transmission device 710 may be a dedicated or general-purpose circuit, chip, or device. The embodiments of the present disclosure do not limit the specific implementation of each of the above-mentioned units.
[0142] There are a few points to note:
[0143] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0144] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0145] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A data transmission method for transmitting data from a first memory to a second memory, wherein the first memory comprises a memory bank with a depth of N, a first memory cell of the memory bank at each depth having a width of K bits, and the second memory comprises a plurality of second memory cells arranged in a continuous manner, and a width of the second memory cell having a width of K bits, the method comprising: Sequentially read the data stored in the first storage unit at each depth of N depths, and sequentially store the data of the N depths into the buffer unit, wherein the data is m bits, and each first storage unit is divided into K / m sub-data units according to m bits from low to high, wherein N, K and m are all positive integers; and writing the data in the memory bank from the buffer unit to the second memory in a specified order, The specified order includes: K / m pieces of data at the same position in adjacent K / m first storage units in the storage bank are stored in the same second storage unit of the second memory.
2. The method according to claim 1, wherein: Data stored in a first storage unit with a low depth among K / m adjacent first storage units in the storage bank is located at a low position in the second storage unit.
3. The method according to claim 1 or 2, wherein: Sequentially reading data stored in a first storage unit at each of the N depths, including: The data stored in the first storage unit at each of the N depths are read in sequence according to the depth order, and the data in the K / m sub-data units stored in the same first storage unit are read in parallel.
4. The method according to any one of claims 1 to 3, wherein: The buffer unit includes a plurality of buffer zones, each buffer zone is used to store one of the m-bit data, and the number of the plurality of buffer zones is greater than or equal to K / m.
5. The method according to claim 4, wherein: The buffer unit includes (K / m) 2 A continuous array of buffers, Writing the data in the storage body from the buffer unit to the second memory in a specified order includes: After the data in the first storage units of K / m depths are sequentially stored in the buffer unit, the data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body are read from the buffer unit.
6. The method according to claim 5, wherein: After sequentially storing the data in the K / m first storage units of depth into the buffer unit, starting to read the data in the K / m sub-data units located at the same position in the adjacent K / m first storage units in the storage body from the buffer unit, including: After sequentially storing the data in the first storage units of K / m depths into the buffer unit, reading in parallel the data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body; and The data in the K / m sub-data units at the same position in the adjacent K / m first storage units are written into the same second storage unit of the second memory.
7. The method according to claim 6, wherein: Continuously indexing the plurality of sub-data units in the storage body in the order of depth from small to large and position from small to large, After sequentially storing the data in the first storage units of K / m depths into the buffer unit, reading in parallel the data in the K / m sub-data units at the same position in the adjacent K / m first storage units in the storage body, comprising: From the first cycle to the K / mth cycle, in each cycle, data in the sub-data unit with index numbers (i-1)×K / m, ...K / m×i-1 is read from the storage body and stored in the buffer unit, where i represents a cycle count value; and Starting from the K / m+1th cycle, for each cycle, while reading out the data in the sub-data unit with index number (i-1)×K / m, ...K / m×i-1 and storing them in the buffer unit, K / m data at the same position in the adjacent K / m first storage units in the storage body are read from the buffer unit in parallel.
8. A data transmission device, for transmitting data from a first memory to a second memory, wherein the first memory comprises a memory bank with a depth of N, a first memory cell of the memory bank at each depth having a width of K bits, and the second memory comprises a plurality of second memory cells arranged in a row, a width of the second memory cell having a width of K bits, the device comprising: Buffer unit; A reading unit is configured to sequentially read the first storage unit storage at each depth of N depths , and sequentially storing N depths of data in the buffer unit, wherein the data is m bits, and each first storage unit is divided into K / m sub-data units according to m bits from low to high, wherein N, K and m are all positive integers; and a writing unit configured to write the data in the storage body from the buffer unit to the second memory in a specified order, The specified order includes: K / m pieces of data at the same position in adjacent K / m first storage units in the storage bank are stored in the same second storage unit of the second memory.
9. The device according to claim 8, wherein: The buffer unit comprises: a plurality of registers, each register being configured to store the data read from the first storage unit, The reading unit comprises: a plurality of first multiplexers, the plurality of first multiplexers being configured to read data stored in a first storage unit at each depth of N depths of each of the memory banks, and sequentially store the data stored in the first storage unit at each depth into K / m registers of the plurality of registers, The write unit includes: multiple second multiplexers, which are configured to select K / m data located at the same position in adjacent K / m first storage units in the storage body from the multiple registers, and write them into the same second storage unit of the second memory.
10. The device according to claim 8 or 9, wherein: The first memory includes P storage bodies, the number of the buffer units is P, each buffer unit includes a preset number of registers, and the P storage bodies correspond to the P buffer units in a one-to-one manner, wherein P is a positive integer. The device according to claim 10 , wherein the preset number is a positive integer greater than or equal to K / m.
12. An electronic device comprising: The data transmission device according to any one of claims 8 to 11; the first memory; as well as the second memory, The first memory is coupled to the data transmission device, and the data transmission device is coupled to the second memory.
Citation Information
Patent Citations
Data transmission method, data transmission device and electronic device
CN117725002B
Memory controller and memory access control method
CN102567241A
Data reading method and data reading circuit
CN112506567A
Data transmission method, data transmission device and electronic equipment
CN117725002A
Operation circuit and method of operation
US20200278798A1