Data transfer method, data transfer circuit, and computing chip

CN122450387BActive Publication Date: 2026-09-04SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610832929.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-04
Estimated Expiration
2046-06-10

AI Technical Summary

Benefits of technology

[0009] According to the method proposed in this disclosure, firstly, block layout and linear layout can be converted as needed, thus taking into account the advantages of both linear and block layouts. It can ensure the convenience of users' address operations on tensor data under linear layout, and improve the memory access throughput in matrix operation scenarios through block layout. Secondly, by receiving data transfer requests indicating the source layout and target layout, the target address after layout conversion is determined by using multiple parameters related to tensor data and layout indicated by the data transfer request. This solves the address matching problem between the row-by-row scanning logic under linear layout and the block storage logic under block layout, ensuring the data integrity and spatial consistency of tensor data during cross-layout transfer. Furthermore, the complex multi-level address offset calculation is transformed from the software logic of a general-purpose processor to the hardware circuit in the data transfer circuit at the underlying level, significantly reducing the computational latency and instruction overhead during the layout conversion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122450387B_ABST
    Figure CN122450387B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data transfer method, a data transfer circuit and a computing chip. A data transfer method is performed by a data transfer circuit, the method comprising: receiving a data transfer request, the data transfer request being used to request to transfer tensor data to be transferred from a source storage area to a target storage area, the data transfer request indicating a source layout type and a target layout type of the tensor data, wherein one of the source layout type and the target layout type is a linear layout, and the other is a blocked layout; determining one or more source addresses of the tensor data in the source storage area and one or more target addresses in the target storage area based on a transfer parameter indicated by the data transfer request; loading the tensor data from the source storage area based on the one or more source addresses; and writing the tensor data into the target storage area based on the one or more target addresses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a data transfer method, a data transfer circuit, and a computing chip. Background Technology

[0002] In general-purpose computing, computing units typically process large amounts of data when performing computational tasks, and this data is generally stored in memory or buffers. To optimize memory access efficiency and adapt to operator architectures, data will have different layouts in the storage space.

[0003] In some computing scenarios, to improve the efficiency of matrix operations, data may be distributed in a block layout in the storage area. A block layout divides a large two-dimensional or multi-dimensional matrix into multiple blocks, ensuring that the data within each block is stored contiguously in memory. This layout reduces the frequency of memory switching caused by cross-row accesses when performing matrix multiplication or convolution operations.

[0004] In other computing scenarios, to make the data arrangement more intuitive and easier for users to manipulate, data may be stored in a linear layout. The characteristics of a linear layout are logical continuity and consistency of physical addresses: data is arranged sequentially in row-major or column-major order, with the end of the first row immediately following the beginning of the next. This layout is developer-friendly, conforms to the memory access habits of mainstream programming languages, and facilitates addressing operations. Summary of the Invention

[0005] This disclosure provides a data transfer method, a data transfer circuit, and a computing chip.

[0006] According to one aspect of this disclosure, a data transfer method is provided, executed by a data transfer circuit, the method comprising: receiving a data transfer request for requesting the transfer of tensor data to be transferred from a source storage region to a target storage region, the data transfer request indicating a source layout type and a target layout type of the tensor data, wherein one of the source layout type and the target layout type is a linear layout and the other is a block layout; determining one or more source addresses of the tensor data in the source storage region and one or more target addresses in the target storage region based on transfer parameters indicated by the data transfer request; loading the tensor data from the source storage region based on the one or more source addresses; and writing the tensor data into the target storage region based on the one or more target addresses.

[0007] According to another aspect of this disclosure, a data transport circuit is provided, including a circuit module for performing the aforementioned method.

[0008] According to another aspect of this disclosure, a computing chip is provided that includes the aforementioned data transport circuitry.

[0009] According to the method proposed in this disclosure, firstly, block layout and linear layout can be converted as needed, thus taking into account the advantages of both linear and block layouts. It can ensure the convenience of users' address operations on tensor data under linear layout, and improve the memory access throughput in matrix operation scenarios through block layout. Secondly, by receiving data transfer requests indicating the source layout and target layout, the target address after layout conversion is determined by using multiple parameters related to tensor data and layout indicated by the data transfer request. This solves the address matching problem between the row-by-row scanning logic under linear layout and the block storage logic under block layout, ensuring the data integrity and spatial consistency of tensor data during cross-layout transfer. Furthermore, the complex multi-level address offset calculation is transformed from the software logic of a general-purpose processor to the hardware circuit in the data transfer circuit at the underlying level, significantly reducing the computational latency and instruction overhead during the layout conversion process.

[0010] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description

[0011] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of this disclosure. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0012] Figure 1 A schematic diagram of the structure of a general-purpose graphics processing unit (GPU) is shown. Figure 2A A schematic diagram of tensor data stored in a block layout is shown; Figure 2B A schematic diagram of tensor data stored in a linear layout is shown; Figure 3 A schematic flowchart of a data transfer method according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram of a data transport circuit according to an embodiment of the present disclosure is shown; Figure 5 A schematic diagram of a computing chip according to an embodiment of the present disclosure is shown. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0015] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. As used herein, the term "multiple" means two or more, and the term "based on" should be interpreted as "at least partially based on". Furthermore, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations thereof.

[0016] The data transfer methods disclosed herein are related to GPUs. Figure 1 A schematic diagram of the structure of a general-purpose graphics processing unit (GPU) is shown.

[0017] refer to Figure 1 The GPU module 100 shown includes one or more streaming multiprocessors (e.g., streaming multiprocessor 102), and one or more scheduling and distribution modules (e.g., thread block scheduling and distribution module 101) for distributing computing tasks to the streaming multiprocessors. The streaming multiprocessor 102 may internally include one or more execution cores (e.g., core 104, core 105), and a scheduling and distribution module (e.g., thread bundle scheduling and distribution module 103) for issuing instructions to the cores. In some examples, the streaming multiprocessor 102 may also include a tensor memory accelerator (TMA) or a data transport module with similar functionality to transport tensor data from global memory to shared memory. Cores 104 and 105 may internally include one or more arithmetic logic units, floating-point arithmetic units, dedicated tensor computation units (e.g., multiplication units for performing matrix multiplication and accumulation operations), etc.

[0018] GPUs can store and process tensor data. Generally, tensor data can be viewed as a generalization of multidimensional arrays: zero-dimensional tensors are scalars, one-dimensional tensors are vectors, two-dimensional tensors are matrices, and higher-dimensional tensors can capture complex high-dimensional features. The GPU module 100 can have multiple storage levels to store tensor data or related parameters. For example, tensor data can be stored in video memory 109 (e.g., HBM) and moved to the vicinity of cores 104 and 105 via on-chip storage such as L2 cache 108 and L1 cache / shared memory 107 for computation; intermediate results during computation can be temporarily stored in register 106. In some examples, tensor data in video memory 109 or on-chip storage such as L2 cache 108 and L1 cache / shared memory 107 can be arranged in memory or buffers using the aforementioned block layout or linear layout. In some examples, the movement of tensor data from video memory 109 to L1 cache / shared memory 107 can be performed asynchronously by a tensor memory accelerator (TMA) or a data movement unit with similar functionality. In some examples, the data movement unit can move tensor data from video memory 109 to L1 cache / shared memory 107, or from one part of video memory 109 to another part, or from one part of L1 cache / shared memory 107 to another.

[0019] For common matrix operations in GPU computing (such as Generalized Matrix Multiplication (GEMM) or Matrix Multiplication and Accumulation (MACC) calculations), tensor data is typically represented and used in computation in the form of the aforementioned multidimensional array. After the instruction is issued to core 104 or 105, the arithmetic unit in the core can multiply and accumulate the elements at corresponding positions in the tensor data to complete the corresponding operation.

[0020] Figure 1 This is merely one exemplary GPU architecture; however, those skilled in the art will recognize other GPU architectures, and this disclosure is not intended to limit them.

[0021] Tensor data can be stored using either a block layout or a linear layout. For a tensor data matrix A with y rows and x columns, Figure 2A This diagram illustrates a method for storing tensor data in a block layout. Figure 2A Each grid cell represents a block in the block layout 201, which includes multiple blocks (or row blocks) in the row direction and multiple blocks (or column blocks) in the column direction, forming an 8-row, 8-column block layout. In such an example, as... Figure 2AAs shown, multiple blocks in block layout 201 can be stored consecutively in a first-direction priority order. In some examples, the first direction can be the row direction. Blocks are stored in the storage area in a row-direction priority order; for example, they can be arranged first by row number in ascending order, and then by column number in ascending order within the same row. That is, after all blocks in the first row are stored, all blocks in the next row are stored. In other examples, the first direction can be the column direction. Blocks are stored in the storage area in a column-direction priority order; for example, they can be arranged first by column number in ascending order, and then by row number in ascending order within the same column. That is, after all blocks in the first column are stored, all blocks in the next column are stored. The 8-row, 8-column block layout here is only an example. For a tensor data matrix A with y rows and x columns, the tensor data is stored in blocks of a predetermined block size. The block layout can have a corresponding number of blocks in each row and each column, respectively. The granularity of the predetermined block size will be discussed later. Figure 3 The method according to this disclosure will be described in detail. It will be understood that in other examples, the first direction may be a column direction, and this disclosure is not limited thereto.

[0022] For a tensor data matrix A with y rows and x columns of the same data size, Figure 2B A schematic diagram of tensor data stored in a linear layout is shown. As a concrete, non-limiting example, taking tensor data matrix A where each number is in INT4 format, it can be seen that... Figure 2B As shown, each element (or each number in INT4 format) in the tensor data matrix A can be stored contiguously in first-direction priority order. In some examples, the first direction can be the row direction. Figure 2A Similar to the situation in [the text], the elements of tensor data are stored in the storage area in row-major order. For example, they can be arranged first by row number in ascending order, and within the same row, they can be arranged by column number in ascending order. That is, after all the elements in the first row are stored, all the elements in the next row are stored. In other examples, the first direction can be column-major. The elements of tensor data are stored in the storage area in column-major order. For example, they can be arranged first by column number in ascending order, and within the same column, they can be arranged by row number in ascending order. That is, after all the elements in the first column are stored, all the elements in the next column are stored. It is understood that in other examples, the first direction can be column-major, and this disclosure is not limited to this.

[0023] Below, for reference Figure 3 An illustrative flow 300 is provided to describe a data transfer method performed by a data transfer circuit according to an embodiment of the present disclosure.

[0024] At S301, a data transfer request is received. The data transfer request is used to request that tensor data to be transferred be transferred from the source storage area to the target storage area. The data transfer request indicates the source layout type and the target layout type of the tensor data, wherein one of the source layout type and the target layout type is a linear layout and the other is a block layout.

[0025] At S302, based on the transport parameters indicated by the data transport request, one or more source addresses of the tensor data in the source storage area and one or more destination addresses in the destination storage area are determined.

[0026] At S303, tensor data is loaded from the source storage region based on one or more source addresses.

[0027] At S304, tensor data is written to the target storage area based on one or more target addresses.

[0028] According to the method of this disclosure, flexible conversion between linear and block layouts can be achieved using hardware modules. This preserves the intuitive convenience of tensor data addressing and manipulation under linear layout while significantly optimizing memory access throughput in matrix operation scenarios using block layout. By parsing data transfer requests containing source and target layout types, the method of this disclosure can calculate the target address after layout conversion, effectively bridging the row-by-row scanning logic of linear layout and the discrete storage logic of block layout, thus solving the cross-layout addressing problem. Furthermore, according to embodiments of this disclosure, the addressing problem, originally implemented using software logic, is offloaded to dedicated data transfer circuit hardware. This hardware-based low-level acceleration scheme greatly reduces the computational latency and instruction load of layout conversion, eliminating software-level performance bottlenecks.

[0029] Specifically, storing tensor data in a block-based layout in the storage area can improve memory access efficiency. This layout is suitable for GPU matrix operations. Storing tensor data in a linear layout is intuitive and facilitates addressing and accessing tensor data. However, in some cases, users may face more complex problems. For example, after model inference, the result matrix stored in GPU memory in a block-based layout needs to be exported to host memory for user viewing or post-processing. In this case, it needs to be converted to a linear layout to adapt to post-processing algorithms that process data in a linear order. Another example is that user-input image data might be stored as scan rows in a linear layout, but before entering operator calculations (such as convolution or matrix multiplication), it needs to be converted to a block-based layout. This allows for rapid extraction of pixel data from adjacent rows within a small contiguous address space during matrix multiplication, effectively reducing the overhead of frequent memory switching caused by cross-row access. Therefore, the layout type of tensor data may change when moving it from the source storage area to the target storage area.

[0030] In related technologies, the data transport module is only used to transport tensor data to the target storage area in its original layout type, and cannot perform layout type conversion on the tensor data. In related technologies, when layout conversion is required, the rearrangement of tensor data is performed independently of the transport process and is executed at the software level.

[0031] However, this software rearrangement method requires a significant amount of general-purpose computing core resources to execute tedious coordinate transformations and memory access instructions. Modern GPU matrix operations can involve millions or even hundreds of millions of parameters. If the cores performing mathematical operations are tasked with handling these intricate address transformations, they cannot focus on performing matrix multiplication and accumulation tasks, thus drastically reducing the system's effective computing power utilization.

[0032] Furthermore, in early computing chips, the computing power of computing units was relatively small, and the time spent on the computation itself far exceeded the time spent moving tensor data. Therefore, the overhead of software spending extra time rearranging tensor data was tolerable. However, in modern computing chips, due to the introduction of computing modules with strong matrix computation capabilities such as tensor cores, the throughput of computing units has increased exponentially. This leads to the need to send more tensor data to computing units for matrix computation. At this point, the overhead of software layout transformations significantly restricts the system's computing efficiency, resulting in excess computing power but delayed data movement. Because the software layout transformation speed is very slow, high-performance computing units frequently idle, greatly reducing the actual utilization rate of the chip.

[0033] According to one or more embodiments of this disclosure, a data transfer method is proposed, and a data transfer circuit and computing chip for performing such a data transfer method are provided, so that the layout conversion process can be completed within a very short clock cycle, realizing the operation of transfer as conversion.

[0034] In some examples, the data transfer method of this disclosure can be executed by a Tensor Memory Accelerator (TMA) or a data transfer unit with similar functionality. The TMA is capable of asynchronously transferring data from a source storage area to a target storage area based on the data transfer request, using the layout type and data size of the tensor data as transfer parameters. Therefore, it is suitable as the execution entity of the data processing method described in this disclosure. In such examples, the aforementioned data transfer request can be a data transfer request sent to the TMA, which then performs subsequent operations such as parsing the request, determining the source address, and determining the target address.

[0035] In some examples, a data transfer request can be a data transfer request directed to the TMA, which instructs the TMA to perform operations such as determining the target address and performing layout transformation during the data transfer process.

[0036] The above steps will be described in detail below with reference to some non-limiting embodiments.

[0037] At S301, a data transfer request is received. The data transfer request is used to request that tensor data to be transferred be transferred from the source storage area to the target storage area. The data transfer request indicates the source layout type and the target layout type of the tensor data, wherein one of the source layout type and the target layout type is a linear layout and the other is a block layout.

[0038] Compared to traditional data transfer requests in related technologies that only indicate which tensor data to transfer, the data transfer requests according to embodiments of this disclosure can also indicate the source layout type and target layout type of the tensor data. This can be achieved through cooperation between the software and hardware levels. For example, a software-level layout transformation request interface can be configured to a hardware-level data transfer request, thereby scheduling the data transfer circuitry at the GPU hardware level to enable the execution of the data transfer method according to this disclosure for layout transformation between block layout and linear layout.

[0039] Exemplarily, a data transfer request can be physically represented as a control descriptor or register configuration set generated by a hardware logic unit. Exemplarily, according to embodiments of this disclosure, an updated software instruction parsing logic is provided, so that when a layout transformation is required, the processor's instruction scheduling unit can parse the software-level layout transformation request into a low-level transaction request for the hardware engine (such as a Tensor Memory Accelerator, TMA), and integrate this layout transformation request with data transfer. Exemplarily, such a request may not only contain the basic memory address, but may also carry predefined or configurable hardware parameters characterizing layout features, as described in the various examples herein. Exemplarily, upon receiving the request, the hardware-level address mapping circuit can automatically calculate a non-linear offset based on the source and target layout types. This mechanism achieves a deep coupling of software-level logical flexibility and hardware-level high-bandwidth execution efficiency through software instruction task issuance and independent hardware logic execution.

[0040] In some examples, a data transfer request can be a request issued by the kernel to a data transfer module such as TMA that has data transfer capabilities. In some examples, a data transfer request may include multiple transfer parameters, such as: the source layout type of the tensor data to be transferred is either a block layout or a linear layout, and the target layout type is the other; the tensor data to be transferred has a data size of length x and width y; the base address of the tensor data stored in the source memory region; and other parameters associated with the transfer and layout transformation of the tensor data. The data transfer circuitry can determine the source and target addresses at the hardware level in response to a data transfer request indicating one or more of these transfer parameters, thereby loading the tensor data with the source layout type and writing the tensor data with the target layout type.

[0041] In some examples, the source and destination memory regions can be buffers or memory regions. For instance, the source and destination memory regions can be different memory regions within HBM. Another example is that the source memory region is off-chip HBM, and the destination memory region is on-chip shared memory.

[0042] In some examples, such as Figure 2AAs described, in the block layout, tensor data is stored in blocks of a predetermined block size. The tensor data is an array comprising multiple blocks in both the row and column directions, and these blocks are stored consecutively in a first-direction priority order. In some examples, the first direction can be the row direction. Blocks are stored in the storage area in a row-direction priority order; for example, they can be arranged first by row number in ascending order, and within the same row, they can be arranged in column number in ascending order. That is, after all blocks in the first row are stored, all blocks in the next row are then stored. In other examples, the first direction can be the column direction. Blocks are stored in the storage area in a column-direction priority order; for example, they can be arranged first by column number in ascending order, and within the same column, they can be arranged in row number in ascending order. That is, after all blocks in the first column are stored, all blocks in the next column are then stored.

[0043] For example, the predetermined block size can be related to one or more of the following: the bit width of each element in the tensor data, the data width of each block, the number of rows in each block, and the number of columns in each block. As a specific non-limiting example, the data width of each block can be predetermined to be 512 bytes, and each block can be predetermined to have 8 rows. Then, for an element of tensor data with a 4-bit width (e.g., INT4), each block will have 128 cols = 512 bytes / 8 rows / 4 bits; for an element of tensor data with an 8-bit width (e.g., INT8), each block will have 64 cols = 512 bytes / 8 rows / 8 bits; for an element of tensor data with a 16-bit width (e.g., FP16), each block will have 32 cols = 512 bytes / 8 rows / 16 bits; and for an element of tensor data with a 32-bit width (e.g., FP32), each block will have 16 cols = 512 bytes / 8 rows / 32 bits. In some examples, the predetermined block size can be determined in advance by the data transport circuitry as a predetermined parameter. In other examples, the predetermined block size can also be indicated by transport parameters in the data transport request.

[0044] According to some examples, the predetermined block size can be based on the hardware buffer size, for example, it can correspond to the capacity of the hardware buffer. In some examples, the total amount of data in the blocks within the predetermined block size can correspond to the amount of data carried by a single memory access of the buffer used by the hardware circuitry performing the data transfer method. For example, the total amount of data in the blocks can be configured to match the number of bytes carried by a single burst transfer of the hardware buffer (e.g., 512B), so that each block can be completed in a single hardware memory access operation during loading or writing. In some examples, the number of rows in the second direction of the block (e.g., 8 rows) can be determined by the total amount of data in the blocks and the alignment granularity of a single hardware memory access.

[0045] In some examples, such as Figure 2B As described, in a linear layout, elements of tensor data are stored contiguously in a first-direction-first order, where the first direction can be either row-oriented or column-oriented. In some examples, the first direction can be row-oriented. For instance, tensor data elements are stored in the storage area in row-oriented order, such as first arranged in ascending order of row number, and then in ascending order of column number within the same row; that is, after storing all elements in the first row, all elements in the next row are stored. In other examples, the first direction can be column-oriented. Tensor data elements are stored in the storage area in column-oriented order, such as first arranged in ascending order of column number, and then in ascending order of row number within the same column; that is, after storing all elements in the first column, all elements in the next column are stored.

[0046] At S302, based on the transport parameters indicated by the data transport request, one or more source addresses of the tensor data in the source storage region and one or more destination addresses in the destination storage region are determined.

[0047] In some examples, determining one or more source addresses of tensor data in a source storage region and one or more target addresses in a target storage region may include: determining the granularity of a data segment of the tensor data based on the source layout type; determining a first data segment of the tensor data corresponding to the granularity, the first data segment being stored contiguously under the source layout type; and determining a first source address corresponding to the first data segment of the tensor data.

[0048] For example, if a data transfer request indicates that the source layout type is a chunked layout, then based on this source layout type, it can be determined that the data fragments of the tensor data are chunked data fragments. In this case, for example in... Figure 2AAs shown, tensor data in block layout 201 comprises multiple blocks. The first data segment of tensor data corresponding to a granularity can be a block, and each block is stored contiguously with a predetermined block size. In the block layout, exemplarily, tensor data within each block can share a source address, and the tensor data within the block is address-contiguous. In such an example, determining one or more source addresses of the tensor data in the source storage area can be done by determining source addresses corresponding to one or more blocks respectively. For example, each block in block layout 201 has its own corresponding source address, and tensor data within each block can be loaded through the same source address. Exemplarily, the addresses between blocks can be non-contiguous. In such an example, when the data transport module transports tensor data in block layout 201, the data transport module can determine the first source address of one or more blocks (e.g., referred to as the first block) among the multiple blocks and load the tensor data within that first block. For example, the data transfer module can load the source address corresponding to each block one by one, and load the tensor data within each block. For two adjacent blocks in the first direction, the source address of the preceding block and the source address of the following block can differ by the data width of one block. For example, if the data width of a block is 512 bytes, then the source addresses of the two blocks differ by 512 bytes.

[0049] For example, if a data transfer request indicates that the source layout type is linear, then the granularity of the tensor data fragments can be determined based on this source layout type. This granularity can be multiple elements determined based on transfer parameters. For instance, under a row-major linear layout, the data fragments can be determined as a single-row continuous sequence of elements. This granularity corresponds to one or more block widths in the row direction of the target block layout (e.g., in the aforementioned predetermined block size, a 4-bit element width corresponds to one or more 128-cols block widths). In such an example, in a linear layout, the physical addresses between rows are not contiguous, and adjacent rows are separated by the row-direction data width of the tensor data under the linear layout (e.g., an image width, such as...). Figure 2BIn the target block layout (x cols), each row in each block has an independent source address under the source layout of the linear layout. To construct a complete block, corresponding data fragments can be loaded and aggregated from multiple non-contiguous source addresses. For example, under a column-major linear layout, this granularity corresponds to the target block layout, where data fragments can be defined as a single-column contiguous sequence of elements, corresponding to a block height in the column direction of the target block layout. This is because in a column-major linear layout, the physical addresses between columns are not contiguous, and adjacent columns are separated by tensor data along the column direction (e.g., the total height of an image). Since each block in the target block layout can be a region consisting of multiple rows and columns, where each column fragment has an independent source address under the source linear layout, in such an example, to construct a complete block, corresponding data fragments can be loaded and aggregated from multiple non-contiguous source addresses.

[0050] According to embodiments of this disclosure, by dynamically determining the data segment granularity based on the source layout type, it is possible to achieve adaptive addressing under different layout types, ensuring that each load operation can obtain a physical continuous data segment corresponding to the source layout type, thereby improving the utilization of the data bus.

[0051] In some examples, the first source address can be determined based on the data size of the tensor data to be moved, the source storage location of the tensor data to be moved in the source storage region, and the position of the first data segment in the tensor data according to the source layout type. In such examples, the data moving request can include the data size of the tensor data to be moved and the source storage location of the tensor data to be moved in the source storage region. The data moving circuit can determine the source address of the corresponding data segment based on the parameters included in the data moving request and by determining the position of the currently computed data segment in the tensor data according to the source layout type. When the first data segment is the initial data segment in the tensor data, the position of the first data segment in the tensor data can be represented as a zero offset; when the first data segment is a subsequent data segment in the tensor data, the position of the first data segment in the tensor data according to the source layout type can be determined based on the length of the preceding data segment and the source layout type.

[0052] For example, the data size of the tensor data to be transferred may include, for example, the length and width of the tensor data matrix. For instance, in... Figure 2A and Figure 2BIn this context, the size of the tensor data to be moved corresponds to the y-th row and x-th column of the tensor data matrix A. Furthermore, the tensor data matrix A contains a total of x*y numbers, each corresponding to an element in the tensor data. In the calculation of the first source address, the size of the tensor data determines the wrapping rules in memory. For example, in a row-major linear layout, determining the tensor data size determines the row step. Without data size information, the hardware cannot determine how many bytes to skip to reach the next row.

[0053] For example, the source storage location can be indicated by the base address of the tensor data to be transferred in the source storage region, or by another address corresponding to the location in the source storage region where the tensor data to be transferred begins to be stored. In the calculation of the first source address, the source storage location can determine the absolute physical location of the tensor data in memory. The memory allocated to the tensor data by the system does not start from address 0, but has a starting address; the address corresponding to the source storage location is the anchor point for calculating the source address. Only by adding the calculated relative offset to the address corresponding to the source storage location can the data transfer circuit accurately access the tensor data from the source storage region.

[0054] For example, the position of the first data fragment in the tensor data according to the source layout type can include the position of the first data fragment under the source layout type. In the calculation of the first source address, the position of the first data fragment can determine the address relative offset of the first data fragment relative to the source storage location under the source layout type. This parameter can determine which logical coordinate of the tensor the currently transferred tensor data is located at (e.g., which block, or which row and column element). For example, if the source layout type is a block layout, with... Figure 2A For example, Figure 2A There are n rows of blocks. The position of each block in block layout 201 can be indicated by the block index (x_block_id, y_block_id) in the row and column directions. If the first data fragment corresponds to the lower left block in block layout 201, then the position of the first data fragment in the tensor data according to the source layout type is (n-1, 0). For example, return to reference Figure 2B In an example where the source layout type is a linear layout, such as Figure 2B As shown, the tensor data contains x*y elements. The position of each data segment in the linear layout can be indicated by the coordinates (x_coord, y_coord) in the row and column directions. The first data segment can correspond to the element at the bottom left corner of the linear layout 202. If the position of each data segment in the linear layout is indicated by the coordinates (x_coord, y_coord) in the row and column directions, and the first direction corresponds to the row direction, then the position of the first data segment in the tensor data with the source layout type is (x-1, 0).

[0055] For example, the data transport circuit can determine the first source address based on the aforementioned parameters. Specific examples of determining the source address corresponding to a data segment based on different source layout types and block layout types will be discussed below, and will not be repeated here.

[0056] For example, by coupling the data size of the tensor data to be transferred, the source storage location of the tensor data to be transferred in the source storage area, and the position of the first data segment in the tensor data according to the source layout type, the edge alignment problem can be accurately handled.

[0057] In some examples, determining one or more source addresses of tensor data in a source storage region and one or more target addresses in a target storage region may include: splitting a first data segment into two or more sub-segments corresponding to a target layout, such that each of the two or more sub-segments is stored contiguously under the target layout, and the two or more sub-segments are not stored contiguously with respect to each other under the target layout; and determining two or more target addresses corresponding to the two or more sub-segments, wherein writing tensor data to the target storage region based on the one or more target addresses includes: writing the two or more sub-segments to the target storage region based on the two or more target addresses.

[0058] For example, a first data segment may correspond to multiple target addresses, and each sub-segment within the first data segment may correspond to a single target address. The reason for splitting the first data segment into multiple sub-segments is that the target address calculation logic differs between block layouts and linear layouts. Specifically, for example, if the source layout type is a block layout and the target layout type is a row-major linear layout, then the first data segment can be determined as one of multiple blocks in a block layout, as described in S302. A linear layout requires tensor data to be continuous along scan lines, while a block in a block layout spans multiple lines. Therefore, a block in a block layout will be split and distributed across different lines in the linear address space. To meet the requirement of continuous writing in the target layout, according to examples of this disclosure, a block can be split into two or more sub-segments corresponding to different scan lines. Similarly, if the target layout type is a block layout, which requires tensor data within a block to be stored continuously, a data segment in a linear layout may span multiple blocks. Therefore, a data segment in a row-first linear layout will be split and distributed across different blocks as a single row of tensor data within a block in the address space of a block layout. Thus, according to the examples of this disclosure, a data segment in a linear layout can be split into sub-segments corresponding to different blocks. For example, when the source layout type is a block layout and the target layout type is a linear layout, the first data segment can be a block, which can be split into two or more sub-segments corresponding to the linear layout. Each sub-segment can be a single row of tensor data within a block.

[0059] As can be seen, in this example, after splitting the first data segment into two or more sub-segments, each sub-segment itself is stored contiguously in the target layout, while the sub-segments are not stored contiguously with each other in the target layout. For example, when the first data segment is a block, a sub-segment can be a row of tensor data within that block. The tensor data in this sub-segment is stored contiguously in the linear layout, but the sub-segments are separated from each other, that is, between two rows of tensor data in the block, by the width of the entire tensor data in the linear layout. Similarly, when the first data segment is a consecutive sequence of several elements in a row in the linear layout, a sub-segment can be, for example, several elements in the data segment corresponding to one block width in the row direction of the predetermined block size. The tensor data in this sub-segment is stored contiguously in the block layout, but the sub-segments are separated from each other, that is, between several elements of the data segment, by the width of one block in the block layout.

[0060] In such an example, determining one or more target addresses in a target storage region may include determining two or more target addresses corresponding to two or more sub-segments. In such an example, the sub-segments may be non-contiguous, with each sub-segment having its own corresponding target address. Exemplarily, writing tensor data to a target storage address may include writing each sub-segment to the target storage region based on the target address of each sub-segment.

[0061] For example, by further splitting the first data fragment into multiple sub-fragments that are physically contiguous under the target layout and assigning an independent target address to each, the dimensional mapping conflict between the two-dimensional aggregated block layout and the one-dimensional continuous scanning linear layout can be effectively resolved. This splitting of sub-fragments ensures that each sub-fragment written to the target storage area meets the continuity requirements of the target layout, thereby avoiding the overlap or omission of tensor data during the conversion process.

[0062] In some examples, the method may include converting a chunked layout to a linear layout. In such examples, the source layout type can be a chunked layout, and the target layout type can be a linear layout. In such examples, the first data fragment can be one of multiple chunks in a chunked layout, and each of two or more sub-fragments is continuous along a first direction in the first data fragment. For example, the first direction can be a row direction, in which case the source layout type is a row-oriented chunked layout, and the target layout type is a row-oriented linear layout. As described in the example in S302, the first data fragment can be one of multiple chunks in a chunked layout, and each of two or more sub-fragments can be a row of tensor data in that chunk, in which case each sub-fragment itself is continuous along the row direction in the first data fragment.

[0063] In this example, the transport parameters may include the data size of the tensor data to be transported and the base address of the tensor data to be transported in the source storage region. Determining the first source address corresponding to the first data segment of the tensor data includes: determining a first address offset in the first direction and a second address offset in the second direction of the first data segment based on the block index of the first data segment in a first direction and the block index in the second direction, respectively, wherein the block index of the first data segment in the first direction and the block index in the second direction are based on a predetermined block size and the data size of the tensor data to be transported; and determining the first source address of the first data segment in the source storage region based on the base address, the first address offset, and the second address offset of the tensor data to be transported in the source storage region.

[0064] In addition to the data size of the tensor data to be transferred and the base address of the tensor data to be transferred in the source storage region, the transfer parameters may also indicate a predetermined block size. However, if the predetermined block size has already been determined in advance, the transfer parameters may not be indicated. In other words, as will be further explained later in conjunction with the circuit module, one or more parameters of the data transfer circuit for performing the data transfer method according to embodiments of the present disclosure may be configurable, and one or more of these parameters may also be fixed in the circuit in hardware, and the present disclosure is not limited thereto. As mentioned above, the predetermined block size may be related to the bit width of each element in the tensor data, the data width of each block, the number of rows in each block, and the number of columns in each block, and will not be repeated here.

[0065] Continuing with the example where the first direction is row-wise, in such examples, the second direction can be column-wise. For example... Figure 2A As described above, the position of each block in the block layout can be indicated by block indices in the row and column directions, respectively. Therefore, based on the block index of the first data segment in the first direction (the row direction in this example) and the block index in the second direction (the column direction in this example), the first address offset of the first data segment in the row direction and the second address offset in the column direction can be determined, respectively. Here, "in the row direction" can include the row direction of the block layout, that is, the row direction composed of multiple rows of blocks. Similarly, "in the column direction" can include the column direction of the block layout, that is, the column direction composed of multiple rows of blocks. Furthermore, based on the block index of the first data segment in the column direction, the row (or, in which row block) of the first data segment in the block layout can be determined. And based on the block index of the first data segment in the row direction, the column (or, in which column block) of the first data segment can be determined. Furthermore, since the first data segment can be located within the source layout type using its block indexes in the row and column directions, its first address offset in the row direction and its second address offset in the column direction can be determined. The first address offset in the row direction can be determined using the block index in the column direction. Since the block index in the column direction locates the row of the first data segment, the first address offset of the first data segment in the row direction, preceding its current row in the block layout, can be determined. Similarly, the second address offset in the column direction can be determined using the block index in the row direction. Since the block index in the row direction locates the column of the first data segment, the second address offset of the first data segment in the column direction, preceding its current column in the block layout, can be determined.

[0066] In some examples, if the block index in the column direction is represented as y_block_id, the block index in the row direction is represented as x_block_id, the data size of the tensor data to be moved is x*y, and the predetermined block size is (8row*m col, element bit width is data_format, total block data size is 512B), then the first address offset can include y_block_id*8row*x*data_format / 8bit, and the second address offset can include x_block_id*512B.

[0067] Furthermore, since the predetermined block size limits the size of the first data segment, and the data size of the tensor data to be transferred limits the overall size of the tensor data, the block index of the first data segment in the first direction and the block index in the second direction can be determined by dividing the two or performing a right shift. For example, if the data size of the tensor data to be transferred is x*y, and the predetermined block size is (8row*128col, element width 4bit, total block size 512B), then the tensor data can be split into rows of x>>7 blocks and columns of y>>3 blocks.

[0068] In such an example, since the first address offset and the second address offset corresponding to the position of the first data fragment in the tensor data have been determined, the first source address of the first data fragment in the source storage area can be determined based on the aforementioned address offsets and the base address of the tensor data to be transferred in the source storage area.

[0069] In this example, determining two or more target addresses corresponding to two or more sub-fragments may include: determining the total data width of the tensor data along a first direction in a block layout based on the data size of the tensor data; determining the address offset of the first data fragment across the second direction based on the block index of the first data fragment in a second direction and the total data width of the tensor data along the first direction in a block layout; determining the address offset of the sub-fragment across the second direction based on the first direction data width of the tensor data along the first direction in a linear layout; determining the address offset of the first data fragment along the first direction based on the block index of the first data fragment in the first direction and the fragment data width of the first data fragment along the first direction; and determining two or more target addresses of the two or more sub-fragments in the target storage region based on the address offset of the first data fragment along the first direction, the address offset of the first data fragment across the second direction, and the address offset of the sub-fragment across the second direction.

[0070] As mentioned earlier, in this example, the first direction can be the row direction, the second direction can be the column direction, the first data segment can include a block, and two or more sub-segments can include tensor data for each row in that block. To determine the target addresses corresponding to each sub-segment, the approach to calculating the target addresses can include: First, based on the block index in the column direction and the total data width of the tensor data along the first direction in the block layout, the address offset of the first data segment across the column direction can be determined; in other words, the starting position of the row containing the first data segment in the block layout within the linear layout can be determined. This is because the block index of the first data segment in the column direction indicates how many row blocks the first data segment skips. Furthermore, combined with the total data width of the tensor data along the row direction in the block layout, the address offset corresponding to the skipped rows can be determined, which is equivalent to determining the starting position of the row containing the first data segment within the linear layout. The total data width of the tensor data along the first direction in the block layout can include the total data width extending along the row direction in the block layout, that is, the total number of bytes corresponding to each row block. The total data width can be determined based on the data size of the tensor data.

[0071] In some examples, if the block index in the column direction is represented as y_block_id, the block index in the row direction is represented as x_block_id, the data size of the tensor data to be moved is represented as x*y, and the predetermined block size is represented as (8row*m col, element bit width is data_format, total block data size 512B), then the total data width of the tensor data along the first direction can include 8row*x*data_format / 8bit. The address offset of the first data segment across the column direction can include y_block_id*8row*x*data_format / 8bit.

[0072] In this example, continuing the approach to calculating the target address, after determining that the row containing the first data segment in the block layout is at the beginning of the linear layout, we can determine the address offset of each sub-segment of the first data segment across the second direction (the column direction in this example). In other words, after locating the corresponding row of the first data segment in the block layout, we need to determine which row each sub-segment is located in within the linear layout. Specifically, after calculating the address offset of the first data segment across the column direction, we can only determine which row block the first data segment is in. However, a row block in the block layout can contain several rows of tensor data. Therefore, we also need to determine which row within that row block each sub-segment of the first data segment is located in. Although the source addresses of the sub-segments in the first data segment are contiguous in the block layout, in the target linear layout, the sub-segments are separated by the width of the tensor data along the first direction in the linear layout. Therefore, we must multiply the row number of the sub-segment within the block by the width of the first direction data to move it to the correct row. The first-direction data width of the tensor data in a linear layout can include the row-direction data width of the tensor data extending along the row direction in the linear layout, that is, the number of bytes included in each row of the tensor data to be transferred. Based on this first-direction data width, the address offset of each sub-segment in the first data segment across the second direction can be determined.

[0073] In some examples, if the block index in the column direction is represented as y_block_id, the block index in the row direction is represented as x_block_id, the data size of the tensor data to be moved is represented as x*y, and the predetermined block size is represented as (8row*m col, element bit width is data_format, total block data size is 512B), then the data width in the first direction can include x*data_format / 8bit. Therefore, the address offset across the second direction for the 8 sub-segments in the first data segment (corresponding to 8 rows of tensor data in a block) can include x*data_format / 8bit*row_in_block.

[0074] In this example, continuing the approach to calculating the target address, after determining which row each sub-fragment is located in the linear layout, the address offset of the first data segment along the first direction can be determined based on the block index in the row direction and the segment data width of the first data segment along the first direction. In other words, after determining the row in which each sub-fragment is located in the linear layout, it can be determined which block within that row the sub-fragment belongs to. The segment data width of the first data segment along the first direction can include the data width corresponding to a row of tensor data extending from a block along the row direction; that is, the number of bytes corresponding to the width of a block. Specifically, since the row in which each sub-fragment is located in the linear layout has been determined, but each block in that row only contributes the amount of data corresponding to the segment data width, the amount of data corresponding to the segment data width of several blocks in the row direction preceding the current block can be skipped based on the block index in the row direction.

[0075] In some examples, if the block index in the column direction is represented as y_block_id, the block index in the row direction is represented as x_block_id, the data size of the tensor data to be moved is represented as x*y, and the predetermined block size is represented as (8row*m col, element bit width is data_format, total block data size is 512B), the data width of the first data segment along the first direction can include (512B / 8row / data_format)*data_format / 8bit, then the address offset of the first data segment along the first direction can include x_block_id*(512B / 8row / data_format)*data_format / 8bit.

[0076] Therefore, based on the address offset of the first data segment along the first direction, the address offset of the first data segment across the second direction, and the address offset of the sub-segment across the second direction, two or more target addresses of two or more sub-segments in the target storage region can be determined. In some examples, by adding the aforementioned offsets and then based on the base address of the tensor data in the target storage region as indicated in the transport parameters, the target address of each sub-segment in the target storage region can be determined.

[0077] In some examples, the conversion from a linear layout to a block layout may be included: the source layout type is a linear layout, and the target layout type is a block layout. In such an example, the first data fragment is a contiguous data fragment stored in the linear layout; each of the two or more sub-fragments is contiguous along a first direction within a single block in the block layout. For example, the first direction may be a row direction, in which case the source layout type is a row-oriented linear layout, and the target layout type is a row-oriented block layout. The first data fragment may be a contiguous data fragment stored in the linear layout, and may include a single-row contiguous sequence of elements, as described in the example in S302. Each of the two or more sub-fragments may correspond to a block width in the row direction of the target block layout, and each sub-fragment itself is contiguous along the row direction within a single block in the block layout.

[0078] In this example, the transport parameters include the data size of the tensor data to be transported and the base address of the tensor data to be transported in the source storage area. Determining the first source address corresponding to the first data segment of the tensor data includes: determining the address offset of the first data segment in a linear layout based on the coordinates of the first data segment in the first direction and the second direction respectively; and determining the first source address of the first data segment in the source storage area based on the base address and address offset of the tensor data to be transported in the source storage area.

[0079] For example, in addition to the data size of the tensor data to be transferred and the base address of the tensor data to be transferred in the source storage area, the transfer parameters may also indicate a predetermined block size. However, if the predetermined block size has already been determined in advance, the transfer parameters may not indicate it. As mentioned above, the predetermined block size may be related to the bit width of each element in the tensor data, the data width of each block, the number of rows in each block, and the number of columns in each block.

[0080] In this example, since the first direction is the row direction, the second direction can be the column direction. For example... Figure 2B As described above, the position of each data segment in the linear layout can be indicated by coordinates in the row and column directions, respectively. Therefore, based on the coordinates of the first data segment in the first direction (the row direction in this example) and the coordinates in the second direction (the column direction in this example), the address offset of the first data segment in the linear layout can be determined. Specifically, based on the coordinates of the first data segment in the column direction, the row in the linear layout can be determined. And based on the coordinates of the first data segment in the row direction, the column in which the first data segment is located can be determined. Furthermore, since the first data segment can be located in the linear layout using its coordinates in the row and column directions, the address offset of the first data segment in the linear layout can be determined.

[0081] In some examples, if the coordinates of the first data segment in the column direction are represented as y_coord and the coordinates in the row direction are represented as x_coord, the data size of the tensor data to be moved is x*y, and the predetermined block size is (8row*mcol, element bit width is data_format, total block data size is 512B), then the address offset corresponding to the row where the first data segment is located can include y_coord*x*data_format / 8bit, and the address offset corresponding to the column where the first data segment is located can include x_coord*data_format / 8bit.

[0082] For example, since the address offset of the first data segment in the linear layout has been determined, the first source address of the first data segment in the source storage area can be determined based on the aforementioned address offset and the base address of the tensor data to be transferred in the source storage area.

[0083] In some examples, determining two or more target addresses corresponding to the two or more sub-fragments may include: for each sub-fragment in the first data segment: determining the total data width of the tensor data along a first direction under a block layout based on the data size of the tensor data; determining the address offset of the sub-fragment across the second direction under a block layout based on the coordinates of the sub-fragment in a second direction and the total data width; determining the address offset of the sub-fragment along the first direction under a block layout based on the coordinates of the sub-fragment in the first direction and the number of elements corresponding to the segment data width corresponding to a predetermined block size; determining the intra-block address offset of the sub-fragment within the corresponding block of the block layout based on the coordinates of the sub-fragment in the second direction and the predetermined block size; and determining the target address of the sub-fragment in the target storage region based on the address offset of the sub-fragment across the second direction under a block layout, the address offset of the sub-fragment along the first direction under a block layout, and the intra-block address offset of the sub-fragment within the corresponding block of the block layout.

[0084] As mentioned earlier, in this example, the first direction can be the row direction, the second direction can be the column direction, and the first data segment can include a data segment of a single row of elements stored contiguously in a linear layout. Two or more sub-segments can include sub-segments within the first data segment that correspond to a block width in the row direction of the target block layout. To determine the target addresses corresponding to each sub-segment, the approach to calculating the target addresses can include: First, based on the coordinates of the sub-segment in the second direction and the total data width, the address offset of the sub-segment across the second direction in the block layout can be determined. In other words, it can be determined which row block in the block layout the sub-segment can be located in. This is because the coordinates of the first data segment in the column direction can indicate how many rows of blocks the sub-segment in the first data segment skips. Furthermore, combined with the total data width of the tensor data along the row direction in the block layout, the address offset of the rows crossed by the sub-segment in the block layout can be determined, which is equivalent to determining which row in the block layout the sub-segment is in. If the coordinates of the first data segment in the column direction are divided into groups according to the number of rows in the block, then it is possible to calculate which row block the sub-segment belongs to under the block layout. As mentioned earlier, the total data width of the tensor data along the first direction under the block layout can include the total data width extending along the row direction under the block layout, that is, the total number of bytes corresponding to each row block. The total data width can be determined based on the data size of the tensor data.

[0085] In some examples, if the coordinates of the sub-fragment in the column direction are represented as y_coord and the coordinates in the row direction are represented as x_coord, the data size of the tensor data to be moved is x*y, and the predetermined block size is (8row*m col, element bit width is data_format, total block data size is 512B), then the total data width of the tensor data along the first direction under the block layout can still be represented as 8row*x*data_format / 8bit, while the address offset of the sub-fragment across the second direction under the block layout can include y_coord / 8*(8row*x*data_format / 8bit).

[0086] Continuing with the example of calculating the target address, after determining that the sub-fragment can be located in the corresponding row block under the block layout, the address offset of the sub-fragment along the first direction under the block layout can be determined based on the coordinates of the sub-fragment in the first direction and the number of elements corresponding to the fragment data width. Determining the address offset of the sub-fragment along the first direction under the block layout can include determining the address offset of the sub-fragment along the direction of the row extension under the block layout, or in other words, it can include determining which block of the row block of the block layout the sub-fragment belongs to. As mentioned earlier, the fragment data width can include the data width corresponding to a row of tensor data extending along the row direction of a block, that is, the number of bytes corresponding to the width of a block. The number of elements corresponding to the fragment data width can include the number of tensor data elements included in a row of a block. Specifically, by using the coordinates of the sub-fragment in the row direction, combined with the number of tensor data elements included in a row of a block, it is possible to determine how many blocks to the left of the sub-fragment are in the row block of the block layout where the sub-fragment is located. Since each block is stored contiguously in a block layout, the address offset of the corresponding block in the block layout can be determined by finding the number of blocks to the left of the sub-fragment.

[0087] In some examples, if the coordinates of the sub-fragment in the column direction are represented as y_coord and the coordinates in the row direction are represented as x_coord, the data size of the tensor data to be moved is x*y, and the predetermined block size is (8row*m col, element bit width is data_format, total block data size is 512B), then the number of elements corresponding to the sub-fragment and the fragment data width can include 512B / 8row / data_format, and the address offset of the sub-fragment along the first direction under the block layout can include x_corrd / (512B / 8row / data_format)*512B.

[0088] Continuing with the example of calculating the target address, after determining which block of the block layout the sub-fragment belongs to, the intra-block address offset of the sub-fragment within the corresponding block of the block layout can be determined, or in other words, which row of the sub-fragment within that block. Depending on the predetermined block size, a block can include multiple rows of tensor data. Based on the coordinates of the sub-fragment in the second direction (the column direction in this example), the row number of the sub-fragment within a block can be determined by taking the remainder with the number of rows in the predetermined block size. Furthermore, based on the predetermined block size, the data width of each row of tensor data within a block is determined; therefore, combined with the row number of the sub-fragment within a block, the intra-block address offset of the sub-fragment within the corresponding block of the block layout can be determined.

[0089] In some examples, if the coordinates of the sub-fragment in the column direction are represented as y_coord and the coordinates in the row direction are represented as x_coord, the data size of the tensor data to be moved is x*y, and the predetermined block size is (8row*m col, element bit width is data_format, total block data size is 512B), then the intra-block address offset of the sub-fragment in the corresponding block of the block layout can include (y_coord (mod 8))*512B / 8row.

[0090] In such an example, based on the address offset of the sub-fragment across the second direction in the block layout, the address offset of the sub-fragment along the first direction in the block layout, and the intra-block address offset of the sub-fragment within the corresponding block in the block layout, two or more target addresses of two or more sub-fragments in the target memory region can be determined. In some examples, by adding the aforementioned offsets and then based on the base address of the tensor data in the target memory region as indicated in the transport parameters, the target address of each sub-fragment in the target memory region can be determined.

[0091] Although the first direction has been exemplified as the row direction in the foregoing discussion, the present disclosure is not limited thereto, and the first direction may also include the column direction. Furthermore, in column-oriented linear and block layouts, the method for determining the target address of each sub-segment can be similar to the methods described above. The present disclosure does not restrict the preferred storage direction of linear and block layouts.

[0092] Return to reference Figure 3 At S303, tensor data is loaded from the source storage region based on one or more source addresses.

[0093] In such an example, tensor data can be loaded from the source storage region based on one or more source addresses (or source addresses corresponding to individual sub-segments) determined in the preceding steps corresponding to the first data segment. In this step, the data transfer module can load tensor data from the source storage region using the source layout type.

[0094] Return to reference Figure 3 At S304, tensor data is written to the target storage area based on one or more target addresses.

[0095] In such an example, tensor data can be written to the target storage area based on one or more target addresses (or target addresses corresponding to each sub-segment) determined in the preceding steps corresponding to the first data segment. In this step, the data transfer module can write the tensor data to the target storage area according to the target layout type.

[0096] The data transport scheme proposed in the embodiments of this disclosure achieves seamless conversion between linear layout and block layout, effectively balancing the ease of data operation at the logical level and the high efficiency of memory access at the execution level. According to the embodiments of this disclosure, while ensuring a user-friendly and intuitive framework for programming and data interaction using linear scan logic, it also takes into account the advantages of memory access throughput in block layout when processing tensor data. According to the embodiments of this disclosure, the complex multi-level address offset calculation is implemented at the low-level by hardware circuitry in the data transport circuit, significantly reducing computational latency and instruction overhead during the layout conversion process.

[0097] The following describes a data transfer circuit according to an embodiment of the present disclosure, which includes circuit modules for performing the aforementioned method.

[0098] The data transfer circuit 400 may include a data transfer request receiving circuit module 401, an address mapping circuit module 402, a source memory area loading circuit module 403, and a target memory area writing circuit module 404. In addition to these units, the data transfer circuit may also include other components; however, since these components are not relevant to the embodiments of this disclosure, their illustrations and descriptions are omitted herein. Furthermore, the specific details of the operations performed by the data transfer circuit according to the embodiments of this disclosure are consistent with those described above. Figure 3 The details described are the same, so repeated descriptions of the same details are omitted here to avoid repetition.

[0099] The data transfer request receiving circuit module 401 can be configured to receive a data transfer request for requesting that tensor data to be transferred be moved from a source storage area to a target storage area. The data transfer request indicates the source layout type and target layout type of the tensor data, wherein one of the source layout type and the target layout type is a linear layout and the other is a block layout.

[0100] Address mapping circuit module 402 can be configured to determine one or more source addresses of tensor data in the source storage region and one or more target addresses in the target storage region based on the transport parameters indicated by the data transport request.

[0101] The source memory region loading circuit module 403 can be configured to load tensor data from a source memory region based on one or more source addresses.

[0102] The target memory area write circuit module 404 can be configured to write tensor data to a target memory area based on one or more target addresses.

[0103] In some examples, after receiving a data transfer request, the data transfer request receiving circuit module 401 can parse the base address of the tensor data, the direction of the layout transformation, and other indications through its internal register interface. These parameters are then passed to the address mapping logic. The address mapping circuit module 402 can integrate a dedicated address generator and a coordinate transformation pipeline consisting of multi-stage adders and shifters. When linear layout to block layout conversion is involved, this circuit module 402 can calculate the aforementioned address offset in hardware parallel. The source memory region loading circuit module 403 can include a read buffer controller that initiates a memory access request to the source memory region based on one or more source addresses generated by the address mapping circuit module 402. The target memory region writing circuit module 404 can be equipped with mask control logic to achieve write accuracy during cross-layout data rearrangement. When the target address calculated by the address mapping circuit module 402 involves unaligned boundaries or tensor edge regions, the mask control logic can generate corresponding byte enable signals to control the storage location of each data segment in the target memory region, avoiding erroneous writes to non-target memory spaces.

[0104] In some examples, this data transport circuitry can be located within a tensor memory accelerator.

[0105] Integrating the aforementioned layout transformation method into the hardware circuitry of the tensor memory accelerator allows complex address offset calculations to be performed by dedicated hardware logic within the TMA, without consuming computation cycles of a general-purpose processor, enabling the core to focus on matrix operations.

[0106] Below, for reference Figure 5 This is a schematic diagram of a computing chip according to an embodiment of the present disclosure.

[0107] For example, such as Figure 5 As shown, the computing chip 500 may include the aforementioned data transport circuit 501.

[0108] Integrating the hardware circuitry for data transfer at the computing chip level significantly reduces invalid data round trips and repeated memory accesses on the bus compared to software-implemented layout conversion schemes, effectively reducing the dynamic power consumption of the chip during data transfer. Moreover, it enables layout conversion to be completed during tensor data transfer, allowing for efficient matching with the core's computing power and timely delivery of the converted data to the core for computation.

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0110] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on. The names of the units described in the embodiments of this disclosure do not, in some cases, constitute a limitation on the unit itself.

[0111] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0112] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

Claims

1. A data transfer method, executed by a data transfer circuit, the method comprising: A data transfer request is received. The data transfer request is used to request that tensor data to be transferred be transferred from a source storage area to a target storage area. The data transfer request indicates the source layout type and the target layout type of the tensor data. One of the source layout type and the target layout type is a linear layout and the other is a block layout. In the linear layout, the tensor data is preferentially and continuously stored along a first direction. In the block layout, the tensor data is preferentially and continuously stored along the first direction with a predetermined block size as the granularity. Based on the transfer parameters indicated by the data transfer request, determine one or more source addresses of the tensor data in the source storage area and one or more target addresses in the target storage area; Based on the one or more source addresses, the tensor data is loaded from the source storage area; as well as Based on the one or more target addresses, the tensor data is written to the target storage area. The total data volume of the blocks in the predetermined block size is consistent with the number of bytes carried by the hardware buffer used by the data transport circuit in a single burst transfer. The number of rows of the blocks of the predetermined block size in a second direction different from the first direction is determined by the total data volume of the blocks and the alignment granularity of a single hardware memory access. Loading the tensor data from the source storage area includes loading the blocks through a single hardware memory access operation during the loading process.

2. The method according to claim 1, wherein, In the linear layout, the elements in the tensor data are stored sequentially along a first direction, which is either the row direction or the column direction.

3. The method according to claim 2, wherein, In the block layout, the tensor data is stored in blocks of a predetermined block size. The tensor data is an array that includes multiple blocks in both the row and column directions, and the multiple blocks are stored consecutively in a priority order along the first direction.

4. The method according to claim 3, wherein, Determining one or more source addresses of the tensor data in the source storage region and one or more destination addresses in the destination storage region includes: The granularity of the data segments of the tensor data is determined based on the source layout type; Determine a first data segment of the tensor data corresponding to the granularity, wherein the first data segment is stored contiguously under the source layout type; and Determine the first source address corresponding to the first data segment of the tensor data.

5. The method according to claim 4, wherein, The first source address is determined based on the data size of the tensor data to be transferred, the source storage location of the tensor data to be transferred in the source storage area, and the position of the first data segment in the tensor data according to the source layout type.

6. The method according to claim 4, wherein, Determining one or more source addresses of the tensor data in the source storage region and one or more destination addresses in the destination storage region includes: The first data segment is split into two or more sub-segments corresponding to the target layout, such that each of the two or more sub-segments is stored contiguously under the target layout, and the two or more sub-segments are not stored contiguously with each other under the target layout; and Determine two or more target addresses corresponding to the two or more sub-segments, and Specifically, writing the tensor data into the target storage area based on the one or more target addresses includes writing the two or more sub-fragments into the target storage area based on the two or more target addresses.

7. The method according to claim 6, wherein, The source layout type is the block layout, the target layout type is the linear layout, and wherein the first data segment is one of a plurality of blocks of the block layout, and each of the two or more sub-segments is continuous along the first direction in the first data segment.

8. The method according to claim 7, wherein, The transport parameters include the data size of the tensor data to be transported and the base address of the tensor data to be transported in the source storage area, wherein determining the first source address corresponding to the first data segment of the tensor data includes: Based on the block index of the first data segment in the first direction and the block index in the second direction (different from the first direction), a first address offset in the first direction and a second address offset in the second direction are determined for the first data segment, wherein the block index of the first data segment in the first direction and the block index in the second direction (different from the first direction) are based on the predetermined block size and the data size of the tensor data to be transferred; and Based on the base address of the tensor data to be transferred in the source storage area, the first address offset, and the second address offset, the first source address of the first data fragment in the source storage area is determined.

9. The method according to claim 8, wherein, Determining the two or more target addresses corresponding to the two or more sub-segments includes: The total data width of the tensor data along the first direction under the block layout is determined based on the data size of the tensor data. The address offset of the first data segment across the second direction is determined based on the block index of the first data segment in the second direction and the total data width of the tensor data along the first direction under the block layout. Based on the tensor data and the first directional data width along the first direction in the linear layout, determine the address offset of the sub-segment across the second direction; Based on the block index of the first data segment in the first direction and the segment data width of the first data segment along the first direction, determine the address offset of the first data segment along the first direction; and Based on the address offset of the first data segment along the first direction, the address offset of the first data segment across the second direction, and the address offset of the sub-segment across the second direction, determine the two or more target addresses of the two or more sub-segments in the target storage region.

10. The method of claim 6, wherein, The source layout type is the linear layout, the target layout type is the block layout, and wherein the first data segment is a data segment stored continuously in the linear layout; each of the two or more sub-segments is continuous along the first direction within a single block in the block layout.

11. The method of claim 10, wherein, The transport parameters include the data size of the tensor data to be transported and the base address of the tensor data to be transported in the source storage area, wherein determining the first source address corresponding to the first data segment of the tensor data includes: Based on the coordinates of the first data segment in the first direction and the second direction, respectively, determine the address offset of the first data segment under the linear layout; and Based on the base address and the address offset of the tensor data to be transferred in the source storage area, the first source address of the first data fragment in the source storage area is determined.

12. The method according to claim 11, wherein, Determining the two or more target addresses corresponding to the two or more sub-segments includes: For each sub-segment in the first data segment: The total data width of the tensor data along the first direction under the block layout is determined based on the data size of the tensor data. Based on the coordinates of the sub-fragment in the second direction and the total data width, determine the address offset of the sub-fragment across the second direction under the block layout; Based on the coordinates of the sub-fragment in the first direction and the number of elements corresponding to the fragment data width, the address offset of the sub-fragment along the first direction under the block layout is determined, wherein the fragment data width corresponds to the predetermined block size; Based on the coordinates of the sub-fragment in the second direction and the predetermined block size, determine the intra-block address offset of the sub-fragment within the corresponding block of the block layout; and Based on the address offset of the sub-fragment across the second direction in the block layout, the address offset of the sub-fragment along the first direction in the block layout, and the intra-block address offset of the sub-fragment within the corresponding block in the block layout, the target address of the sub-fragment in the target storage area is determined.

13. The method according to any one of claims 1-12, wherein, The source storage area and the target storage area are storage areas in buffers or memory.

14. A data transport circuit comprising a circuit module for performing the method of any one of claims 1-13.

15. The data transfer circuit according to claim 14, wherein, The data transfer circuit is located in the tensor memory accelerator.

16. A computing chip, the computing chip comprising the data transport circuit according to claim 14 or 15.

Citation Information

Patent Citations

  • Fusion operator optimization method and device and storage medium

    CN120894218A

  • Tensor memory accelerator and tensor memory acceleration method

    CN121143876A