Data transfer method, data transfer circuit, and computing chip
Patent Information
- Application Number
- CN202610833406.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-10
AI Technical Summary
[0004]为了解决数据搬运电路中集成张量数据转置引发的访存延迟和指令开销等问题,本公开提供了一种数据搬运方法、数据搬运电路及计算芯片
[0008]根据本公开提出的方法,首先,通过在数据搬运电路中集成张量数据转置的逻辑,实现在数据搬运的同时完成维度变换,能够显著减少对内存的访问次数,提高总线利用率;其次,由于转置操作由硬件电路基于地址映射直接实现,本公开无需在片上配置额外的寄存器组来暂存多轮读取的中转数据,从而降低了芯片的硬件开销和功耗;另外,通过将源存储区域中非转置的数据,按照转置后的布局写入目标存储区域,使得后续的执行单元可以通过少量的突发传输直接读取到所需的转置数据,消除了计算过程中的访存等待延迟。
Smart Images

Figure CN122470126B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a data transfer method, a data transfer circuit, and a computing chip. Background Technology
[0002] In the field of general computing, tensor data often needs to be transposed in order to adapt to the architectural requirements of specific mathematical algorithms or hardware accelerators.
[0003] However, in the field of general computing, data transport circuits move tensor data to the buffer in the original non-transposed layout. When transposition is required for computation, the rearrangement of tensor data is independent of the transport process and is performed by the software. The software needs to frequently perform non-continuous read and write operations in the buffer, resulting in significant instruction overhead and memory access latency. Summary of the Invention
[0004] To address the issues of memory access latency and instruction overhead caused by the transposition of integrated tensor data in data transfer circuits, this disclosure provides a data transfer method, a data transfer circuit, and a computing chip.
[0005] According to one aspect of this disclosure, a data transfer method executed by a data transfer circuit is provided, comprising: receiving a data transfer request for transferring tensor data to be transferred from a source storage area to a target storage area, the data transfer request including transpose instruction information indicating whether to perform transpose transfer on the tensor data to be transferred; in response to the transpose instruction information indicating that transpose transfer should be performed on the tensor data, determining one or more source addresses of the tensor data in the source storage area and one or more target addresses in the target storage area based on transfer parameters indicated by the data transfer request; loading the tensor data from the source storage area based on the one or more source addresses; and writing the tensor data into the target storage area based on the one or more target addresses.
[0006] According to another aspect of this disclosure, a data transport circuit is provided, including a circuit module for performing the aforementioned method.
[0007] According to another aspect of this disclosure, a computing chip is provided that includes the aforementioned data transport circuitry.
[0008] According to the method proposed in this disclosure, firstly, by integrating tensor data transposition logic into the data transport circuit, dimensional transformation can be completed simultaneously with data transport, which can significantly reduce the number of memory accesses and improve bus utilization. Secondly, since the transposition operation is directly implemented by the hardware circuit based on address mapping, this disclosure does not require configuring an additional register set on-chip to temporarily store the intermediate data for multiple rounds of reading, thereby reducing the chip's hardware overhead and power consumption. In addition, by writing the non-transposed data in the source storage area into the target storage area according to the transposed layout, subsequent execution units can directly read the required transposed data through a small number of burst transfers, eliminating memory access waiting latency during the calculation process.
[0009] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description
[0010] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of this disclosure. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0011] Figure 1 A schematic diagram of the structure of a general-purpose graphics processing unit (GPU) is shown. Figure 2A A schematic diagram of tensor data stored in a non-transposed layout is shown; Figure 2B A schematic diagram of tensor data stored in a transposed layout is shown; Figure 3 A schematic flowchart of a data transfer method according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram of a data transport circuit according to an embodiment of the present disclosure is shown; Figure 5 A schematic diagram of a computing chip according to an embodiment of the present disclosure is shown. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0014] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. As used herein, the term "multiple" means two or more, and the term "based on" should be interpreted as "at least partially based on". Furthermore, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations thereof.
[0015] The data transfer methods disclosed herein are related to GPUs. Figure 1 A schematic diagram of the structure of a general-purpose graphics processing unit (GPU) is shown.
[0016] refer to Figure 1 The GPU module 100 shown includes one or more streaming multiprocessors (e.g., streaming multiprocessor 102), and one or more scheduling and distribution modules (e.g., thread block scheduling and distribution module 101) for distributing computational tasks to the streaming multiprocessors. The streaming multiprocessor 102 may internally include one or more execution cores (e.g., core 104, core 105), and a scheduling and distribution module (e.g., thread bundle scheduling and distribution module 103) for issuing instructions to the cores. In some examples, the streaming multiprocessor 102 may also include a tensor memory accelerator (TMA) or a data transport module with similar functionality to transport tensor data between various storage areas. Exemplarily, the data transport module can transport tensor data from video memory 109 to a buffer in on-chip storage such as L1 cache / shared memory 107, allowing cores 104 and 105 to directly fetch data from the buffer for computation. Cores 104 and 105 may include one or more arithmetic logic units, floating-point arithmetic units, dedicated tensor computation units (such as multiplication units for performing matrix multiplication and accumulation operations), etc.
[0017] GPUs can store and process tensor data. Generally, tensor data can be viewed as a generalization of multidimensional arrays: zero-dimensional tensors are scalars, one-dimensional tensors are vectors, two-dimensional tensors are matrices, and higher-dimensional tensors can capture complex high-dimensional features. GPU module 100 can have multiple storage levels to store tensor data or related parameters. For example, tensor data can be stored in video memory 109 (e.g., HBM) and moved to the vicinity of cores 104 and 105 via on-chip storage such as L2 cache 108 and L1 cache / shared memory 107 for computation; intermediate results during computation can be temporarily stored in register 106. In some examples, tensor data in on-chip storage such as L1 cache / shared memory 107 can be arranged in a buffer with the aforementioned non-transposed or transposed layout for use by cores 104 and 105 during matrix calculations. In some examples, the movement of tensor data between different storage areas can be done asynchronously by a tensor memory accelerator (TMA) or a data movement unit with similar functionality.
[0018] For common matrix operations in GPU computing (such as Generalized Matrix Multiplication (GEMM) or Matrix Multiplication and Accumulation (MACC) calculations), tensor data is typically represented and used in computation in the form of the aforementioned multidimensional array. After the instruction is issued to core 104 or 105, the arithmetic unit in the core can multiply and accumulate the elements at corresponding positions in the tensor data to complete the corresponding operation.
[0019] Figure 1 This is merely one exemplary GPU architecture; however, those skilled in the art will recognize other GPU architectures, and this disclosure is not intended to limit them.
[0020] Figure 2A This diagram illustrates a method for storing tensor data in a non-transposed layout 201. Figure 2A In the example, tensor data is stored contiguously in channel-major order. Horizontally, an address indicates the elements of a contiguous row of tensor data. Each row represents all elements across all channels within the same pixel dimension; for example, the first row stores elements p0c0 to p0c7 within the same pixel dimension. Vertically, each column represents the distribution of the same channel dimension across different pixel dimensions; for example, the first column shows elements c0p0, c1p0, c2p0, etc., of channel c0 across different pixel dimensions. In the non-transposed layout 201, the elements in each column are discrete in physical memory, separated by all channel widths within the same pixel dimension.
[0021] Figure 2B This diagram illustrates a method for storing tensor data using a transposed layout 202. Figure 2BIn the example, tensor data is stored contiguously in pixel-priority order. Horizontally, an address indicates the elements of a contiguous row of tensor data. Each row represents all elements of the tensor data across all pixel dimensions within the same channel dimension. For example, the row starting at address `addr` stores elements c0p0 to c0p7 across the same channel dimension. Vertically, each column represents the distribution of the same pixel dimension across different channel dimensions. In this transposed layout 202, elements across different pixel dimensions within the same channel dimension are linearly contiguous in physical storage space.
[0022] In matrix computation, computational units may frequently use transposed tensor data for matrix operations. When the tensor data required by the computational unit is arranged in a non-transposed layout 201 in the buffer, because elements of the same channel dimension are discretely distributed across pixel dimensions in physical memory, and conventional data transfer methods can only read one row of data continuously at a time based on one address, the computational unit cannot obtain the required transposed data organized by the channel dimension in a single continuous read. For example, if the parallelism (dp) of the dot product operation required by the computational unit requires inputting 16 pixel-dimensional data of the same channel (e.g., elements p0c0, p1c0 to p15c0) at once, the data transfer circuit must access 16 different physical addresses, extract only a single element at the 0th channel position from each physical storage row, and then perform multiple rounds of concatenation in additional register space to form the transposed data. This approach not only leads to low utilization of bus burst transmission bandwidth and a large amount of address switching overhead for each transfer request, but also increases hardware latency and register resource consumption. In contrast, when tensor data is arranged in the buffer in transposed layout 202, when the computing unit initiates a request for tensor data for the same channel dimension, the read circuit can directly address the consecutive address of the same row, avoiding cross-row reading and number concatenation operations.
[0023] In related technologies, tensor memory accelerators are not transpose-aware. Regardless of whether subsequent calculations require transposed data, they move tensor data to the buffer in the original untransposed layout. When transposition is required for calculation operations, the rearrangement of tensor data is independent of the moving process and is performed by the software. The software needs to frequently perform non-continuous read and write operations in the buffer, resulting in significant instruction overhead.
[0024] To this end, this disclosure integrates transpose-aware logic into the data transport circuit, so that in scenarios requiring transpose computation, the data transport circuit can simultaneously complete the layout transformation during the process of transporting tensor data from a non-computation region (e.g., HBM) to the buffer, and directly write the tensor data into the buffer with transpose layout 202, so that subsequent computation units can obtain the required continuous transpose data through one or a few burst transmissions.
[0025] Below, for reference Figure 3 An illustrative flow 300 is provided to describe a data transfer method performed by a data transfer circuit according to an embodiment of the present disclosure.
[0026] At S301, a data transfer request is received. The data transfer request is used to transfer tensor data to be transferred from the source storage area to the target storage area. The data transfer request includes transpose instruction information indicating whether to perform transpose transfer on the tensor data to be transferred.
[0027] At S302, in response to the transpose instruction information indicating that a transpose transport of tensor data is to be performed, the transport parameters indicated by the data transport request are used to determine one or more source addresses of the tensor data in the source storage area and one or more destination addresses in the destination storage area.
[0028] At S303, tensor data is loaded from the source storage region based on one or more source addresses.
[0029] At S304, tensor data is written to the target storage area based on one or more target addresses.
[0030] Compared to traditional data transfer requests in related technologies that only indicate which tensor data to transfer, the data transfer requests according to embodiments of this disclosure can also indicate whether the tensor data needs to be transposed. This can be achieved through cooperation between the software and hardware levels. For example, the transpose operation in a software-level computation request can be configured into a hardware-level data transfer request, thereby enabling the data transfer unit (e.g., a tensor memory accelerator (TMA)) to be aware of whether the data is transposed. It is understood that the target address can be determined based on the position of the source element after dimension swapping.
[0031] Exemplarily, a data transfer request can be physically represented as a control descriptor or register configuration set generated by a hardware logic unit. Exemplarily, according to embodiments of this disclosure, an updated software instruction parsing logic is provided, so that when a transposition operation is required, the processor's instruction scheduling unit can parse the software-level transposition operation request into a low-level transaction request for the hardware engine (such as a Tensor Memory Accelerator, TMA), and integrate this transposition operation request with data transfer. Exemplarily, such a request can not only contain the basic memory address, but also carry predefined or configurable hardware parameters characterizing layout features, as described in the various examples herein. Exemplarily, upon receiving the request, the hardware-level address mapping circuit can automatically calculate the nonlinear offset based on the source and target layout types. This mechanism achieves deep coupling between the logical flexibility of the software layer and the high bandwidth execution efficiency of the hardware layer through software instruction task issuance and independent execution of hardware logic.
[0032] According to the method of this disclosure, hardware modules can be used to convert tensor data from a non-transposed layout to a transposed layout. This allows the computing unit to acquire the transposed continuous data through one or a few burst transfers when subsequently reading tensor data for computation, eliminating memory access waits caused by the mismatch between the non-transposed layout of the data and the transposed calculation of the matrix. According to the method of this disclosure, by parsing data transfer requests containing transpose indication information, the target address of each element of the tensor data after transposition can be calculated, thereby completing the transposition synchronously during data transfer and avoiding the bandwidth loss caused by multiple rounds of data reading and number concatenation operations required in the non-transposed layout. The method of this disclosure offloads the addressing problem, which originally relied on software logic, to dedicated data transfer circuit hardware, reducing operational complexity. Furthermore, the address mapping relationship automatically generated by the hardware circuit greatly reduces the complexity of software configuration, thus significantly reducing hardware overhead while improving transposition efficiency.
[0033] According to embodiments of this disclosure, the hardware module for data transport is made aware of transpose transport. As a specific, non-limiting example, after the data transport module performs transpose transport on tensor data, the computing unit is able to read the transposed, continuous data from the buffer at once.
[0034] According to the data transport method proposed in the embodiments of this disclosure, by transforming the dimensional transformation requirements in computational operations into transposition transport in a hardware-level data transport circuit, the data transport circuit can natively perceive the transposition mapping and perform transposition processing on tensor data. In this way, the data transport circuit can capture transposition indication information with extremely low latency, without the need for complex software protocol stack parsing, thereby ensuring seamless integration of transposition logic with the data transport pipeline. Moreover, upon receiving a data transport request, the data transport circuit can transport tensor data from non-computational regions (e.g., remote memory or HBM) to regions near the computational core (e.g., on-chip buffers), so that the computational core can retrieve the transposed tensor data from the nearby buffer within a very short period when performing matrix operations, improving the execution efficiency of computational tasks.
[0035] In some examples, the data transfer method disclosed herein may be executed by a tensor memory accelerator (TMA) or a data transfer unit with similar functionality.
[0036] The above steps will be described in detail below with reference to some non-limiting embodiments.
[0037] At S301, a data transfer request is received. The data transfer request is used to transfer tensor data to be transferred from the source storage area to the target storage area. The data transfer request includes transpose instruction information indicating whether to perform transpose transfer on the tensor data to be transferred.
[0038] In some examples, the target storage region can be a buffer used for computational operations. In such examples, the transpose indication information is determined based on whether the tensor data is involved in the computational operation in transposed form.
[0039] In such an example, instead of simply indicating which tensor data to move, the data move request also includes transpose indication information, which can indicate whether the tensor data should be used in subsequent computations in a transposed form. Thus, the data move unit (e.g., a tensor memory accelerator (TMA)) can be aware of whether data is transposed.
[0040] According to embodiments of this disclosure, software-level data computation operations can be transformed into hardware-level data transfer requests, thereby enabling the data transfer circuitry at the GPU hardware level to execute the transpose-aware data transfer method according to this disclosure.
[0041] In some examples, the data transfer request may be a hardware request generated by the instruction scheduling unit during GPU computing kernel execution and sent to a data transfer module such as TMA via a hardware interface. In some examples, the data transfer request may include multiple transfer parameters, such as the base address of the tensor data to be transferred in the source storage region, the base address of the tensor data in the target storage region, and other parameters associated with the transpose transfer of the tensor data. By indicating one or more of these transfer parameters, the data transfer request can trigger a data transfer circuit for executing the method according to this disclosure, thereby determining the source address and the target address, loading the tensor data in the source layout type, and writing the tensor data in the target layout type. Exemplarily, the source storage region may be a storage region in memory, such as, but not limited to, a storage region in HBM. Thus, a read-and-transpose operation can be implemented, allowing data to be stored in the buffer in the form required by the computing unit, reducing subsequent additional transpose operations.
[0042] In some examples, a data transfer request according to this disclosure can be generated by decoding the data transfer instructions issued at the software level, and the data transfer request can be input into the configuration register or pipeline interface of the data transfer circuit.
[0043] In some examples, the data transfer request can be a hardware handshake signal between the data transfer circuit and the instruction scheduling unit. When the software issues transpose-related calculation instructions, the instruction scheduling unit parses them into a hardware control word (or data transfer request) containing transpose indication information, base address, and other relevant parameters, and directly drives the hardware pipeline of the data transfer module to start.
[0044] For example, according to the transpose transport method of this disclosure, the physical location of each element in the target storage region is the location of the corresponding element in the source storage region after swapping the first-dimensional index and the second-dimensional index. In such an example, this mapping relationship can be directly calculated and determined by the data transport circuit based on the transport parameters, without relying on the software layer to rearrange the element positions one by one.
[0045] In some examples, the tensor data to be transferred may include a first dimension and a second dimension; and in the source storage area, multiple elements of the tensor data to be transferred are stored consecutively in order of priority of the first dimension, and in the target storage area, multiple elements of the tensor data are stored consecutively in order of priority of the second dimension, wherein the first dimension and the second dimension are different dimensions.
[0046] In a non-limiting embodiment, with Figure 2A For example, the first dimension of tensor data can include the channel dimension, and the second dimension can include the pixel dimension. In the source storage area, tensor data can be stored according to a non-transposed layout 201, that is, multiple elements in the tensor data are stored contiguously according to channel dimension priority. Here, contiguous storage according to channel dimension priority can include, in the physical address space, all elements on all channels under the same pixel dimension are arranged in adjacent storage areas, thus logically forming rows of tensor data elements based on the same pixel dimension. For example, Figure 2A In this context, for the same pixel dimension, elements p0c0, p0c1, up to p0c7 in consecutive channel dimensions are aggregated and stored contiguously. This means that when a read is initiated for the same pixel dimension, a slice of that pixel dimension across multiple channel dimensions is obtained. Similarly, in the target storage area, tensor data can be stored according to transpose layout 202, that is, multiple elements in the tensor data are stored contiguously according to pixel dimension priority. Here, storing contiguously according to pixel dimension priority can include, in the physical address space, arranging elements in all pixel dimensions under the same channel dimension in adjacent storage areas, thereby logically forming rows of tensor data elements based on the same channel dimension. For example, Figure 2B In this approach, for the same channel dimension, elements c0p0, c0p1, up to c0p7 in consecutive pixel dimensions are aggregated and stored contiguously. This means that when a read is initiated for the same channel dimension, what is obtained is a slice of that channel dimension across multiple pixel dimensions.
[0047] However, it is understandable that tensor data can also have dimensions in the NDHWC format. In some examples, the tensor data is stored in the source storage area in NDHWC format, where N is the batch dimension, D is the depth dimension, H is the height dimension, W is the width dimension, and C is the channel dimension. In such examples, the first dimension can be the C dimension, and the second dimension can be a dimension formed by expanding one or more of the N, D, H, and W dimensions.
[0048] In such an example, the NDHWC format can refer to tensor data being linearly arranged in a stepwise order within a storage region, where the channel dimension (C) changes the fastest and the batch dimension (N) changes the slowest. In this format, different channel data belonging to the same spatial location (defined by D, H, and W) are physically contiguous.
[0049] For example, in the NDHWC format, dimension C is located at the innermost position where physical addresses change the most rapidly. Elements that change along dimension C occupy contiguous physical addresses in the source storage region. In this case, the first dimension can correspond to dimension C, and in such an example, contiguous storage in the source storage region according to the first dimension priority can be represented as the physical contiguity of dimension C in the NDHWC format. During transpose transport, elements along dimension C are redistributed to positions across multiple contiguous address segments in the target storage region. The hardware circuitry can directly calculate the address offset of an element across rows in the target storage region using its index along dimension C.
[0050] In some examples, the N, D, H, and W dimensions collectively constitute the location information of tensor data in batches and space. During transpose transport, these dimensions can be treated as a single pixel or row dimension. By logically flattening one or more of N, D, H, and W, they can be merged into a unified index space, simplifying address calculation to a single index variable. A second dimension can correspond to the flattened dimension, and in such examples, contiguous storage in the target storage region according to the second dimension priority can be represented as spatial dimension elements originally distributed across large strides in the source storage region, which, after flattening and mapping, are rearranged into physically contiguous consecutive address intervals along the same channel in the target storage region. This flexible dimension flattening and mapping method makes this disclosure applicable not only to simple two-dimensional matrix transposes but also to handling layout transformation requirements of high-dimensional tensors in complex formats such as NDHWC, ensuring the versatility of the hardware transport logic for diverse algorithmic scenarios.
[0051] Based on this example, by storing tensor data contiguously in the target storage area according to the second dimension priority, the physically discrete pixel dimensions in the source storage area are linearly and continuously arranged for a single channel in the target storage area. This transposed layout allows the computing unit to obtain continuous transposed tensor data when performing matrix calculations such as matrix transpose operations, without having to frequently jump between different rows for addressing.
[0052] To perform transpose transport on tensor data in the source storage region, the data transport request can be split. In some examples, based on a predetermined granularity in the first dimension, the second dimension of the tensor data to be transported indicated by the data transport request is split at least once to obtain two or more split data transport requests; wherein the split data transport request indicates a data segment of the tensor data to be transported in the second dimension, the data segment consisting of one or more elements of the tensor data to be transported, the number of elements corresponding to the predetermined granularity in the first dimension.
[0053] The reason for splitting data transfer requests is that the transpose operation can distribute elements that are spatially contiguous in the source storage region along the first dimension to different physical locations in the target storage region according to the second dimension's priority order. For example, ... Figure 2A The p0c0-p0c7 data points are distributed across eight non-contiguous physical addresses in the target storage region. If the data transfer request is not split and attempts are made to process multiple rows of elements in different second dimensions simultaneously, the hardware addressing logic will need to compute and open multiple non-contiguous addresses within a single cycle when loading multiple rows of elements from the source storage region. This may exceed the hardware's parallel processing capabilities. For example, simultaneous transposition... Figure 2A Rows p0, p1, and p2 in the array each contain elements from 8 channels. After transposition, the hardware needs to open 24 addresses simultaneously. Therefore, by splitting the data processing requests, the mapping overhead in the transposition process can be simplified.
[0054] Furthermore, the second dimension of the tensor data to be moved, as indicated by the data moving request, can be split at least once based on a predetermined granularity in the first dimension to obtain two or more split data moving requests. In a non-limiting example, if the first dimension includes a channel dimension and the second dimension includes a pixel dimension, then the pixel dimension can be split based on a predetermined granularity in the channel dimension. In a non-limiting example, if there are 16 consecutive elements p0c0-p0c15 in the channel dimension at the same index p0 corresponding to the pixel dimension, then the pixel dimension can be split once based on a predetermined granularity in the channel dimension (e.g., 8 elements in the channel dimension) to obtain two split data moving requests, where one data moving request may include requests p0c0-p0c7 and the other data moving request may include requests p0c8-p0c15.
[0055] Correspondingly, the split data transfer request can indicate a data segment of tensor data in the second dimension, the number of elements in the data segment corresponding to the predetermined granularity in the first dimension. Here, the predetermined granularity in the first dimension can include, if the second dimension is the same, the number of elements in the first dimension increasing one by one to the same number as the predetermined granularity. Here, the tensor data in the second dimension can include tensor data within the same second dimension. Similarly, in the aforementioned example, one data transfer request can include requests p0c0-p0c7, and another data transfer request can include requests p0c8-p0c15. The two split data transfer requests respectively indicate two data segments of tensor data along the pixel dimension (indexed as p0), and each data segment contains 8 elements, which corresponds to the predetermined granularity in the channel dimension.
[0056] In some examples, the predetermined granularity in the first dimension can be determined based on the bus width. Depending on the performance requirements of the actual application scenario, the predetermined granularity can be embedded in the hardware logic of the data transfer circuit (e.g., implemented through hardwiring), or it can be a programmable parameter dynamically set by the software layer through configuration instructions or by rewriting configuration registers, thereby achieving a balance between transfer efficiency and hardware overhead under different task loads. Matching the ability of memory to simultaneously allocate address spaces, the predetermined granularity of splitting data transfer requests can be determined based on the bus width. This is because matching the splitting granularity to the hardware bus width ensures that each split data transfer request can be aligned with the bus access boundary, thereby achieving full utilization of the bus transmission bandwidth and allowing each element in the data segment along the second dimension to be mapped to different physical rows in the target storage area. In a non-limiting example, if the bus width is 16 bytes, the predetermined granularity in the first dimension is determined to be 16 bytes. If the elements of the tensor data have bf16 format, the predetermined granularity in the first dimension can include 8 elements in the first dimension.
[0057] In some examples, the splitting of the tensor data to be moved may also involve the selection of splitting dimensions, which can be determined based on the physical continuity of the tensor data in the source storage area. As mentioned earlier, when the tensor data to be moved is stored in the source storage area in the NDHWC format, the C dimension is located in the innermost position where the physical address changes the fastest, and elements that change along the C dimension occupy contiguous physical addresses in the source storage area; while the N, D, H, and W dimensions together constitute physically discontinuous dimensions.
[0058] When elements in dimension C are physically contiguous in the source storage area, data transfer requests can be split using dimension W as the splitting dimension. In such an example, under the NDHWC format, the entire segment of elements in dimension C corresponding to two adjacent W coordinates is physically stored contiguously. The data fragment indicated by each sub-request obtained along the W dimension (i.e., a predetermined granularity of consecutive elements expanded along the C dimension at the same W coordinate) can be read from the source storage area by the hardware circuit through a single continuous burst access, thereby aligning bus access boundaries and avoiding address switching overhead. In such an example, each completed sub-request processes a consecutive element at the W coordinate position that spans a predetermined granularity along the C dimension, and the corresponding W coordinate counter is incremented. When the W coordinate counter reaches the boundary value of the W dimension, it triggers the incrementing of the coordinate counters in higher dimensions (H dimension, and further D and N dimensions) and resets the W coordinate counter. This splitting method by the W dimension can adapt to the physical arrangement of "contiguous within C, contiguous within W" in the NDHWC format, making the splitting granularity match the physical continuity of the source storage area.
[0059] In other examples, the tensor data to be moved may not be physically contiguous along the C-axis in the source storage region. For example, the tensor data may be a slice of another tensor, or there may be spans between valid elements along the C-axis. In the case of discontinuity along the C-axis, elements spanning a predetermined granularity along the C-axis cannot be read from the source storage region in a single continuous burst access. In such examples, the splitting granularity can be further reduced by splitting the tensor data pixel by pixel. That is, each split data moving request indicates a data segment that corresponds to an element within the range to be moved along the C-axis at a single pixel coordinate. Pixel-by-pixel splitting allows the physical arrangement of the data segment corresponding to each sub-request in the source storage region to return to a continuous access or single-element access form supported by the hardware, avoiding a single sub-request covering access requirements that span non-contiguous address segments. Correspondingly, the calculation logic of address mapping in the target storage region does not need to be changed, and it is still performed according to the aforementioned offset superposition method based on the first and second indices. Only the number of elements processed by each sub-request is reduced compared to the aforementioned splitting along the W-axis.
[0060] The specific selection of the split dimension can be automatically determined by the data transport circuit based on the physical layout information of the tensor data carried in the data transport request (e.g., whether there is a C-dimensional span, the size and effective range of each dimension); in other examples, the split dimension can also be statically specified by the software layer through configuration instructions or by rewriting the configuration register of the data transport circuit to adapt to the inherent physical layout characteristics of the tensor data in a specific application scenario.
[0061] At S302, in response to the transpose instruction information indicating that a transpose transport of tensor data is to be performed, the transport parameters indicated by the data transport request are used to determine one or more source addresses of the tensor data in the source storage area and one or more destination addresses in the destination storage area.
[0062] In response to the transpose instruction information indicating transpose transport, it is necessary to determine one or more source addresses from which tensor data is loaded from the source storage region, and also to determine that after loading the tensor data, the transposed tensor data will be written to one or more target addresses in the target storage region.
[0063] Since a data transfer request is split into two or more split data transfer requests, the source addresses corresponding to each split data transfer request may be different. In some examples, the transfer parameters include the base address of the tensor data to be transferred. In such examples, determining one or more source addresses of the tensor data in the source storage region may include: determining multiple source addresses corresponding to multiple data fragments indicated by multiple split data transfer requests, based on the base address of the tensor data to be transferred and a predetermined granularity in a first dimension.
[0064] For example, the transport parameters may include the base address of the tensor data to be transported, corresponding to the source storage region. This base address may indicate the starting position of the tensor data within the source storage region. Since the granularity of each split data transport request is predetermined, the source address of the data segment indicated by each split data transport request can be determined by combining the base address of the tensor data in the source storage region. In some examples, the product of the index of each split data transport request and the predetermined granularity may correspond to the address offset of the data segment indicated by that split data transport request; therefore, the source address of the data segment indicated by each split data transport request may include the base address plus the address offset.
[0065] In some examples, one or both of the base address of the tensor data in the source storage region and the base address in the target storage region can be included as transport parameters in the data transport request. The data transport request is parsed from the data transport instructions issued at the software level by the instruction scheduling unit when the data transport request is generated and written into the configuration register of the data transport circuit. The hardware circuit reads these two base addresses from the configuration register each time the transport pipeline is started, and uses them as the starting point for calculating the source address and the target address, respectively.
[0066] In some examples, determining one or more target addresses for tensor data in the target storage region may include: determining one or more target addresses in the target storage region for one or more elements within the data fragment, based on the split data transfer request. In other words, the target address in the target storage region can be determined for each element in the split data transfer request.
[0067] In some examples, the transport parameters may include the base address of the tensor data in the target storage region. In such an example, determining one or more target addresses corresponding to one or more elements in the data fragment in the target storage region based on the split data transport request may include: for each element in the split data transport request: determining a first address offset across the first dimension after transposition of the element based on a first index of the element in a first dimension; determining a second address offset along the second dimension after transposition of the element based on a second index of the element in a second dimension; determining the target address corresponding to the element in the target storage region based on the first address offset across the first dimension after transposition of the element, the second address offset along the second dimension after transposition of the element, and the base address of the tensor data in the target storage region; wherein the first index of the element in the first dimension and the second index in the second dimension are determined based on the coordinates of the element in the first dimension and the second dimension, respectively.
[0068] To determine the target addresses corresponding to each element in a data segment, the approach to calculating these addresses can include: Each element in the data segment has indices in a first dimension and a second dimension. These indices indicate the row and column of the element in the source storage region, or in other words, the pixel and channel of the element in the source storage region. Since the transpose operation involves swapping the row and column positions of the element, in the target storage region, the first-dimensional index of the element instead indicates its row position, while the second-dimensional index indicates its column offset within that row. Therefore, based on the first and second indices of the element in the source storage region, combined with the spans corresponding to the rows and columns of the tensor data in the source storage region, the address offsets corresponding to the first and second indices after transpose can be determined. That is, the first address offset across the first dimension and the second address offset along the second dimension after transpose can be determined.
[0069] Here, the first address offset across the first dimension after the element is transposed can include the address offset resulting from the element crossing one or more first dimensions after transposition. In other words, the first address offset represents a span jump determined by the first dimension index, used to locate the element at the start of the row in the target storage area where its first dimension occupies. Figure 2A and Figure 2B In the example, if the first dimension is the channel dimension and the second dimension is the pixel dimension, the element p0c1 in the source storage area is transposed and its row and column positions are swapped. The first address offset of this element across the first dimension includes the address offset corresponding to the entire c0 channel before crossing c1.
[0070] In some examples, determining the first address offset across the first dimension after transposition of the element, based on the element's first index in the first dimension, may include: determining the total data width of the tensor data to be moved along the second dimension based on the data size of the tensor data to be moved; and determining the first address offset across the first dimension after transposition of the element based on the total data width of the tensor data to be moved along the second dimension and the element's first index in the first dimension. Here, the total data width of the tensor data to be moved along the second dimension may include the physical storage space occupied by a single complete first dimension when expanded along the second dimension in the transposed layout. Since the transposition operation reassembles elements that were originally discrete in the second dimension into elements that are continuously arranged in the second dimension, the total data width corresponding to the total number of elements expanded along the second dimension constitutes the logical row length of each first dimension in the target storage area. Therefore, by multiplying the element's first index in the first dimension by the total data width along the second dimension, the jump span from the starting position of the tensor data to the starting position corresponding to the element's first dimension can be calculated.
[0071] In some examples, if the first index of the element in the first dimension is represented as Col_cnt, the total number of elements along the second dimension is represented as row_total, and the bit width of each element is represented as data_format, then the total data width along the second dimension can be represented as row_total * data_format, and the first address offset can be represented as Col_cnt * row_total * data_format.
[0072] In some examples, the total number of elements along the second dimension is equal to the product of the dimensions along the N, D, H, and W dimensions, row_total = copy_n × copy_d × copy_h × copy_w.
[0073] In some examples, the computational logic along the total data width of the second dimension can be embedded in the arithmetic logic unit of the data transport circuit. Specifically, the hardware circuit can automatically perform multiplication or shift operations based on preset tensor shape parameters. For example, in a computing chip, if the total number of elements in the second dimension of the tensor data is known to always be fixed at 256 and the data format is int4, the multiplier array in the hardware circuit can be simplified to a fixed shift circuit, directly mapping the operation of multiplying 256 by 0.5 bytes to shift logic on the physical interconnect. This approach enables zero-latency address step generation, minimizing logic gate overhead.
[0074] Alternatively, the total number of elements along the second dimension can also be used as a programmable parameter, which can be dynamically set by the software layer through configuration instructions or by rewriting the configuration register of the data transport circuit to adapt to transport tasks under different tensor shapes; the hardware circuit reads this parameter from the configuration register and participates in the real-time calculation of the address step size each time it receives a data transport request.
[0075] Here, the second address offset along the second dimension after the element is transposed can include the address offset resulting from the element being expanded along multiple second dimensions after transposition. In other words, the second address offset indicates a jump determined by the second index of the second dimension, assuming the first dimension is the same, to position the element at the beginning of the column occupied by its second dimension in the target storage area. Figure 2A and Figure 2B In the example, if the first dimension is the channel dimension and the second dimension is the pixel dimension, after the element p0c1 in the source storage area is transposed and its row and column positions are swapped, the second address offset of the element along the second dimension can include the offset of the element within the c1 channel relative to the starting address of the c1 channel. That is, the second index p0 on the pixel dimension indicates the address offset corresponding to the column position with index p0 in the row of the c1 channel.
[0076] In some examples, determining the second address offset along the second dimension of an element after transposition, based on its second index in the second dimension, involves determining the second address offset along the second dimension based on the element's second index in the second dimension and the element's bit width. This is because, in the transposed layout, the second index in the second dimension has been transformed from indicating a row offset in the source storage region to a column offset within a row of the target storage region. Therefore, by multiplying this second index by the element's bit width, the address offset within the row corresponding to the element in the first dimension, caused by the second index in the second dimension, can be calculated, assuming the first dimension remains the same.
[0077] In some examples, if the second index of an element in the second dimension is represented as row_cnt, and the bit width of each element is represented as data_format, then the second address offset of the transposed element along the second dimension can include row_cnt*data_format.
[0078] In some examples, the calculation logic for multiplying by the element's bit width can be programmable or embedded in the arithmetic logic unit of the data transfer circuit. For instance, to address different precision requirements, the hardware circuit can automatically switch the parameters of the internal multiplier or shifter based on the data format information indicated in the data transfer request. For example, when processing the int4 data format, since its bit width is 0.5 bytes, the hardware circuit can quickly obtain the element's address offset within the row by directly shifting the second index one bit to the right (i.e., logically dividing by 2). When switching to the fp16 data format (bit width of 2 bytes), the hardware circuit performs the multiplication calculation by shifting the index one bit to the left. In some examples, for applications with a fixed element bit width over a long period (e.g., computing chips with full-process int4 or full-process fp16), the shift logic can also be embedded in the hardware logic of the data transfer circuit (e.g., by hard-wiring the second index directly to the predetermined shifter input); while for applications that need to support multiple precision switching, the shift logic can be dynamically switched according to the data format information indicated in the data transfer request or the bit width parameter stored in the configuration register of the data transfer circuit.
[0079] After the first address offset and the second address offset are determined, the target address of the element in the target storage area can be determined based on the base address of the tensor data in the target storage area indicated in the transport parameters.
[0080] The first index of the element in the first dimension and the second index in the second dimension are determined based on the coordinates of the element in the first dimension and the second dimension, respectively.
[0081] In some examples, the coordinates of each element in a data fragment may include coordinates in a first dimension and coordinates in a second dimension, which can be determined based on a state machine approach. For example, the data transport circuit may include a hardware state machine capable of automatically maintaining a set of multidimensional coordinate counters based on the dimensional information of the tensor data carried in each split data transport request (e.g., the dimensions of each of the N, D, H, W, and C dimensions) and the source address of each request. When the transport process starts, the hardware state machine generates the multidimensional coordinates of the currently processed element in the original tensor space in real time through nested loop counting logic. For example, after processing an element in a channel dimension, the channel coordinate counter is incremented; when the channel counter reaches the boundary value of that dimension, the pixel dimension coordinate counter is incremented and the channel counter is reset. In some examples, the boundary values of the multidimensional coordinate counter can be dynamically loaded by the data transport circuit based on the dimension information carried in the data transport request at the start of each transport. In other examples, for application scenarios where the tensor shape has been determined in the chip design stage, the boundary values can also be hardwired into the hardware state machine, thereby further reducing the overhead of the control logic without sacrificing versatility.
[0082] After determining the coordinates of an element in the first and second dimensions, its first and second indices in the first and second dimensions, respectively, can be determined. Specifically, the element's coordinates reflect its absolute position in the tensor space of the source storage region, while the indices determined based on these coordinates reflect its relative logical position within the current data fragment. For example, the coordinates generated by the state machine are based on a global count of all tensor data in the source storage region, but the actual split data transfer request only includes data fragments within the tensor data. To calculate the transposed target address, the global coordinates can be converted into indices relative to the starting element of the current data fragment. For example, the first index is determined by the difference between the coordinates in the first dimension and the coordinates of the starting element of the data fragment in the first dimension. This is because, when processing data fragment transfer, the global first-dimensional position in the tensor space can be mapped to a relative offset within the current data fragment, solving the problem of locating which first dimension to target. As another example, the second index can be determined based on the accumulation of coordinates in the second dimension. This is because the second index reflects the order in which the second dimension is arranged along spatial dimensions (e.g., NDHW and other dimensional information) within the current data fragment. The element coordinates of tensor data may include multi-dimensional coordinates such as NDHW in space. After being converted into a second index in the second dimension, they become an index indicating a one-dimensional linear position. This second index can indicate the one-dimensional linear position of an element in the data segment according to a predetermined spatial scan path (such as W followed by H). Determining the element's index in the data segment based on its coordinates simplifies complex dimensional permutations into linear operations based on the index and the tensor data span. This ensures that the relative span of the element with respect to the base address of the target storage region can be calculated regardless of where the data segment is located in the tensor space.
[0083] In some examples, for two adjacent elements in a split data transfer request, the target address of the latter element can be determined based on the target address of the former element offset from the total data width of the tensor data to be transferred along the second dimension. Here, two adjacent elements in a split data transfer request can include elements that are physically contiguous in the source storage region but logically belong to the same second dimension but different first dimensions. Figure 2AFor example, a split data transfer request may involve multiple channels in the pixel dimension at index p0. In this case, adjacent elements can refer to the elements p0c0 corresponding to channel c0 and p0c1 corresponding to channel c1, which are stored consecutively in the source storage area. As mentioned earlier, the total data width of the tensor data to be transferred along the second dimension can include the physical storage space occupied by a single complete first dimension when expanded along the second dimension in the transposed layout. In some examples, if the total number of elements along the second dimension is represented as row_total, and the bit width of each element is represented as data_format, then the total data width along the second dimension can be represented as row_total * data_format. The reason why the target address of the next element can be determined based on the target address of the previous element offset from the total data width of the tensor data to be transferred along the second dimension is that in the transposed mapping logic proposed in this disclosure, consecutive elements in the first dimension of the same second dimension row in the source storage area are mapped to different rows in the transposed target storage area. Figure 2A In the example, if the starting address of p0c0 to be written to the target storage area is denoted as Gsm_addr_tp, since the transposed layout requires elements of the same channel to be physically contiguous, all elements in the pixel dimension under c0 will occupy a contiguous address space, the length of which is the total data width along the second dimension. Therefore, the adjacent element p0c1 after p0c0 belongs to channel c1, and p0c1 should be the starting element of the c1 channel row in the transposed layout, stored at the starting position immediately following the c0 channel row. Therefore, the target address of p0c1 is at the starting address Gsm_addr_tp of p0c0, skipping an entire data storage block occupied by the c0 channel with a width of row_total*data_format. Through this equidistant offset calculation method based on the total data width along the second dimension, the hardware circuit can very quickly calculate the target address of all contiguous elements in the data segment after transposition, thereby achieving efficient discrete writing.
[0084] At S303, tensor data is loaded from the source storage region based on one or more source addresses.
[0085] Based on one or more source addresses corresponding to the data fragments indicated by the split data transfer request, as determined in step S302 above, tensor data can be loaded from the source storage area. In this step, the data transfer module can load the transposed tensor data from the source storage area.
[0086] At S304, tensor data is written to the target storage area based on one or more target addresses.
[0087] In some examples, writing tensor data to a target storage region may include writing one or more elements to one or more target addresses within the target storage region. That is, after splitting a data transfer request into two or more split data transfer requests, one or more elements from the data segments indicated in the split data transfer requests can be written to their corresponding target addresses in the target storage region. In some examples, for example, one can... Figure 2A Data segment p0c0 to p0c7 written Figure 2B The consecutive row of storage locations indicated by addr+1.
[0088] In some examples, at least a portion of the tensor data can be written to the same target address in the target storage region so that at least a portion of the data written to the same target address can be read in a single burst transfer.
[0089] Burst transfers here can include the continuous transfer of adjacent data segments with only a provided starting address, without needing to send address and control signals for each data element separately. The reason tensor data written to the target storage area can be read through a single burst transfer is that the data transfer method proposed in this disclosure recalculates the transposed storage location of each element, physically rearranging spatially discrete but logically highly related elements (e.g., elements of different pixels in the same channel) into a transposed layout and writing them into contiguous storage cells in the target storage area. Thus, at the physical level, these elements are aggregated within the same address range covered by the same bus width (i.e., the same target address or its contiguous ranges). Figure 2A In the non-transposed layout shown, if the execution unit needs to acquire elements c0p0, c0p1, up to c0p7 in consecutive pixels within the same channel, since these elements span multiple different pixel rows and their physical addresses are highly discrete, the computation unit must initiate multiple non-burst random accesses. Each read can only acquire a very small portion of the effective data, resulting in extremely low bandwidth utilization. However, using the method disclosed in this invention, this tensor data is remapped to the transposed layout during data transfer, such as... Figure 2B As shown, the aforementioned elements c0p0, c0p1, up to c0p7 were continuously written into the same physical memory space starting at address addr. When the computing unit needs this data again, it only needs to initiate a burst transfer request to the target memory area starting at address addr, and can linearly and continuously read all elements in the same pixel dimension within one or a very small number of clock cycles.
[0090] By proposing the data transfer scheme in this disclosure, the complex transpose addressing and data concatenation processes are completed in advance during the transfer stage. This allows the data transfer circuit to natively perceive the transpose mapping and achieve transfer-integration, enabling subsequent computing units to obtain the transposed tensor data using a burst-transfer-efficient memory access mode. Furthermore, this disclosure significantly reduces the overhead of on-chip cache resources. By using hardware circuitry to complete the transpose address mapping and reassembly during the tensor data flow to the target storage area, the intermediate storage stage is eliminated, and no additional on-chip storage resources are required.
[0091] Below, for reference Figure 4 The present disclosure describes a data transfer circuit that includes circuit modules for performing the aforementioned method.
[0092] The data transfer circuit 400 may include a data transfer request receiving circuit module 401, an address mapping circuit module 402, a source memory area loading circuit module 403, and a target memory area writing circuit module 404. In addition to these units, the data transfer circuit may also include other components; however, since these components are not relevant to the content of this disclosure embodiment, their illustrations and descriptions are omitted here. Furthermore, the specific details of the following operations performed by the data transfer circuit according to the present disclosure embodiment are consistent with those described above. Figure 3 The details described are the same, so repeated descriptions of the same details are omitted here to avoid repetition.
[0093] The data transfer request receiving circuit module 401 can be configured to receive a data transfer request for transferring tensor data to be transferred from a source storage area to a target storage area. The data transfer request includes transpose instruction information indicating whether to perform transpose transfer on the tensor data to be transferred.
[0094] Address mapping circuit module 402 can be configured to, in response to a transpose indication message indicating that a transpose transport of tensor data is to be performed, determine one or more source addresses of the tensor data in the source storage region and one or more target addresses in the target storage region based on the transport parameters indicated by the data transport request.
[0095] The source memory region loading circuit module 403 can be configured to load tensor data from a source memory region based on one or more source addresses.
[0096] The target memory area write circuit module 404 can be configured to write tensor data to a target memory area based on one or more target addresses.
[0097] In some examples, the data transfer request received in the data transfer request receiving circuit module 401 can be a hardware handshake signal between the data transfer circuit and the instruction scheduling unit. This data transfer request can be obtained by decoding the data transfer instructions issued at the software level. The address mapping circuit module 402 can integrate a multi-level adder array and multiplication logic unit, enabling real-time parsing of tensor dimension information in the transfer parameters. Specifically, this module can include a set of programmable step registers to store parameters such as indices calculated from the N, D, H, and W dimensions. When the transpose indication information is true, the address mapping circuit module 402 completes the coordinate transformation mapping in a single cycle through a hardware pipeline and calculates the offset of the target address using combinational logic circuits, thereby avoiding cycle loss caused by table lookup operations. The source memory area loading circuit module 403 can have multi-channel concurrent prefetching capability, enabling continuous burst transmission based on the source address sequence, and is equipped with a temporary buffer to align data of different bit widths. The target storage area write circuit module 404 can be equipped with mask control logic, which can control the storage location of each piece of data according to the calculated target address when writing tensor data to the target storage area, ensuring that the data is arranged in a predetermined pixel-first or channel-first order in the transposed layout.
[0098] In some examples, this data transport circuitry can be located within a tensor memory accelerator.
[0099] Integrating the aforementioned transposition and transport method into the hardware circuitry of the tensor memory accelerator allows complex address offset calculations to be performed by dedicated hardware logic within the TMA, without consuming computation cycles of a general-purpose processor, enabling the core to focus on matrix operations.
[0100] Below, for reference Figure 5 To illustrate a computing chip according to this disclosure.
[0101] For example, such as Figure 5 As shown, the computing chip 500 may include the aforementioned data transport circuit.
[0102] Integrating this data transport hardware circuit at the computing chip level firstly significantly improves the chip's efficiency in processing computational tasks. Because the transpose transport logic is embedded within the computing chip's data transport circuitry, the chip can simultaneously complete the complex transpose address mapping during the movement of tensor data from the source memory area to the target memory area with minimal hardware overhead, avoiding the instruction overhead caused by frequent data rearrangement within the computing chip. Secondly, this integrated design achieves a high degree of parallelism between memory access and computation. During computing chip operation, the data transport circuitry can operate independently of the computation unit, pre-preparing the transposed tensor data in a buffer. This means that when the computation unit needs tensor data, it can directly and efficiently read it without any preprocessing, thereby greatly shortening the waiting cycle of the computation pipeline and ensuring that the chip's computing power is fully utilized.
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0104] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on. The names of the units described in the embodiments of this disclosure do not, in some cases, constitute a limitation on the unit itself.
[0105] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0106] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
Claims
1. A data transfer method executed by a data transfer circuit, characterized in that, include: Receive a data transfer request, the data transfer request being used to transfer tensor data to be transferred from a source storage area to a target storage area, the data transfer request including transpose instruction information indicating whether to perform transpose transfer on the tensor data to be transferred; In response to the transpose instruction information indicating that transpose transport should be performed on the tensor data, based on the transport parameters indicated by the data transport request, one or more source addresses of the tensor data in the source storage area and one or more target addresses in the target storage area are determined; Based on the one or more source addresses, the tensor data is loaded from the source storage area; as well as Based on the one or more target addresses, the tensor data is written to the target storage area. The tensor data to be transferred includes a first dimension and a second dimension. In the source storage area, multiple elements of the tensor data to be transferred are stored consecutively according to the priority order of the first dimension, and in the target storage area, multiple elements of the tensor data are stored consecutively according to the priority order of the second dimension. The first dimension and the second dimension are different dimensions. The method further includes: splitting the second dimension of the tensor data to be transferred indicated by the data transfer request at least once based on a predetermined granularity in the first dimension to obtain two or more split data transfer requests, wherein the split data transfer request indicates a data segment of the tensor data to be transferred in the second dimension, the data segment consisting of one or more elements of the tensor data to be transferred, the number of elements corresponding to the predetermined granularity in the first dimension.
2. The method according to claim 1, characterized in that, The target storage area is a buffer for computational operations, and the transpose indication information is determined based on whether the tensor data participates in the computational operation in a transposed form.
3. The method according to claim 1 or 2, characterized in that, The predetermined granularity in the first dimension is determined based on the bus width.
4. The method according to claim 1 or 2, characterized in that, The transport parameters include the base address of the tensor data to be transported, and Determining one or more source addresses of the tensor data in the source storage area includes: Based on the base address of the tensor data to be transported and the predetermined granularity in the first dimension, multiple source addresses are determined for each of the multiple data fragments indicated by the multiple split data transport requests.
5. The method according to claim 4, characterized in that, Determining one or more target addresses of the tensor data in the target storage area includes: Based on the split data transfer request, determine one or more target addresses corresponding to one or more elements in the data fragment in the target storage area.
6. The method according to claim 5, characterized in that, Writing the tensor data to the target storage area includes: Write the one or more elements to the one or more target addresses in the target storage area.
7. The method according to claim 1 or 2, characterized in that, The first dimension is the channel dimension, and the second dimension is the pixel dimension.
8. The method according to claim 1 or 2, characterized in that, The tensor data is tensor data stored in the source storage area in the NDHWC format, where N dimension is the batch dimension, D dimension is the depth dimension, H dimension is the height dimension, W dimension is the width dimension, and C dimension is the channel dimension. The first dimension is the C dimension; The second dimension is a dimension formed by expanding one or more of the N dimension, the D dimension, the H dimension, and the W dimension.
9. The method according to claim 1 or 2, characterized in that, Also includes: At least a portion of the tensor data is written to the same target address in the target storage region, such that the at least a portion of the data written to the same target address can be read in a single burst transfer.
10. The method according to claim 5, characterized in that, The transport parameters include the base address of the tensor data in the target storage area, and wherein, based on the split data transport request, determining one or more target addresses corresponding to the one or more elements in the data fragment in the target storage area includes: For each element in the split data transfer request: Based on the element’s first index in the first dimension, determine the first address offset of the element across the first dimension after transposition. Based on the second index of the element in the second dimension, determine the second address offset of the element after transposition along the second dimension; Based on the first address offset across the first dimension after the element is transposed, the second address offset along the second dimension after the element is transposed, and the base address of the tensor data in the target storage area, the target address corresponding to the element in the target storage area is determined. The first index of the element in the first dimension and the second index in the second dimension are determined based on the coordinates of the element in the first dimension and the second dimension, respectively.
11. The method according to claim 10, characterized in that, Based on the element's first index in the first dimension, the first address offset across the first dimension after transposition of the element is determined as follows: Based on the data size of the tensor data to be transferred, determine the total data width of the tensor data to be transferred along the second dimension; Based on the total data width of the tensor data to be transported along the second dimension and the first index of the element in the first dimension, the first address offset of the element across the first dimension after transposition is determined.
12. The method according to claim 10 or 11, characterized in that, Based on the element's second index in the second dimension, the second address offset along the second dimension after transposition of the element is determined as follows: Based on the second index of the element in the second dimension and the bit width of the element, determine the second address offset of the element after transposition along the second dimension.
13. The method according to claim 10 or 11, characterized in that, For two adjacent elements in the split data transfer request, the target address of the latter element is determined based on the target address of the former element offset from the total data width of the tensor data to be transferred along the second dimension.
14. A data transfer circuit, characterized in that, Includes a circuit module for performing the method according to any one of claims 1-13.
15. The data transfer circuit according to claim 14, characterized in that, The data transfer circuit is located in the tensor memory accelerator.
16. A computing chip, characterized in that, The computing chip includes the data transport circuit according to claim 14 or 15.
Citation Information
Patent Citations
Tensor data transformation method, tensor processing unit, processor, system on chip, and computing device
CN121858501A