A data transmission method, apparatus and electronic device
Patent Information
- Application Number
- CN202611164122.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-03
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]有鉴于此,本申请提供了一种数据传输方法、装置及电子设备,主要目的在于解决现有多维DMA频繁重置维度计数器打断总线连续突发传输、带宽利用率低且硬件动态功耗高的技术问题
[0013] By employing the above technical solutions, this application provides a data transmission method, apparatus, and electronic device. Compared with existing technologies, this application can divide continuous data arrangement dimension groups and independent data arrangement dimensions based on the size and step information of multiple data arrangement dimensions contained in the descriptor information of the target tensor. It then merges the addressing paths corresponding to the continuous data arrangement dimension groups and generates a target addressing path by combining the addressing paths corresponding to the independent data arrangement dimensions. Tensor data addressing and transmission are then completed using this target addressing path. By merging continuous dimension addressing paths, the frequency of dimension counter resets is reduced, ensuring continuous burst transmission of the bus, improving bandwidth utilization, and simultaneously reducing the dynamic power consumption caused by redundant counter flipping.
Smart Images

Figure CN122654053A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chip technology, and in particular to a data transmission method, apparatus and electronic device. Background Technology
[0002] AI chips require massive tensor data transfer during neural network inference and training phases. This transfer task is handled by a multi-dimensional direct memory access (DMA) engine. The DMA engine calculates memory access addresses dimension-by-dimensionally using multi-layered nested loop logic based on the size and step size information of multiple data arrangement dimensions carried by the target descriptor information, thus completing the data read and write operations of the target tensor.
[0003] Existing multidimensional DMA directly performs nested loop addressing based on all the original data arrangement dimensions of the target tensor. After traversing each data arrangement dimension, the corresponding dimension counter needs to be reset. However, frequent counter resets interrupt continuous bus burst transfers, splitting long bursts into multiple short bursts, significantly reducing bus bandwidth utilization. At the same time, the continuous toggling of multi-level counters and frequent switching of loop states generate high dynamic power consumption, which restricts the overall data transmission performance. Summary of the Invention
[0004] In view of this, this application provides a data transmission method, apparatus and electronic device, the main purpose of which is to solve the technical problems of existing multidimensional DMAs that frequently reset the dimension counter, interrupting continuous burst transmission of the bus, resulting in low bandwidth utilization and high dynamic power consumption of hardware.
[0005] In a first aspect, this application provides a data transmission method, including: Obtain the descriptor information of the target tensor. The descriptor information includes the size information and step size information of the multiple data arrangement dimensions corresponding to the target tensor.
[0006] Based on the size and step information of multiple data arrangement dimensions, adjacent data arrangement dimensions that meet the continuity condition between dimensions are divided into continuous data arrangement dimension groups, while data arrangement dimensions that do not meet the continuity condition between dimensions are divided into independent data arrangement dimensions.
[0007] The addressing paths corresponding to each data arrangement dimension within the continuous data arrangement dimension group are merged. Based on the merged addressing path and the addressing paths corresponding to the independent data arrangement dimensions, the target addressing path of the target tensor is generated. The target addressing path is used to address and transmit the data within the target tensor.
[0008] Secondly, this application provides a data transmission apparatus, comprising: The acquisition module is configured to acquire the descriptor information of the target tensor. The descriptor information includes the size information and step size information of multiple data arrangement dimensions corresponding to the target tensor.
[0009] The judgment module is configured to classify adjacent data arrangement dimensions that meet the continuity condition between dimensions into continuous data arrangement dimension groups, and data arrangement dimensions that do not meet the continuity condition between dimensions into independent data arrangement dimensions, based on the size information and step size information of multiple data arrangement dimensions.
[0010] The generation module is configured to merge the addressing paths corresponding to each data arrangement dimension within the continuous data arrangement dimension group, and generate the target addressing path of the target tensor based on the merged addressing path and the addressing path corresponding to the independent data arrangement dimension. The target addressing path is used to address and transmit the data within the target tensor.
[0011] Thirdly, this application provides a computer storage medium on which a computer program is stored, and when the computer program is executed by a processor, it implements the data transmission method of the first aspect described above.
[0012] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the data transmission method described in the first aspect.
[0013] By employing the above technical solutions, this application provides a data transmission method, apparatus, and electronic device. Compared with existing technologies, this application can divide continuous data arrangement dimension groups and independent data arrangement dimensions based on the size and step information of multiple data arrangement dimensions contained in the descriptor information of the target tensor. It then merges the addressing paths corresponding to the continuous data arrangement dimension groups and generates a target addressing path by combining the addressing paths corresponding to the independent data arrangement dimensions. Tensor data addressing and transmission are then completed using this target addressing path. By merging continuous dimension addressing paths, the frequency of dimension counter resets is reduced, ensuring continuous burst transmission of the bus, improving bandwidth utilization, and simultaneously reducing the dynamic power consumption caused by redundant counter flipping. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A schematic flowchart of a data transmission method provided in an embodiment of this application is shown; Figure 2 A schematic diagram of a DMA channel architecture applicable to a data transmission method provided in an embodiment of this application is shown; Figure 3 This paper shows a schematic diagram of the structure of a continuous detection unit provided in an embodiment of the present application; Figure 4 This illustration shows a schematic diagram of the structure of a data transmission device according to an embodiment of this application; Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0017] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0018] This application's solution is applicable to AI chip neural network training and inference scenarios, covering dynamic shape operations such as variable-length sequence processing, token distribution in Mixture of Experts (MoE) models, key-value cache read / write, and sparse tensor operations. When the operation runs, the software sends descriptor information carrying the target tensor to the multidimensional DMA. The descriptor information records the size and step size information of each data arrangement dimension. After reading the descriptor information, the DMA engine traverses all data arrangement dimensions layer by layer, calculates the memory address by accumulating the step size dimension by layer using multi-layered nested loop counters, and initiates segmented bus burst transfers to complete the data transfer between on-chip cache and off-chip memory. In such scenarios, the dimensional continuity of tensors cannot be predicted during the compilation stage, making it a common scenario for performance defects in traditional multidimensional DMA.
[0019] In the aforementioned Dynamic Shape tensor transport scenario, traditional multidimensional DMA always builds multi-layer cyclic addressing logic based on the entire original data arrangement dimensions of the tensor. Each time a dimension traversal is completed, the corresponding dimension counter must be reset. Frequent counter resets forcibly terminate the current continuous bus burst transfer, splitting the complete long burst into multiple short bursts. Each time the bus initiates a burst, it needs to transmit address and control handshake signals. In short burst scenarios, the address handshake overhead increases significantly, compressing the effective data transmission rate and directly reducing bus bandwidth utilization. Simultaneously, the multi-layer dimension counters need to continuously accumulate and reset in a loop. Each counter value flip triggers a timing jump in the corresponding register, generating a large amount of switching current. Combined with the power consumption from combinational logic switching due to frequent state machine transitions, the overall dynamic power consumption of the hardware increases significantly, severely limiting the efficiency of tensor data transmission.
[0020] Currently, there are various targeted optimization solutions in the industry, but all of them have obvious shortcomings: One approach is the static descriptor fixation scheme, which only adapts to tensors of fixed size. In dynamic scenarios, the CPU needs to rewrite the descriptor, introducing a significant amount of computational latency. The second is the instruction set extension scheme, which requires modification of the binary interface of software and hardware applications, resulting in high software adaptation costs. Third is the compile-time dimension folding scheme, which can only optimize tensors with known shapes at compile time, and fails in dynamic shape scenarios; Fourth is the software runtime rescheduling scheme, which relies on CPU intervention and incurs significant scheduling overhead. Fifth, the DMA bandwidth preprocessing optimization scheme only optimizes the data transformation process and does not solve the problem of dimensional addressing redundancy. Sixth is the matrix rearrangement DMA scheme, which requires synchronous modification of tensor layout metadata and lacks software transparency.
[0021] None of the above solutions can simplify the multidimensional addressing path in real time through DMA hardware without modifying the software descriptor or requiring additional CPU involvement, and they are unable to simultaneously improve bus bandwidth utilization and hardware power consumption defects.
[0022] To address the technical issues of frequent dimension counter resets interrupting continuous burst transmissions on the bus, low bandwidth utilization, and high dynamic power consumption in existing multidimensional DMA, this application reads the data arrangement dimension size information and step size information contained in the descriptor information during the DMA hardware operation phase, automatically determines the continuity of dimensions, divides continuous dimension groups into independent dimensions, merges the addressing paths corresponding to continuous dimensions to generate a simplified target addressing path, and completes tensor data addressing and transmission based on this path, thereby reducing the number of counter resets and redundant flip operations at the hardware addressing link level.
[0023] This embodiment provides a data transmission method applied to a multidimensional DMA engine (hereinafter referred to as DMA engine). Figure 1 The diagram shows a flowchart of a data transmission method, which includes: S101. Obtain the descriptor information of the target tensor.
[0024] The target tensor is, for example, the multidimensional tensor of the Dynamic Shape service mentioned above. The descriptor information of the target tensor is the configuration information configured by the upper-layer inference / training software and sent to the multidimensional DMA engine. The descriptor information contains the size information and stride information of multiple data arrangement dimensions corresponding to the target tensor. The size information (also known as Size) represents the total number of data elements in the corresponding data arrangement dimension, and the stride information (also known as Stride) represents the address offset difference between two adjacent data elements in memory.
[0025] For example, the descriptor information may also carry parameters such as tensor base address, data bit width, read / write transfer identifier, on-chip cache interaction identifier, and transfer length threshold. This application does not modify the descriptor information structure or the distribution logic; upper-layer software does not need to be adapted or modified, maintaining compatibility with the application's binary interface.
[0026] In one implementation, the upper-layer software writes the descriptor into the command first-in-first-out queue (Cmd FIFO) at the front end of the DMA channel. The DMA engine reads the descriptor corresponding to the target tensor sequentially from the Cmd FIFO, and then the command parser built into the engine completes the parsing of basic parameters such as size information, step size information, and base address to obtain the descriptor information.
[0027] S102. Based on the size information and step size information of multiple data arrangement dimensions, adjacent data arrangement dimensions that meet the continuity condition between dimensions are divided into continuous data arrangement dimension groups, and data arrangement dimensions that do not meet the continuity condition between dimensions are divided into independent data arrangement dimensions.
[0028] The continuity condition between dimensions refers to the fact that the data element addresses of two adjacent data arrangement dimensions are arranged consecutively in memory space, and there are no gaps in memory addresses after the dimensions are superimposed. For example, if the total address span generated by a complete traversal of the inner data arrangement dimension between adjacent data arrangement dimensions is equal to the step size corresponding to a single element of the outer data arrangement dimension, then no invalid blank memory intervals will be generated when the two dimensions are superimposed, thus satisfying the continuity condition between dimensions.
[0029] In one implementation, based on the size information and step size information corresponding to each data arrangement dimension obtained by parsing, the continuity determination of each group of adjacent data arrangement dimensions is carried out in sequence; multiple data arrangement dimensions that are adjacent to each other and all satisfy the continuity condition between dimensions are aggregated to form a continuous data arrangement dimension group; when any adjacent dimension does not satisfy the continuity condition, the group boundary is broken and the corresponding data arrangement dimension is divided into an independent data arrangement dimension.
[0030] For example, the target tensor is a 3D feature map tensor in a neural network inference scenario, with the original dimensions arranged from the inside out as width W (D1), height H (D2), and channel C (D3). Adjacent dimensions are traversed pairwise in the order from the inside out for continuity checks. If the inner layer D1 and the middle layer D2 satisfy the continuity condition, they are first grouped into the same group to be merged, and then the continuity of this group as a whole is checked against the outer layer D3. If the middle layer D2 and the outer layer D3 do not satisfy the continuity condition, the group boundary is broken. After the determination, D1 and D2 are aggregated into a continuous data arrangement dimension group, and the outer layer D3 is divided into an independent data arrangement dimension.
[0031] In another example, the target tensor is a four-dimensional WCHN format feature tensor for dynamic shape business. The original dimensions are arranged from the inside out as follows: width dimension W (D1), height dimension H (D2), channel dimension C (D3), and batch dimension N (D4). Adjacent dimensions are checked pairwise from the inside out. D1 and D2 satisfy the inter-dimensional continuity condition, as do D3 and D4. D2 and D3 do not satisfy the inter-dimensional continuity condition, and the dimension grouping boundary is broken between D2 and D3. After the determination, D1 and D2 are aggregated into the first group of continuous data arrangement dimensions, and D3 and D4 are aggregated into the second group of continuous data arrangement dimensions, with the two groups of dimensions separated from each other.
[0032] S103. Merge the addressing paths corresponding to each data arrangement dimension in the continuous data arrangement dimension group. Based on the merged addressing path and the addressing path corresponding to the independent data arrangement dimension, generate the target addressing path of the target tensor. The target addressing path is used to address and transmit the data in the target tensor.
[0033] Here, the addressing path refers to the single-level loop addressing logic corresponding to a single data arrangement dimension. When performing merging processing on a group of consecutive data arrangement dimensions, the multiple single-level loop addressing logics corresponding to each dimension within the group are merged and converted into a single single-level linear addressing path; while independent data arrangement dimensions do not undergo merging operations and directly retain their own corresponding single-level loop addressing paths. Several merged single-level linear addressing paths and several single-level loop addressing paths corresponding to independent dimensions are sequentially concatenated according to the original dimensions of the tensor from the inside out, forming a target addressing path with a multi-level structure.
[0034] For example, a WHC format three-dimensional feature tensor in a neural network inference scenario is selected as an example. The original dimension arrangement of this tensor from the inside out is width dimension D1 (W), height dimension D2 (H), and channel dimension D3 (C). After the continuity determination in S102, D1 and D2 satisfy the continuity condition between dimensions and are divided into a continuous data arrangement dimension group. D2 and D3 do not satisfy the continuity condition, and D3 is treated as an independent data arrangement dimension. D1, D2, and D3 each correspond to an independent single-layer loop addressing logic. The three single-layer loop addressing logics are nested inside and outside, forming a complete three-layer nested native addressing link. During the merging process, the size and step information of D1 and D2 are read, the overall continuous address length of the dimension group is calculated, the two independent loop counters of D1 and D2 are removed, and the original two independent single-level loop addressing logics of D1 and D2 are merged into a single single-level linear addressing path. The entire continuous data of D1×D2 is traversed in one go using the linear address offset. Then, the original single-level loop addressing path of the independent dimension D3 is retained. Following the original dimension order from inside to outside, the merged single-level linear addressing path is placed in the inner layer, and the single-level loop addressing path of D3 is placed in the outer layer. The two single-level addressing logics are sequentially concatenated to form a hybrid hierarchical addressing logic, which is the target addressing path of the target tensor. The number of loop counters in the target addressing path obtained in this way is less than the number of loop counters corresponding to the original three-level nested addressing link of the target tensor. This reduces the hardware overhead caused by counter toggling and resetting, while maintaining bus-length burst transmission and improving data transmission bandwidth utilization.
[0035] Following the target addressing path of the target tensor, the DMA engine completes address calculation and bus data transfer based on this hybrid-level target addressing path: First, it executes the inner single-layer linear addressing path, initiating an Advanced eXtensible Interface (AXI) long burst transfer based on continuous linear addresses, completely outputting all data of D1×D2 under a single channel. The address is continuous and uninterrupted throughout the transfer, and the bus burst is not interrupted by the inner dimension counter reset. After a single inner linear data transfer is completed, only the single-layer loop counter corresponding to the outer D3 is updated, and the memory starting address of the next set of width and height planes is calculated by adding the channel dimension step size. The inner linear long burst transfer is then started again. The above process is repeated cyclically until the D3 counter has traversed all channel dimensions. The entire transfer process only requires step updates to the outer single counter, avoiding the problem of frequent multi-layer counter clearing and resetting in the native three-layer nested addressing. The bus can continuously maintain a full-length continuous burst, reducing bus handshake overhead and further reducing hardware dynamic power consumption.
[0036] like Figure 2The diagram illustrates a DMA channel architecture applicable to a data transfer method. This DMA channel architecture includes a task scheduling layer, a semantic reinterpretation layer, an instruction generation layer, and a bus backend transmission layer. The composition and corresponding functions of each layer are as follows: The task scheduling layer includes a multi-task queue, a task arbitration unit, and a task parameter parsing unit. The multi-task queue is used to cache multiple tensor transmission tasks carrying descriptor information to be executed. The task arbitration unit selects and outputs a single task from the multi-task queue according to a preset priority rule. The task parameter parsing unit reads the descriptor information corresponding to the target tensor in the current task, extracts parameters such as the size information, step size information, and tensor base address of each data arrangement dimension from the descriptor information, and then outputs the parsed dimension parameters to the semantic reinterpretation layer.
[0037] The semantic reinterpretation layer comprises a continuous detection unit, a recursive folding unit, and a decision encoder. The continuous detection unit receives the parsed dimension parameters and checks each pair of adjacent data arrangement dimensions from the inside out to determine if they satisfy the continuity condition, dividing the target tensor into groups of continuously arranged data dimensions and independent data arrangement dimensions. The recursive folding unit fuses and transforms multiple original single-layer cyclic addressing logics within each group of continuously arranged data dimensions, folding and merging multiple single-layer cyclic addressing paths into a single single-layer linear addressing path. Then, it splices the linear addressing path and the independent-dimensional single-layer addressing path according to the original internal and external order of the tensor's dimensions to generate a hybrid-level target addressing path. The decision encoder encodes and marks the target addressing path type, outputting standardized addressing description information.
[0038] The instruction generation layer includes a hardware execution instruction generator; it receives addressing description information output by the decision encoder, translates the target addressing path of the hybrid layer into hardware-recognizable low-level control instructions, completes the decoupling and adaptation between the upper-level semantic logic and the lower-level execution hardware, and distributes the control instructions down to the bus back-end transmission layer.
[0039] The bus backend transport layer includes a unified high-level extensible interface (AXI) / network on chip (NoC) backend. Internally, the unified AXI / NOC backend integrates three types of finite state machines: Linear Finite State Machine (LinearFSM), Collapsed Finite State Machine (CollapsedFSM), and Multidimensional Finite State Machine (MultidimFSM), as well as an AXI initiator, an outstanding tracker, a write-merge buffer, and a data path. This layer receives control commands from the upper layer, matches the corresponding finite state machine according to the type of the target addressing path, performs address step calculation, batch data caching, continuous bus burst scheduling, and continuously initiates uninterrupted AXI4 / AXI5 / CHI / NoC bus transmissions. It relies on the underlying physical bus interface to complete the addressing, reading, writing, and transport of all data in the target tensor.
[0040] Compared with existing technologies, this application can divide continuous data arrangement dimension groups and independent data arrangement dimensions based on the size and step information of multiple data arrangement dimensions contained in the descriptor information. It then merges the addressing paths corresponding to the continuous data arrangement dimension groups and combines them with the addressing paths corresponding to the independent data arrangement dimensions to generate a target addressing path. Tensor data addressing and transmission are then completed using this target addressing path. By merging continuous dimension addressing paths, the frequency of dimension counter resets is reduced, ensuring continuous burst transmission on the bus, improving bandwidth utilization, and simultaneously reducing the dynamic power consumption caused by redundant counter flipping.
[0041] The data transmission method provided in the embodiments of this application will now be described in detail.
[0042] Optionally, this application also provides an embodiment for determining whether the data arrangement dimensions satisfy the inter-dimensional continuity condition. In this embodiment, based on the size information and step size information of multiple data arrangement dimensions, the layout attributes of the target tensor are determined, and a target determination logic that adapts to the layout attributes is selected. According to the target determination logic, the inter-dimensional continuity condition between two adjacent data arrangement dimensions is determined sequentially according to the original arrangement order of the multiple data arrangement dimensions. Based on the determination results, the multiple data arrangement dimensions are divided into continuous data arrangement dimension groups and independent data arrangement dimensions.
[0043] The layout attributes include either packed canonical row-major or non-canonical row-major. Packed canonical row-major means the tensor data has no padding gaps, the step size of adjacent inner dimensions is strictly equal to the inner dimension size, and the memory address increases linearly with the dimension index. In actual AI accelerator operation, most target tensor descriptor information belongs to this layout attribute. Non-canonical row-major means the tensor has memory padding segments, or the step size of adjacent inner dimensions does not match the inner dimension size, and there are gaps in the memory addresses corresponding to adjacent dimension indices. This layout attribute is compatible with special tensor scenarios with padding and custom step sizes, and includes complete judgment logic with offset correction.
[0044] In one implementation, all data arrangement dimensions are traversed and verified one by one, comparing the correspondence between the outer dimension step size and the inner dimension size, and the product of the inner dimension step size. If all adjacent dimensions satisfy the correspondence, the current target tensor is determined to be a tightly packed canonical row-major tensor. If any set of adjacent dimensions does not satisfy the correspondence, or if the tensor is configured with a padding offset parameter, the current target tensor is determined to be a non-canonical row-major tensor.
[0045] For example, the above traversal verification operation is performed by the continuity detection unit inside the semantic reinterpretation layer, and the corresponding continuity determination calculation formula is: stride [i+1] == size [i] × stride [i]; In the formula, i represents the index of the current inner layer data arrangement dimension; size[i] represents the size information corresponding to the i-th inner layer data arrangement dimension; stride[i] represents the stride information corresponding to the i-th inner layer data arrangement dimension; stride[i+1] represents the stride information corresponding to the outer layer data arrangement dimension immediately adjacent to the i-th layer. If the stride value of the outer layer dimension is equal to the product of the inner layer dimension size and the inner layer dimension stride, it means that the two adjacent dimensions are consecutive and without gaps in physical memory address.
[0046] Optionally, in addition to relying on dimension and stride information to determine layout attributes, various configuration channels can be used to pre-specify the target tensor layout attributes. For example, three optional configuration channels can be included: operator-bound static layout parameters, task queue-attached configuration tags, and on-chip system global mode configuration switches.
[0047] During actual hardware operation, configuration information is read layer by layer in a hierarchical retrieval manner, and valid configurations are matched from top to bottom to determine the final effective layout attributes. An example of the retrieval and execution process is as follows: The layout flag field reserved in the header of the descriptor information is read first. If the flag has a valid configuration, the layout attribute is determined directly based on the flag. This field is only read and will not modify or destroy the architecture state visible to the software side, so it has the highest priority.
[0048] If there is no valid layout flag in the descriptor header, continue reading the channel-level register LAYOUT_MODE_CSR[1:0] dedicated to the current DMA channel. Select the corresponding working mode through the binary code stored in the register. The code 00 corresponds to adaptive mode, 01 corresponds to force-packed mode, and 10 corresponds to force-strided mode, which is also known as non-standard row main order.
[0049] If the current channel register is not forcibly configured, the global policy register shared by multiple DMA channels is read, and the pre-set unified layout judgment policy in the register is used as the global default rule. The hardware resource overhead caused by the independent register of each DMA channel is reduced by relying on the global shared configuration.
[0050] If the above descriptor flags, channel registers, and global registers have no effective layout constraints, the above-mentioned discrimination logic of traversing dimensions, comparing size and step size values can be executed according to the adaptive default strategy to distinguish between tightly packed normal line order and non-normal line order.
[0051] After obtaining the layout attributes of the target tensor, the continuous detection unit selects the target judgment logic that matches the layout attributes, and then performs continuity judgment on each pair of adjacent data layout dimensions according to the original arrangement order of the multiple data layout dimensions of the target tensor. Combining the judgment results of each pair of adjacent dimensions, all data layout dimensions are divided into several groups of continuous data layout dimensions and independent data layout dimensions.
[0052] Optionally, if the layout attribute of the target tensor is a tightly packed, normal row order, a comparison-based decision logic is selected as the target decision logic to adapt to the layout attribute; if the layout attribute of the target tensor is a non-normal row order, a multiplication-based decision logic is selected as the target decision logic to adapt to the layout attribute.
[0053] The comparison and determination logic is used to adapt target tensors with a tightly packed, row-major layout. The hardware path is configured with a numerical comparator and a first-level AND gate. The theoretical outer-layer step size is calculated by performing a shift operation on the dimensions and step size of the inner data arrangement dimension. This theoretical outer-layer step size is then fed into the numerical comparator and compared with the actual outer-layer dimension step size. The numerical comparator outputs an equality flag signal and a tensor no-padding offset flag signal, both of which are fed into the first-level AND gate. Only when both signals are valid simultaneously can the memory addresses of two adjacent data arrangement dimensions be determined to be contiguous. This logic has a shorter overall computational chain, lower hardware timing overhead, and is suitable for most tensor transport scenarios with a tightly packed, row-major layout.
[0054] The multiplication operation logic is used to adapt target tensors with non-normal row-major layout attributes. The hardware path is configured with a multiplier, a numerical comparator, and a multi-level AND gate. Step size compensation is achieved by adjusting the tensor filling offset parameter. Then, the multiplier calculates the product of the inner data arrangement dimension and the inner step size, and the product result is sent to the numerical comparator for comparison with the corrected outer dimension step size. The equality flag, offset correction completion flag, and filling parameter validity flag output by the numerical comparator are all connected to the multi-level AND gate. When all constraint signals are valid simultaneously, it is determined that the memory addresses of two adjacent sets of data arrangement dimensions are continuous, which is compatible with various special tensors with memory filling and custom step sizes.
[0055] In one implementation, the continuous detection unit relies on multi-level configuration to perform step-by-step retrieval to select and switch the judgment logic. The retrieval priority from high to low is as follows: descriptor header layout flag bit, channel-level LAYOUT_MODE_CSR [1:0] register, and global policy register. If there is a valid forced configuration at any of the aforementioned levels, the corresponding layout attribute is directly matched and the matching judgment logic is selected. If none of the above channels have layout restrictions, the adaptive default strategy is enabled. By traversing the dimensions and step size values of each dimension, the target tensor is automatically distinguished as belonging to the tightly packed normal row main order or the non-normal row main order. The comparison judgment logic or the multiplication operation judgment logic is adaptively switched to perform dimensional continuity verification on the target tensor.
[0056] Optionally, if the target determination logic is a comparison determination logic, a shift operation is performed based on the size information and step size information of the inner data arrangement dimension; if the shift operation result is equal to the step size information of the outer data arrangement dimension, it is determined that the two adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition.
[0057] In one implementation, the continuous detection unit reads the inner layer data arrangement dimension size information and the inner layer step size information, performs a shift operation, compares the operation result with the outer layer data arrangement dimension step size information, and if the two values are equal, it can be determined that the adjacent two sets of data arrangement dimensions meet the continuity condition between dimensions.
[0058] For example, for a tightly packed main-order tensor with no padding offset and regular arrangement, a complete verification process is performed: The continuous detection unit first reads the dimension parameters in the descriptor and extracts the inner layer data arrangement dimension size information and inner layer step size information; determines the left shift number based on the inner layer size, and performs a left shift operation on the inner layer step size to obtain the theoretical outer layer step size; compares the theoretical outer layer step size with the outer layer step size, and generates an equality flag signal if the values are equal; then the equality flag and the no-padding offset flag are jointly judged, and when both flags are valid at the same time, it is determined that the dimensions of two adjacent sets of data arrangement satisfy the continuity condition between dimensions.
[0059] Optionally, if the target determination logic is a multiplication operation determination logic, a product operation is performed based on the size information and step size information of the inner data arrangement dimension; if the product operation result is equal to the step size information of the outer data arrangement dimension, it is determined that the two adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition.
[0060] In one implementation, the continuous detection unit reads the inner layer data arrangement dimension size information and the inner layer step size information, performs a product operation, compares the operation result with the outer layer data arrangement dimension step size information, and if the two values are equal, it can be determined that the adjacent two sets of data arrangement dimensions meet the continuity condition between dimensions.
[0061] For example, for non-standard row main order tensors with padding offsets and custom strides, a complete verification process is executed: The continuous detection unit first reads the padding offset parameters in the descriptor and compensates and corrects the original step size information of the inner and outer data arrangement dimensions respectively; it then multiplies the corrected inner size and inner step size to obtain the theoretical outer step size; it compares the theoretical outer step size with the corrected outer step size, and generates an equality flag signal if the values are equal; then the equality flag, the offset correction completion flag, and the padding parameter validity flag are jointly judged. When all three flags are valid, it is determined that the adjacent two sets of data arrangement dimensions meet the continuity condition between dimensions.
[0062] like Figure 3 The diagram shows a structural schematic of a continuous detection unit. The continuous detection unit includes two detection paths: a general path DP-G and a preferred path DP-S. The two paths operate in parallel and independently to generate their respective continuity determination vectors. The two determination vectors are synchronously input into a multiplexer. The multiplexer dynamically selects one of the paths based on the path selection signal corresponding to the layout attribute and outputs the continuity determination result.
[0063] The general path DP-G deploys N-1 parallel decision branches. Each branch carries an unsigned multiplier, a numerical comparator, and two sets of non-zero dimension guard gate circuits. The multiplier bit width matches the sum of the maximum dimension size bit width and the maximum step size bit width. The N-1 branches output single-dimensional continuous decision results, and all single-dimensional continuous decision results are combined to form a continuous decision vector.
[0064] Next, the continuity decision vector is fed into the AND tree circuit and OR tree circuit for parallel computation. The AND tree circuit performs a global AND operation on the continuity decision vector, generating a full_contiguous identifier. This identifier is set to valid only when all decision results within the vector are valid; otherwise, it is set to invalid. An invalid identifier indicates that not all data arrangement dimensions of the target tensor satisfy the inter-dimensional continuity condition. The OR tree circuit performs a global OR operation on the continuity decision vector, generating a local_contiguous identifier. This identifier is set to valid only when any one of the decision results within the vector is valid. A valid identifier indicates that the target tensor has at least two adjacent data arrangement dimensions that satisfy the inter-dimensional continuity condition. When all decision results in the vector are invalid, this identifier is set to invalid, indicating that none of the target tensor's adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition.
[0065] For example, when the total number of dimensions N of the target tensor is less than or equal to 8, all combinational logic in the general path completes the operation and converges in a single cycle. When the total number of dimensions N of the target tensor is greater than 8, the circuit can add a pipeline register at the multiplier output node to alleviate timing pressure. The hardware architecture of this path is compatible with all tensor arrangement forms and adapts to various operating conditions such as tightly packed standard row-major order, straddle view, slice view, and layout with padding.
[0066] The preferred path DP-S deploys N-1 parallel decision branches. Each branch only carries a numerical comparator and a single-dimensional non-zero guard gate circuit; no multiplication unit is integrated within the branch. The N-1 branches output single-dimensional continuous decision results, and all results are combined to form a continuous decision vector. The continuous decision vector is synchronously fed into the AND tree circuit and OR tree circuit for parallel computation. It can be seen that the operation logic and execution flow of the AND tree circuit and OR tree circuit at the back end of the preferred path are completely consistent with the corresponding circuits at the back end of the general path. The only difference between the two types of paths is the decision hardware of the front-end branches; the generation rules for the global continuousness identifier are the same.
[0067] The optimized path has lower overall hardware footprint, lower dynamic power consumption, and lower signal transmission latency than the general path. This path can be deployed independently in hardware-space-sensitive scenarios.
[0068] The continuity detection unit outputs a continuity determination vector `contiguous[N-2:0]`, a complete continuity identifier `full_contiguous`, and a local continuity identifier `any_contiguous`. These three sets of signals are directly sent to the recursive folding unit. Based on the three sets of continuity determination results, the recursive folding unit divides the target tensor into continuous data arrangement dimension groups, thereby determining the simplified target addressing path and reducing the computational overhead caused by multi-dimensional independent addressing.
[0069] The continuity detection unit also synchronously outputs a safety flag (safety_flag), which is connected to the decision encoder. When the safety flag is at a valid level, the decision encoder configures the hardware execution mode to multidimensional tensor processing mode. The safety flag is a protection control signal output by the continuity detection unit. When the tensor dimensions are irregularly arranged or do not meet the simplified folding processing conditions, the continuity detection unit sets the safety flag to a valid level. After recognizing a valid safety flag, the decision encoder abandons the low-latency single-dimensional fast transport mode and initiates the complete multidimensional tensor processing flow, avoiding hardware malfunctions such as memory address errors and data read / write misalignments.
[0070] Optionally, this application also provides an embodiment for generating a target addressing path for a target tensor. The addressing path for the target tensor is generated based on whether all data arrangement dimensions in the target tensor satisfy the inter-dimensional continuity condition, partially satisfy the inter-dimensional continuity condition, or neither satisfies the inter-dimensional continuity condition. Specifically, if all data arrangement dimensions in the target tensor satisfy the inter-dimensional continuity condition, that is, the target tensor has only one set of continuous data arrangement dimensions. If some data arrangement dimensions partially satisfy the inter-dimensional continuity condition, that is, the target tensor has both a set of continuous data arrangement dimensions and independent data arrangement dimensions. If neither data arrangement dimensions satisfy the inter-dimensional continuity condition, that is, the target tensor has only independent data arrangement dimensions.
[0071] When the target tensor has both continuous data arrangement dimension groups and independent data arrangement dimensions, the multi-level cyclic addressing logic corresponding to the continuous data arrangement dimension group is converted into a single-level linear addressing logic. The single-level linear addressing logic corresponding to the continuous data arrangement dimension group and the multi-level cyclic addressing logic corresponding to the independent data arrangement dimension are logically concatenated to obtain a hybrid hierarchical addressing logic. The hybrid hierarchical addressing logic is determined as the target addressing path of the target tensor.
[0072] In one implementation, the recursive folding unit receives the continuity determination vector, the local continuity identifier "any_contiguous", and the full continuity identifier "full_contiguous" output by the continuity detection unit. When the local continuity identifier is valid and the full continuity identifier is invalid, it determines that the current tensor contains both continuous data arrangement dimension groups and independent data arrangement dimensions. The recursive folding unit first traverses each group of continuous data arrangement dimension groups, extracts the size information and step size information of each dimension within the group, calculates the complete continuous address interval corresponding to the dimension group, removes the multi-level independent loop counters within the group, and merges the multiple single-level loop addressing logic segments within the group into a single single-level linear addressing logic; it completely preserves the original multi-level loop addressing logic of each independent data arrangement dimension; then, according to the original arrangement order of the target tensor data arrangement dimensions from the inside to the outside, it sequentially concatenates the single-level linear addressing logic and the independent dimension multi-level loop addressing logic to generate a hybrid hierarchical addressing logic, and outputs the hybrid hierarchical addressing logic to the decision encoder to complete the addressing path encoding mark, which serves as the target addressing path of the target tensor.
[0073] When the target tensor has only independent data arrangement dimensions, the multi-layer circular addressing logic corresponding to each independent data arrangement dimension is determined as the target addressing path of the target tensor.
[0074] In one implementation, when the local continuation flag `any_contiguous` output by the continuous detection unit is invalid, it is determined that none of the adjacent data arrangement dimensions of the target tensor satisfy the inter-dimensional continuation condition, and there is no available group of continuous data arrangement dimensions. The recursive folding unit skips the dimension folding and merging process, fully retains the original single-level circular addressing logic corresponding to each independent data arrangement dimension, and nests and combines all independent dimension circular addressing logics according to the original dimensions of the tensor from the inside out, forming a complete multi-level nested addressing logic. This multi-level nested addressing logic is output as the target addressing path of the target tensor to the decision encoder to generate standardized addressing description information. At the same time, the continuous detection unit outputs a valid safety flag `safety_flag`. After recognizing the safety flag, the decision encoder switches to the complete multi-dimensional tensor processing mode to avoid hardware operation abnormalities such as memory address errors and data read / write misalignments caused by directly folding and simplifying irregular tensors.
[0075] When the target tensor has only one set of continuous data arrangement dimensions, the multi-level circular addressing logic corresponding to the continuous data arrangement dimension set is converted into a single-level linear addressing logic, and the single-level linear addressing logic is determined as the target addressing path of the target tensor.
[0076] In one implementation, when the full_contiguous flag output by the continuous detection unit is valid, it is determined that all adjacent data arrangement dimensions of the target tensor satisfy the continuity condition between dimensions, and the entire tensor is aggregated into a single continuous data arrangement dimension group, with no independent data arrangement dimensions. The recursive folding unit extracts the size parameters and step size parameters of all dimensions within this unique continuous dimension group, calculates the total length of the tensor's global continuous address, removes all multi-level loop counters within the group, and converts the multi-level nested loop addressing logic within the group into a single single-level linear addressing logic, directly using this single-level linear addressing logic as the target addressing path of the target tensor.
[0077] After receiving the target addressing path output by the recursive folding unit, the decision encoder encodes and marks the type of the target addressing path, and outputs standardized addressing description information to the instruction generation layer. Based on the standardized addressing description information, the instruction generation layer translates the hybrid hierarchical addressing logic, pure multi-layer nested addressing logic, or pure single-layer linear addressing logic into hardware-recognizable low-level control instructions, and distributes the control instructions down to the bus back-end transmission layer. The bus back-end transmission layer matches the corresponding type of finite state machine according to the control instructions to complete address step calculation, batch data buffering, and continuous bus burst scheduling: if the target addressing path is single-layer linear addressing logic, the linear finite state machine LinearFSM is enabled; if the target addressing path is hybrid hierarchical addressing logic, the collapsed finite state machine CollapsedFSM and the multi-dimensional nested finite state machine MultidimFSM are enabled in tandem; if the target addressing path is pure multi-layer nested addressing logic, only the multi-dimensional nested finite state machine MultidimFSM is enabled, and continuous uninterrupted AXI / NoC bus transmission is continuously initiated based on the corresponding finite state machine to complete the addressing, reading, writing, and transport of all data of the target tensor.
[0078] For example, after obtaining the target addressing path, the number of folding dimensions M, the full_contiguous flag, and the any_contiguous flag from the recursive folding unit, the decision encoder completes the addressing path type encoding mark and synchronously outputs the execution mode encoding signal exec_mode. exec_mode is a multi-bit encoded flag used to distinguish the addressing execution path corresponding to the current tensor. The hardware divides the working conditions into three categories based on the number of folding dimensions M and matches a dedicated exec_mode encoding, execution mode, and address generation path for each working condition. The first type of working condition: the entire data arrangement dimension of the target tensor is continuous, M is 1, exec_mode is configured as Linear mode, the LinearFSM is scheduled to work, and continuous burst linear address copying is completed by relying on a single counter. The second type of working condition: The target tensor has a locally continuous data arrangement dimension, M satisfies 2≤M<N, where N is the original total number of dimensions of the target tensor, exec_mode is configured to Collapsed mode, CollapsedFSM is scheduled to work, dimension folding is completed based on the M-layer counter and the overall step size after folding is recalculated, and hybrid hierarchical addressing transmission is performed. The third type of working condition: The target tensor does not have a continuous data arrangement dimension, the value of M is equal to the total number of tensor dimensions N, the exec_mode is configured as Multidim mode, MultidimFSM is scheduled to work, and the complete N-layer raw counter is called to perform multi-layer nested addressing and transmission.
[0079] Furthermore, in this embodiment, the three types of address generation finite state machines—LinearFSM, CollapsedFSM, and MultidimFSM—share the same bus back-end processing hardware structure. The address signals output by LinearFSM, CollapsedFSM, and MultidimFSM are uniformly connected to the same set of AXIISsuer, OutstandingTracker, Write-MergeBuffer, and Datapath. The three types of finite state machines are not allocated dedicated bus ports or independent multi-task queues.
[0080] Furthermore, the number of counters configured internally in CollapsedFSM is equal to the number of folding dimensions M obtained during runtime parsing, and the value of M ranges from 2≤M≤N-1; for the NM groups of idle counters that do not participate in folding, the clock of the corresponding circuit is turned off through the clock-gatingenable signal.
[0081] Compared with existing technologies, this application can divide continuous data arrangement dimension groups and independent data arrangement dimensions based on the size and step information of multiple data arrangement dimensions contained in the descriptor information. It then merges the addressing paths corresponding to the continuous data arrangement dimension groups and combines them with the addressing paths corresponding to the independent data arrangement dimensions to generate a target addressing path. Tensor data addressing and transmission are then completed using this target addressing path. By merging continuous dimension addressing paths, the frequency of dimension counter resets is reduced, ensuring continuous burst transmission on the bus, improving bandwidth utilization, and simultaneously reducing the dynamic power consumption caused by redundant counter flipping.
[0082] Furthermore, as Figures 1 to 3 To provide a specific implementation of the method shown, this embodiment offers a data transmission device, such as... Figure 4 The diagram shows a data transmission device 400, which includes an acquisition module 410, a judgment module 420, and a generation module 430; wherein: The acquisition module 410 is configured to acquire the descriptor information of the target tensor. The descriptor information includes the size information and step size information of multiple data arrangement dimensions corresponding to the target tensor.
[0083] The judgment module 420 is configured to divide adjacent data arrangement dimensions that meet the continuity condition between dimensions into continuous data arrangement dimension groups, and data arrangement dimensions that do not meet the continuity condition between dimensions into independent data arrangement dimensions, based on the size information and step size information of multiple data arrangement dimensions.
[0084] The generation module 430 is configured to merge the addressing paths corresponding to each data arrangement dimension in the continuous data arrangement dimension group, and generate the target addressing path of the target tensor based on the merged addressing path and the addressing path corresponding to the independent data arrangement dimension. The target addressing path is used to address and transmit the data in the target tensor.
[0085] Optionally, the judgment module 420 is configured to determine the layout attributes of the target tensor based on the size information and step information of multiple data arrangement dimensions. The layout attributes include one of tightly packed canonical row master order and non-canonical row master order. Target determination logic for selecting appropriate layout attributes.
[0086] Based on the target determination logic, the system sequentially determines whether two adjacent data arrangement dimensions satisfy the continuity condition between dimensions according to the original arrangement order of multiple data arrangement dimensions. Based on the determination results, the multiple data arrangement dimensions are divided into continuous data arrangement dimension groups and independent data arrangement dimensions.
[0087] Optionally, the judgment module 420 is configured to select the comparison judgment logic as the target judgment logic for adapting the layout attribute when the layout attribute of the target tensor is a tightly packed canonical row main order.
[0088] When the layout attribute of the target tensor is non-normal row-major order, the multiplication operation judgment logic is selected as the target judgment logic to adapt the layout attribute.
[0089] Optionally, the judgment module 420 is configured to perform a shift operation based on the size information and step size information of the inner data arrangement dimension when the target judgment logic is a comparison judgment logic.
[0090] If the result of the shift operation is equal to the step size information of the outer data arrangement dimension, it is determined that the two adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition.
[0091] In this context, for two adjacent data arrangement dimensions, the inner data arrangement dimension precedes the outer data arrangement dimension in the original arrangement order.
[0092] Optionally, the judgment module 420 is configured to perform a product operation based on the size information and step size information of the inner data arrangement dimension when the target judgment logic is a multiplication operation judgment logic.
[0093] If the product result is equal to the step size information of the outer data arrangement dimension, it is determined that the two adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition.
[0094] Optionally, the generation module 430 is configured to convert the multi-level circular addressing logic corresponding to the continuous data arrangement dimension group into a single-level linear addressing logic when the target tensor has both continuous data arrangement dimension groups and independent data arrangement dimensions.
[0095] The single-level linear addressing logic corresponding to the continuous data arrangement dimension group and the multi-level cyclic addressing logic corresponding to the independent data arrangement dimension are logically concatenated to obtain the hybrid hierarchical addressing logic. The hybrid hierarchical addressing logic is determined as the target addressing path of the target tensor.
[0096] Optionally, the generation module 430 is configured to determine the multi-level circular addressing logic corresponding to each independent data arrangement dimension as the target addressing path of the target tensor when the target tensor has only independent data arrangement dimensions.
[0097] Optionally, the generation module 430 is configured to convert the multi-level cyclic addressing logic corresponding to the continuous data arrangement dimension group into a single-level linear addressing logic when the target tensor has only one set of continuous data arrangement dimension groups, and to determine the single-level linear addressing logic as the target addressing path of the target tensor.
[0098] It should be noted that other corresponding descriptions of the functional units involved in the data transmission device provided in this embodiment can be found in [reference needed]. Figures 1 to 3 The corresponding descriptions in [the document] will not be repeated here.
[0099] Based on the above, Figures 1 to 3 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figures 1 to 3 The method shown.
[0100] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0101] like Figure 5The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 501; and, A memory 502 is communicatively connected to at least one of the processors 501; wherein, The memory 502 stores instructions that can be executed by at least one of the processors to enable the at least one of the processors to perform the data transfer method as described above.
[0102] Figure 5 Take a processor 501 as an example.
[0103] The electronic device may also include an input device 503 and a display device 504.
[0104] The processor 501, memory 502, input device 503, and display device 504 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0105] The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the data transmission method in the embodiments of this application, for example, Figures 1 to 3 The method flow is shown. The processor 501 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 502, thereby implementing the data transmission method in the above embodiments.
[0106] Memory 502 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created according to the use of the data transmission method, etc. Furthermore, memory 502 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 502 may optionally include memory remotely located relative to processor 501, and these remote memories may be connected to the apparatus performing the data transmission method via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0107] Input device 503 can receive user clicks and generate signal inputs related to user settings and function control for data transmission methods. Display device 504 may include display screens or other display devices.
[0108] When one or more modules are stored in the memory 502, and are run by one or more processors 501, the data transmission method in any of the above method embodiments is executed.
[0109] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0110] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0111] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware platforms, or it can be implemented in hardware. By applying the solution of this embodiment, compared with the prior art, this application can divide continuous data arrangement dimension groups and independent data arrangement dimensions based on the size information and step information of multiple data arrangement dimensions contained in the descriptor information, and merge the addressing paths corresponding to the continuous data arrangement dimension groups. It then combines the addressing paths corresponding to the independent data arrangement dimensions to generate a target addressing path, and completes tensor data addressing and transmission based on this target addressing path. By merging continuous dimension addressing paths, the frequency of dimension counter resets is reduced, ensuring continuous burst transmission of the bus, improving bandwidth utilization, and simultaneously reducing the dynamic power consumption of hardware caused by redundant counter flipping.
[0113] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0114] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A data transmission method, characterized in that, include: Obtain the descriptor information of the target tensor, wherein the descriptor information includes the size information and step information of multiple data arrangement dimensions corresponding to the target tensor; Based on the size information and step size information of the multiple data arrangement dimensions, adjacent data arrangement dimensions that meet the continuity condition between dimensions are divided into continuous data arrangement dimension groups, and data arrangement dimensions that do not meet the continuity condition between dimensions are divided into independent data arrangement dimensions. The addressing paths corresponding to each data arrangement dimension within the continuous data arrangement dimension group are merged. Based on the merged addressing path and the addressing path corresponding to the independent data arrangement dimension, a target addressing path for the target tensor is generated. The target addressing path is used to address and transmit the data within the target tensor.
2. The method according to claim 1, characterized in that, Based on the size and step information of the multiple data arrangement dimensions, adjacent data arrangement dimensions that satisfy the continuity condition between dimensions are divided into continuous data arrangement dimension groups, and data arrangement dimensions that do not satisfy the continuity condition between dimensions are divided into independent data arrangement dimensions, including: Based on the size information and step information of the multiple data arrangement dimensions, the layout attributes of the target tensor are determined, and the layout attributes include one of tightly packed canonical row master order and non-canonical row master order; Select the target determination logic that adapts to the layout attributes; Based on the target determination logic, the original arrangement order of the multiple data arrangement dimensions is used to determine whether two adjacent data arrangement dimensions meet the continuity condition between dimensions. Based on the determination result, the multiple data arrangement dimensions are divided into continuous data arrangement dimension groups and independent data arrangement dimensions.
3. The method according to claim 2, characterized in that, The logic for selecting a target that matches the layout attributes includes: When the layout attribute of the target tensor is the main order of the tightly packed specification row, the comparison and determination logic is selected as the target determination logic to adapt to the layout attribute. When the layout attribute of the target tensor is the non-normal row major order, the multiplication operation judgment logic is selected as the target judgment logic to adapt to the layout attribute.
4. The method according to claim 3, characterized in that, The step of determining whether two adjacent data arrangement dimensions satisfy the continuity condition between dimensions according to the original arrangement order of the multiple data arrangement dimensions based on the target determination logic includes: When the target determination logic is a comparison determination logic, a shift operation is performed based on the size information and step size information of the inner layer data arrangement dimension. If the result of the shift operation is equal to the step size information of the outer data arrangement dimension, it is determined that the two adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition. Specifically, for two adjacent data arrangement dimensions, the inner data arrangement dimension precedes the outer data arrangement dimension in the original arrangement order.
5. The method according to claim 3, characterized in that, The step of determining whether two adjacent data arrangement dimensions satisfy the continuity condition between dimensions according to the original arrangement order of the multiple data arrangement dimensions based on the target determination logic includes: When the target determination logic is the multiplication operation determination logic, the product operation is performed according to the size information and step information of the inner data arrangement dimension. If the product result is equal to the step size information of the outer data arrangement dimension, it is determined that the two adjacent data arrangement dimensions satisfy the inter-dimensional continuity condition.
6. The method according to any one of claims 1 to 5, characterized in that, The process of merging the addressing paths corresponding to each data arrangement dimension within the continuous data arrangement dimension group, and generating the target addressing path of the target tensor based on the merged addressing path and the addressing path corresponding to the independent data arrangement dimension, includes: When the target tensor has both continuous data arrangement dimension groups and independent data arrangement dimensions, the multi-level circular addressing logic corresponding to the continuous data arrangement dimension groups is converted into a single-level linear addressing logic. The single-level linear addressing logic corresponding to the continuous data arrangement dimension group is logically concatenated with the multi-level cyclic addressing logic corresponding to the independent data arrangement dimension to obtain a hybrid hierarchical addressing logic, and the hybrid hierarchical addressing logic is determined as the target addressing path of the target tensor.
7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: When the target tensor has only the independent data arrangement dimension, the multi-layer circular addressing logic corresponding to each independent data arrangement dimension is determined as the target addressing path of the target tensor.
8. The method according to any one of claims 1 to 5, characterized in that, The method further includes: When the target tensor has only one set of continuous data arrangement dimensions, the multi-level cyclic addressing logic corresponding to the continuous data arrangement dimension set is converted into a single-level linear addressing logic, and the single-level linear addressing logic is determined as the target addressing path of the target tensor.
9. A data transmission device, characterized in that, include: The acquisition module is configured to acquire descriptor information of a target tensor, wherein the descriptor information includes size information and step information of multiple data arrangement dimensions corresponding to the target tensor; The judgment module is configured to, based on the size information and step size information of the multiple data arrangement dimensions, divide adjacent data arrangement dimensions that satisfy the continuity condition between dimensions into continuous data arrangement dimension groups, and divide data arrangement dimensions that do not satisfy the continuity condition between dimensions into independent data arrangement dimensions. The generation module is configured to merge the addressing paths corresponding to each data arrangement dimension in the continuous data arrangement dimension group, and generate the target addressing path of the target tensor based on the merged addressing path and the addressing path corresponding to the independent data arrangement dimension. The target addressing path is used to address and transmit the data in the target tensor.
10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data transmission method according to any one of claims 1 to 8.