Physical address generation method and device, processor and chip
By calculating the coordinate increment and cross-dimensional enable signals of multi-dimensional data during the initial operation period, the problem of large hardware overhead is solved, the hardware resource utilization and timing performance are optimized, and data handling efficiency is improved.
Patent Information
- Application Number
- CN202510495926.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, hardware overhead is relatively high when generating physical addresses, especially in the process of high bandwidth memory data handling, hardware resource occupancy is high and timing complexity increases, making it difficult to meet the needs of efficient data handling.
By calculating the coordinate increments of multidimensional data in advance and generating cross-dimensional enable signals during the initial operation cycle, the use of adder resources is reduced and hardware overhead and timing performance is optimized.
It realizes the generation of physical addresses under smaller hardware overhead, improves data handling efficiency, and reduces hardware resource occupation and timing complexity.
Smart Images

Figure CN120407454A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence chips, and particularly to a method, apparatus, processor, and chip for generating physical addresses. Background Art
[0002] In the field of general computing, especially in applications involving large-scale data processing (such as artificial intelligence inference, image processing, etc.), the computing unit needs to frequently load or store data from High Bandwidth Memory (HBM), and these data are usually organized in a multi-dimensional form. In related technologies, the generation of physical addresses is achieved by gradually accumulating the coordinate values of each dimension. However, the hardware overhead required for such methods in hardware implementation needs to be reduced. Summary of the Invention
[0003] This application provides a method, apparatus, processor, and chip for generating physical addresses, which solves the technical problem that the hardware overhead required in hardware implementation in related technologies needs to be reduced, and achieves the technical effect of reducing hardware overhead.
[0004] To achieve the above objective, the main technical solutions adopted in this application include:
[0005] In a first aspect, an embodiment of this application provides a method for generating a physical address, which is applied to a hardware accelerator. The method includes:
[0006] Determine the starting coordinates and coordinate increment data of the multi-dimensional data in each dimension;
[0007] In an initial operation cycle, add the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinate of each dimension and the initial cross-dimension enable signal in the initial operation cycle; where the initial cross-dimension enable signal is used to indicate whether the initial updated coordinate causes a cross-dimension in the adjacent higher dimension of each dimension;
[0008] Based on the initial updated coordinate and the initial cross-dimension enable signal, perform coordinate calculation and address conversion to obtain the physical address of the multi-dimensional data.
[0009] Optionally, the performing coordinate calculation and address conversion based on the initial updated coordinate and the initial cross-dimension enable signal to obtain the physical address of the multi-dimensional data includes: 2]
[0010] In the initial operation cycle, perform coordinate calculation based on the initial cross-dimension enable signal and the initial updated coordinate to obtain the initial coordinate calculation result of each dimension in the initial operation cycle;
[0011] In the first operation cycle after the initial operation cycle, coordinate calculation is performed on the initial coordinate calculation result based on the initial cross-dimension enabling signal to obtain the first coordinate calculation result of the first operation cycle;
[0012] Coordinate calculation and address conversion are performed according to the initial coordinate calculation result and the first coordinate calculation result to obtain the physical address of the multi-dimensional data.
[0013] Optionally, a first cross-dimension enabling signal is generated in the first operation cycle; the performing coordinate calculation and address conversion according to the initial coordinate calculation result and the first coordinate calculation result to obtain the physical address of the multi-dimensional data includes:
[0014] In the second operation cycle after the first operation cycle, coordinate calculation is performed on the first coordinate calculation result based on the first cross-dimension enabling signal to obtain the second coordinate calculation result of the second operation cycle;
[0015] In the second operation cycle, summarization and address conversion are performed based on the initial coordinate calculation result, the first coordinate calculation result, and the second coordinate calculation result to obtain the physical address of the multi-dimensional data.
[0016] Optionally, the size of the multi-dimensional data in each dimension is denoted as the dimension size; the performing coordinate calculation based on the initial cross-dimension enabling signal and the initial updated coordinates to obtain the initial coordinate calculation result of each dimension in the initial operation cycle includes:
[0017] If the initial cross-dimension enabling signal indicates that the initial updated coordinates cause cross-dimension, the calculation result of subtracting the dimension size on the corresponding dimension from the initial updated coordinates is used to obtain the initial coordinate calculation result;
[0018] If the initial cross-dimension enabling signal indicates that the initial updated coordinates do not cause cross-dimension, the initial updated coordinates are used as the initial coordinate calculation result.
[0019] [[ID=2,4]]Optionally, the size of the multi-dimensional data in each dimension is denoted as the dimension size; the adding the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinates of each dimension and the initial cross-dimension enabling signal in the initial operation cycle includes:
[0020] The coordinate increment data of each dimension is added to the starting coordinate of the corresponding dimension to obtain the initial updated coordinates of each dimension;
[0021] Performing cross - dimensional judgment on the initial updated coordinates of each dimension and the dimension size on the corresponding dimension in the adjacent higher dimension to generate the initial cross - dimensional enable signal.
[0022] Optionally, the lowest dimension among all dimensions of the multi - dimensional data is denoted as the first lowest dimension; the step size of the multi - dimensional data on each dimension is denoted as the dimension step size; the coordinate calculation of the initial coordinate calculation result based on the initial cross - dimensional enable signal to obtain the first coordinate calculation result of the first operation period includes:
[0023] Determining whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross - dimensional enable signal to obtain the first coordinate calculation result of the first target dimension within the first operation period; wherein, the first target dimension is any dimension other than the first lowest dimension among all dimensions.
[0024] Optionally, the size of the multi - dimensional data on each dimension is denoted as the dimension size; the determining whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross - dimensional enable signal to obtain the first coordinate calculation result of the first target dimension within the first operation period includes:
[0025] If the initial cross - dimensional enable signal indicates that the initial updated coordinates cause cross - dimension, accumulating the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension to obtain the first updated coordinate of the first target dimension;
[0026] Generating the first cross - dimensional enable signal within the first operation period based on the first updated coordinate to determine the first coordinate calculation result; wherein, the first cross - dimensional enable signal is used to indicate whether the first updated coordinate causes cross - dimension in the adjacent higher dimension of the first target dimension.
[0027] Optionally, the generating the first cross - dimensional enable signal within the first operation period based on the first updated coordinate includes:
[0028] Performing cross - dimensional judgment on the first updated coordinate of the first target dimension and the dimension size on the first target dimension in the adjacent higher dimension of the first target dimension to generate the first cross - dimensional enable signal.
[0029] Optionally, the lowest dimension among multiple first target dimensions is denoted as the second lowest dimension; the method further includes:
[0030] In a second operation period after the first operation period, determine whether to accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal, so as to obtain the second coordinate calculation result of the second target dimension in the second operation period; wherein, the second target dimension is any dimension among the multiple first target dimensions except the second lowest dimension.
[0031] Optionally, the determining whether to accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal to obtain the second coordinate calculation result of the second target dimension in the second operation period includes:
[0032] If the first cross-dimension enable signal indicates that the first updated coordinate causes cross-dimension, accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension to obtain the second updated coordinate of the second target dimension;
[0033] Generate a second cross-dimension enable signal in the second operation period based on the second updated coordinate to determine the second coordinate calculation result; wherein, the second cross-dimension enable signal is used to indicate whether the second updated coordinate causes cross-dimension on the adjacent higher dimension of the second target dimension.
[0034] Optionally, the generating the second cross-dimension enable signal in the second operation period based on the second updated coordinate includes:
[0035] Perform cross-dimension judgment on the adjacent higher dimension of the second target dimension by using the second updated coordinate of the second target dimension and the dimension size on the second target dimension to generate the second cross-dimension enable signal.
[0036] In a second aspect, an embodiment of the present application provides a physical address generation device applied to a hardware accelerator, and the device includes:
[0037] A coordinate data determination module, configured to determine the starting coordinates and coordinate increment data of multi-dimensional data in each dimension;
[0038] A coordinate signal determination module, configured to, in an initial operation period, add the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinate of each dimension and the initial cross-dimension enable signal in the initial operation period; wherein, the initial cross-dimension enable signal is used to indicate whether the initial updated coordinate causes cross-dimension on the adjacent higher dimension of each dimension;
[0039] A physical address generation module, configured to perform coordinate calculation and address conversion based on the initially updated coordinates and the initial cross-dimension enable signal, so as to obtain the physical address of the multi-dimensional data.
[0040] In a third aspect, an embodiment of the present application provides a processor, which includes a logic circuit and a power supply circuit. The power supply circuit is configured to supply power to the logic circuit, and the logic circuit is configured to execute the steps of any one of the above methods.
[0041] In a fourth aspect, an embodiment of the present application provides a chip, which includes the processor as described above.
[0042] In the embodiment of the present application, first, the starting coordinates and coordinate increment data of multi-dimensional data in each dimension are determined; secondly, in the initial operation period, the coordinate increment data of each dimension is advanced to the starting coordinates of the corresponding dimension to obtain the initially updated coordinates of each dimension, and an initial cross-dimension enable signal in the initial operation period is further generated. Finally, coordinate calculation and address conversion are performed based on the initially updated coordinates and the initial cross-dimension enable signal to obtain the physical address of the multi-dimensional data. Compared with the calculation method of cumulative addition by dimension in the related art, the adder resources used in the calculation process in this embodiment are less, so that the area occupied by the adder can be reduced, the physical address can be generated with a small hardware overhead, and the logic resource overhead in the initial operation period can be further reduced. Description of the Drawings
[0043] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1a It is a schematic diagram of the scenario where the tensor core processes the same picture in the embodiment of the present application;
[0045] Figure 1b It is a flow timing diagram of the tensor core executing instructions in the related art;
[0046] Figure 1c It is a partial structural schematic diagram of the artificial intelligence processor in the scenario example of the present application;
[0047] Figure 1d It is a flow timing diagram of the tensor core executing instructions in the scenario example of the present application;
[0048] Figure 1e It is a flow schematic diagram of the method for generating the physical address in the embodiment of the present application;
[0049] Figure 2 It is a schematic flowchart of the method for generating a physical address in an embodiment of the present application;
[0050] Figure 3 It is a schematic flowchart of the method for generating a physical address in an embodiment of the present application;
[0051] Figure 4 It is a schematic framework diagram of the device for generating a physical address in an embodiment of the present application. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0053] With the rapid development of fields such as artificial intelligence and high-performance computing, the requirement for the data transfer efficiency of high-bandwidth memory (HBM) is increasing day by day. In the data transfer operations in related technologies, on the one hand, if coordinate auto-increment is not implemented, after executing one instruction, it is necessary to reconfigure the data block coordinate information of the next instruction, which will, to a certain extent, result in waste of instruction resources and increase software complexity. On the other hand, the coordinate auto-increment method in related technologies adopts a serialized coordinate update scheme, which relies on multi-level adders, increases the hardware overhead to a certain extent, leads to high occupation of hardware resources, increases the timing complexity and area, and is difficult to meet the demand for efficient data transfer.
[0054] Please refer to Figure 1a , Figure 1aThe figure shows a scene in which four tensor cores (tcore) are used to collaboratively process the same image. The four tensor cores are tcore0, tcore1, tcore2, and tcore3. Tcore0, tcore1, tcore2, and tcore3 process different areas of the image in parallel. Tensor cores are dedicated computing cores designed for tensor operations (such as matrix multiplication, convolution, etc.) and are commonly found in GPUs (graphics processing units) or AI accelerators. Tensor cores can perform parallel calculations on multi-dimensional data (such as tensors in NDHW format) and support mixed precision calculations (such as FP16 / FP32), significantly improving AI training / inference performance. The following is an explanation using the NDHW data format as an example. The N dimension represents the batch size, that is, the number of data samples captured in one training session, the D dimension represents the depth, the H dimension represents the height of the input data, and the W dimension represents the width of the input data.
[0055] Please continue reading Figure 1a , each tcore independently processes a data block in the picture (such as tcore0 processes the first blue data block 102, tcore1 processes the first green data block 104, etc.), and improves throughput through multi-core parallelism. After tcore0 executes the first instruction and processes the first blue data block 102, if tcore0 completes the first instruction, and then tcore0 executes the second instruction to process the second blue data block 106, it is necessary to use the coordinate increment data (slope) on each dimension to indicate the starting coordinates of the second instruction. It should be noted that the slope refers to the coordinate increment of each dimension (such as width and height), which is used to determine the starting position of the next instruction. The slope on any dimension cannot exceed the image size of any dimension. The image size can be the maximum allowed value on each dimension, and the slope cannot exceed the image size to avoid crossing the boundary.
[0056] For related technologies, see Figure 1b , Figure 1b The following is a timing diagram of the Tensor Core execution process. The specific coordinate update process includes the initial operation cycle (cycle0), the first operation cycle (cycle1), the second operation cycle (cycle2), and the third operation cycle (cycle3). w_coord_b is the starting coordinate of width, h_coord_b is the starting coordinate of height, d_coord_b is the starting coordinate of depth, and n_coord_b is the starting coordinate of batch.
[0057] w_slope is the width coordinate increment, h_slope is the height coordinate increment, d_slope is the depth coordinate increment, and n_slope is the batch coordinate increment. h_stride is the height stride, and d_stride is the depth stride. The size of the image data in the width dimension is denoted as tensor_w, the size in the height dimension is denoted as tensor_h, and the size in the depth dimension is denoted as tensor_d.
[0058] Exemplarily illustrate the operation process in the initial operation cycle (cycle0), including S01 to S07. In the initial operation cycle, the width updated coordinate is denoted as w_coord_c0, the height updated coordinate is denoted as h_coord_c0, the depth updated coordinate is denoted as d_coord_c0, and the batch updated coordinate is denoted as n_coord_c0.
[0059] S01. Add the width coordinate increment w_slope to the width starting coordinate w_coord_b to obtain w_coord_b + w_slope.
[0060] S02. Compare the size of w_coord_b + w_slope and tensor_w, and generate the width enable signal w_cross_en0 according to the comparison result.
[0061] If it is determined according to the width enable signal w_cross_en0 that w_coord_b + w_slope is not greater than tensor_w, the width updated coordinate w_coord_c0 in the initial operation cycle is equal to w_coord_b + w_slope. It should be noted that the height starting coordinate h_coord_b, the depth starting coordinate d_coord_b, and the batch starting coordinate n_coord_b remain unchanged. That is, the height updated coordinate h_coord_c0 in the initial operation cycle is equal to the height starting coordinate h_coord_b, the depth updated coordinate d_coord_c0 in the initial operation cycle is equal to the depth starting coordinate d_coord_b, and the batch updated coordinate n_coord_c0 in the initial operation cycle is equal to the batch starting coordinate n_coord_b.
[0062] If it is determined according to the width enable signal w_cross_en0 that w_coord_b + w_slope is greater than tensor_w, the width updated coordinate w_coord_c0 in the initial operation cycle is equal to w_coord_b + w_slope - tensor_w.
[0063] S03. If w_coord_b + w_slope is greater than tensor_w, increase the height step h_stride at the height starting coordinate h_coord_b to obtain h_coord_b + h_stride.
[0064] S04. Compare the magnitudes of h_coord_b + h_stride and tensor_h, and generate the height enable signal h_cross_en0 based on the comparison result.
[0065] If it is determined according to the height enable signal h_cross_en0 that h_coord_b + h_stride is not greater than tensor_h, the height updated coordinate h_coord_c0 in the initial operation cycle is equal to h_coord_b + h_stride. It should be noted that in the initial operation cycle, the depth updated coordinate d_coord_c0 is equal to the depth starting coordinate d_coord_b, and the batch updated coordinate n_coord_c0 in the initial operation cycle is equal to the batch starting coordinate n_coord_b.
[0066] If it is determined according to the height enable signal h_cross_en0 that h_coord_b + h_stride is greater than tensor_h, the height updated coordinate h_coord_c0 in the initial operation cycle is equal to h_coord_b + h_stride - tensor_h.
[0067] S05. If h_coord_b + h_stride is greater than tensor_h, increase the depth step d_stride at the depth starting coordinate d_coord_b to obtain d_coord_b + d_stride.
[0068] S06. Compare the magnitudes of d_coord_b + d_stride and tensor_d, and generate the depth enable signal d_cross_en0 based on the comparison result.
[0069] If it is determined according to the depth enable signal d_cross_en0 that d_coord_b + d_stride is not greater than tensor_d, the depth updated coordinate d_coord_c0 in the initial operation cycle is equal to d_coord_b + d_stride. It should be noted that the batch updated coordinate n_coord_c0 in the initial operation cycle is equal to the batch starting coordinate n_coord_b.
[0070] If it is determined according to the depth enable signal d_cross_en0 that d_coord_b + d_stride is greater than tensor_d, the depth updated coordinate d_coord_c0 in the initial operation period is equal to d_coord_b + d_stride - tensor_d.
[0071] S07. If d_coord_b + d_stride is greater than tensor_d, add 1 to the batch start coordinate n_coord_b to obtain the batch updated coordinate n_coord_c0 in the initial operation period.
[0072] Exemplarily illustrate the operation process in the first operation cycle (cycle1), including S11 to S15. In the first operation cycle, the height updated coordinate is denoted as h_coord_c1, the depth updated coordinate is denoted as d_coord_c1, and the batch updated coordinate is denoted as n_coord_c1.
[0073] S11. Add the height coordinate increment h_slope to the height updated coordinate h_coord_c0 in the initial operation period to obtain h_coord_c0 + h_slope.
[0074] S12. Compare the size of h_coord_c0 + h_slope and tensor_h, and generate the height enable signal h_cross_en1 according to the comparison result.
[0075] If it is determined according to the height enable signal h_cross_en1 that h_coord_c0 + h_slope is not greater than tensor_h, the height updated coordinate h_coord_c1 in the first operation cycle is equal to h_coord_c0 + h_slope. It should be noted that the depth updated coordinate d_coord_c0 in the initial operation period and the batch updated coordinate n_coord_c0 in the initial operation period remain unchanged. That is, the depth updated coordinate d_coord_c1 in the first operation cycle is equal to d_coord_c0, and the batch updated coordinate n_coord_c1 in the first operation cycle is equal to the batch start coordinate n_coord_c0.
[0076] If it is determined according to the height enable signal h_cross_en1 that h_coord_c0 + h_slope is greater than tensor_h, the height updated coordinate h_coord_c1 in the first operation cycle is equal to h_coord_c0 + h_slope - tensor_h.
[0077] S13. If h_coord_c0 + h_slope is greater than tensor_h, add the depth stride d_stride to the depth-updated coordinate d_coord_c0 in the initial operation cycle to obtain d_coord_c0 + d_stride.
[0078] S14. Compare the size of d_coord_c0 + d_stride and tensor_d, and generate the depth enable signal d_cross_en1 according to the comparison result.
[0079] If it is determined according to the depth enable signal d_cross_en1 that d_coord_c0 + d_stride is not greater than tensor_d, the depth-updated coordinate d_coord_c1 in the first operation cycle is equal to d_coord_c0 + d_stride. It should be noted that the batch-updated coordinate n_coord_c1 in the first operation cycle is equal to the batch start coordinate n_coord_c0.
[0080] If it is determined according to the depth enable signal d_cross_en1 that d_coord_b + d_stride is greater than tensor_d, the depth-updated coordinate d_coord_c1 in the first operation cycle is equal to d_coord_c0 + d_stride - tensor_d.
[0081] S15. If d_coord_c0 + d_stride is greater than tensor_d, add 1 to the batch-updated coordinate n_coord_c0 in the initial operation cycle to obtain the batch-updated coordinate n_coord_c1 in the first operation cycle.
[0082] Exemplarily illustrate the operation process in the second operation cycle (cycle2), including S21 to S23. In the second operation cycle, the depth-updated coordinate is denoted as d_coord_c2, and the batch-updated coordinate is denoted as n_coord_c2.
[0083] S21. Add the depth coordinate increment d_slope to the depth-updated coordinate d_coord_c1 in the first operation cycle to obtain d_coord_c1 + d_slope.
[0084] S22. Compare the size of d_coord_c1 + d_slope and tensor_d, and generate the depth enable signal d_cross_en2 according to the comparison result.
[0085] If it is determined according to the depth enable signal d_cross_en2 that d_coord_c1 + d_slope is not greater than tensor_d, the depth updated coordinate d_coord_c2 in the second operation cycle is equal to d_coord_c1 + d_slope. It should be noted that the batch updated coordinate n_coord_c2 in the second operation cycle is equal to the batch start coordinate n_coord_c1.
[0086] If it is determined according to the depth enable signal d_cross_en2 that d_coord_c1 + d_slope is greater than tensor_d, the depth updated coordinate d_coord_c2 in the second operation cycle is equal to d_coord_c1 + d_slope - tensor_d.
[0087] S23: If d_coord_c1 + d_slope is greater than tensor_d, add 1 to the batch updated coordinate n_coord_c1 in the first operation cycle to obtain the batch updated coordinate n_coord_c2 in the second operation cycle.
[0088] Exemplarily illustrate the operation process in the third operation cycle (cycle3). In the third operation cycle, the batch updated coordinate is denoted as n_coord_c3. Add the batch coordinate increment n_slope to the batch updated coordinate n_coord_c2 in the second operation cycle to obtain n_coord_c3. And n_coord_c3 is equal to n_coord_c2 + n_slope.
[0089] So far, through the calculations in the initial operation cycle, the first operation cycle, the second operation cycle, and the third operation cycle, the updated coordinates after incrementing slope in each dimension are obtained. It can be understood that the addition operation in the above calculation process can be implemented by an adder, and the comparison operation in the above calculation process can be implemented by a comparator.
[0090] After analysis, it is found that in the related art, the method of accumulating according to the coordinate increments of each dimension in the order from low to high dimension requires more adder resources. It can be seen that the method of accumulating in different operation cycles according to each dimension requires a large hardware overhead, and the area occupied by the adder is large. Further, the hardware timing of the method of accumulating in different operation cycles according to each dimension needs to be further optimized, and the number of cycle levels needs to be further reduced.
[0091] Based on the above analysis, the present application provides a method for generating a physical address. By pre-computing the coordinate increments in each dimension, the hardware resources consumed at each level are reduced, and the number of cycle levels is further reduced. Specifically, the data transfer instruction carries the size of the data block to be transferred and the starting coordinates in each dimension. The coordinate increments in each dimension are calculated based on the size of the data block to be transferred, the dimensions of the image data in each dimension, and the starting coordinates in each dimension. In the initial operation cycle, the corresponding coordinate increments in each dimension are pre-added to the starting coordinates in each dimension to obtain the initial updated coordinates in each dimension. An initial cross-dimensional enable signal in the initial operation cycle is generated based on the initial updated coordinates in each dimension. By pre-generating the initial cross-dimensional enable signal, the number of adders required in the operation cycle can be reduced, thereby achieving efficient utilization of hardware resources and performance optimization.
[0092] In the scenario example of the present application, please refer to Figure 1c , Figure 1c which is a partial structural schematic diagram of an artificial intelligence processor. The artificial intelligence processor can be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), and a GPGPU (General-Purpose Graphics Processing Unit). The Tensor Memory Accelerator (TMA) is dedicated to the transfer of multi-dimensional data (such as tensors). When a computing task needs to load data from the HBM, the instruction scheduler outputs a data transfer instruction. The TMA receives the data transfer instruction, and the data transfer instruction carries the size of the data block to be transferred, the starting coordinates, and the coordinate increment (slope). The TMA generates a physical address based on the data transfer instruction and drives the HBM to complete the data transfer based on the generated physical address.
[0093] In this scenario example, please refer to Figure 1d, The coordinate update process includes an initial operation cycle (cycle0), a first operation cycle (cycle1), and a second operation cycle (cycle2). w_coord_b is the starting coordinate of width, h_coord_b is the starting coordinate of height, d_coord_b is the starting coordinate of depth, and n_coord_b is the starting coordinate of batch. w_slope is the width coordinate increment, h_slope is the height coordinate increment, d_slope is the depth coordinate increment, and n_slope is the batch coordinate increment. h_stride is the height stride, and d_stride is the depth stride. The size of the image data in width is denoted as tensor_w, the size in height is denoted as tensor_h, and the size in depth is denoted as tensor_d.
[0094] Exemplarily illustrate the operation process in the initial operation cycle (cycle0), including S31 to S32. During the initial operation cycle, the updated width coordinate is denoted as w_coord_c0, the updated height coordinate is denoted as h_coord_c0, the updated depth coordinate is denoted as d_coord_c0, and the updated batch coordinate is denoted as n_coord_c0.
[0095] S31. Add the width coordinate increment w_slope to the starting width coordinate w_coord_b to obtain w_coord_b + w_slope; add the height coordinate increment h_slope to the starting height coordinate h_coord_b to obtain h_coord_b + h_slope; add the depth coordinate increment d_slope to the starting depth coordinate d_coord_b to obtain d_coord_b + d_slope; add the batch coordinate increment n_slope to the starting batch coordinate n_coord_b to obtain n_coord_b + n_slope.
[0096] S32. Compare the size of w_coord_b + w_slope with tensor_w, and generate a width enable signal w_cross_c0 according to the comparison result; compare the size of h_coord_b + h_slope with tensor_h, and generate a height enable signal h_cross_c0 according to the comparison result; compare the size of d_coord_b + d_slope with tensor_d, and generate a depth enable signal d_cross_c0 according to the comparison result.
[0097] Specifically, if the width enable signal w_cross_c0 indicates that w_coord_b + w_slope is greater than tensor_w, the width-updated coordinate w_coord_c0 in the initial operation cycle is equal to w_coord_b + w_slope - tensor_w. If the width enable signal w_cross_c0 indicates that w_coord_b + w_slope is not greater than tensor_w, the width-updated coordinate w_coord_c0 in the initial operation cycle is equal to w_coord_b + w_slope.
[0098] If the height enable signal h_cross_c0 indicates that h_coord_b + h_slope is greater than tensor_h, the height-updated coordinate h_coord_c0 in the initial operation cycle is equal to h_coord_b + h_slope - tensor_h. If the height enable signal h_cross_c0 indicates that h_coord_b + h_slope is not greater than tensor_h, the height-updated coordinate h_coord_c0 in the initial operation cycle is equal to h_coord_b + h_slope.
[0099] If the depth enable signal d_cross_c0 indicates that d_coord_b + d_slope is greater than tensor_d, the depth-updated coordinate d_coord_c0 in the initial operation cycle is equal to d_coord_b + d_slope - tensor_d. If the depth enable signal d_cross_c0 indicates that d_coord_b + d_slope is not greater than tensor_d, the depth-updated coordinate d_coord_c0 in the initial operation cycle is equal to d_coord_b + d_slope.
[0100] If the depth enable signal d_cross_c0 indicates that d_coord_b + d_slope is greater than tensor_d, the batch-updated coordinate n_coord_c0 in the initial operation cycle is equal to n_coord_b + n_slope + 1. If the depth enable signal d_cross_c0 indicates that d_coord_b + d_slope is not greater than tensor_d, the batch-updated coordinate n_coord_c0 in the initial operation cycle is equal to n_coord_b + n_slope.
[0101] Exemplarily illustrate the operation process in the first operation cycle (cycle1), including S41 to S46. In the first operation cycle, the height-updated coordinate is denoted as h_coord_c1, the depth-updated coordinate is denoted as d_coord_c1, and the batch-updated coordinate is denoted as n_coord_c1.
[0102] S41. If the width enable signal w_cross_c0 indicates that w_coord_b + w_slope is greater than tensor_w, then in the initial operation cycle, the height step h_stride is added to the height-updated coordinate h_coord_c0 to obtain h_coord_c0 + h_stride.
[0103] S42. Compare the size of h_coord_c0 + h_stride and tensor_h, and generate the height enable signal h_cross_c1 according to the comparison result.
[0104] Specifically, if the height enable signal h_cross_c1 indicates that h_coord_c0 + h_stride is greater than tensor_h, the height-updated coordinate h_coord_c1 in the first operation cycle is equal to h_coord_c0 + h_stride - tensor_h. If the height enable signal h_cross_c1 indicates that h_coord_c0 + h_stride is not greater than tensor_h, the height-updated coordinate h_coord_c1 in the first operation cycle is equal to h_coord_c0 + h_stride.
[0105] It should be noted that if the width enable signal w_cross_c0 indicates that w_coord_b + w_slope is not greater than tensor_w, the height-updated coordinate h_coord_c1 in the first operation cycle is equal to the height-updated coordinate h_coord_c0 in the initial operation cycle.
[0106] S43. If it is determined according to the height enable signal h_cross_c0 that h_coord_b + h_slope is greater than tensor_h, then in the initial operation cycle, the depth step d_stride is added to the depth-updated coordinate d_coord_c0 to obtain d_coord_c0 + d_stride.
[0107] S44. Compare the size of d_coord_c0 + d_stride and tensor_d, and generate the depth enable signal d_cross_c1 according to the comparison result.
[0108] Specifically, if the depth enable signal d_cross_c1 indicates that d_coord_c0 + d_stride is greater than tensor_d, the depth updated coordinate d_coord_c1 in the first operation cycle is equal to d_coord_c0 + d_stride - tensor_d. If the height enable signal d_cross_c1 indicates that d_coord_c0 + d_stride is not greater than tensor_d, the height updated coordinate d_coord_c1 in the first operation cycle is equal to d_coord_c0 + d_stride.
[0109] It should be noted that if the height enable signal h_cross_c0 indicates that h_coord_b + h_slope is not greater than tensor_h, the depth updated coordinate d_coord_c1 in the first operation cycle is equal to the depth updated coordinate d_coord_c0 in the initial operation cycle.
[0110] S45. If it is determined according to the depth enable signal d_cross_c1 that d_coord_c0 + d_stride is greater than tensor_d, the batch updated coordinate n_coord_c1 in the first operation cycle is equal to the batch updated coordinate n_coord_c0 in the initial operation cycle plus 1.
[0111] S46. If it is determined according to the depth enable signal d_cross_c1 that d_coord_c0 + d_stride is not greater than tensor_d, the batch updated coordinate n_coord_c1 in the first operation cycle is equal to the batch updated coordinate n_coord_c0 in the initial operation cycle.
[0112] Exemplarily illustrate the operation process in the second operation cycle (cycle2), including S51 to S54. In the second operation cycle, the depth updated coordinate is denoted as d_coord_c2, and the batch updated coordinate is denoted as n_coord_c2.
[0113] S51. If the height enable signal h_cross_c1 indicates that h_coord_c0 + h_stride is greater than tensor_h, add the depth step d_stride to the depth updated coordinate d_coord_c1 in the first operation cycle to obtain d_coord_c1 + d_stride.
[0114] S52. Compare d_coord_c1 + d_stride with tensor_d, and generate the depth enable signal d_cross_c2 according to the comparison result.
[0115] Specifically, if the depth enable signal d_cross_c2 indicates that d_coord_c1 + d_stride is greater than tensor_d, the depth updated coordinate d_coord_c2 in the second operation cycle is equal to d_coord_c1 + d_stride - tensor_d. If the height enable signal d_cross_c2 indicates that d_coord_c1 + d_stride is not greater than tensor_d, the depth updated coordinate d_coord_c2 in the second operation cycle is equal to d_coord_c1 + d_stride.
[0116] S53. If it is determined according to the depth enable signal d_cross_c2 that d_coord_c1 + d_stride is greater than tensor_d, the batch updated coordinate n_coord_c2 in the second operation cycle is equal to the batch updated coordinate n_coord_c1 in the first operation cycle plus 1.
[0117] S54. If it is determined according to the depth enable signal d_cross_c2 that d_coord_c1 + d_stride is not greater than tensor_d, the batch updated coordinate n_coord_c2 in the second operation cycle is equal to the batch updated coordinate n_coord_c1 in the first operation cycle.
[0118] So far, through the calculations in the initial operation cycle, the first operation cycle, and the second operation cycle, the updated coordinates in each dimension are obtained, and the updated coordinates in all dimensions are aggregated to generate a physical address. Finally, according to the generated physical address, the TMA drives the HBM to perform a read (load) operation or a write (store) operation, completing the efficient transfer of data.
[0119] In this scenario example, by adding the coordinate increments in each dimension to the starting coordinates of the corresponding dimensions in advance during the initial operation cycle, the hardware resources consumed at each level are reduced, and the number of cycle levels is reduced.
[0120] According to an embodiment of the present application, an embodiment of a method for generating a physical address is provided. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0121] In this embodiment, a method for generating a physical address is provided, which can be used in a hardware accelerator. The hardware accelerator can be an integrated circuit module capable of accelerating the design of multi-dimensional data calculation and transfer. For example, the hardware accelerator can be a computing device such as a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), or a tensor memory accelerator TMA in a computing device. Please refer to Figure 1e , and the generation method includes the following steps:
[0122] S110. Determine the starting coordinates and coordinate increment data of the multi-dimensional data in each dimension.
[0123] Among them, the multi-dimensional data can be a structured data set organized in multiple independent dimensions, and each dimension represents the distribution of data in a certain direction. The multi-dimensional data can be two-dimensional data or a four-dimensional tensor, such as including four dimensions of batch (N), depth (D), height (H), and width (W). For example, in an image processing scenario, a two-dimensional image with a size of 100×50 can be regarded as two-dimensional data of height (H = 50) and width (W = 100), while a data set containing multiple images is three-dimensional data (N×H×W). The coordinates of the multi-dimensional data in each dimension are used to locate the logical position of the data block in memory, and its physical address can be calculated and generated through the dimension stride and coordinate increment (slope).
[0124] The starting coordinates can be the initial logical position values of each dimension in the multi-dimensional data. The starting coordinates can define the starting point of data processing. For example, in the width dimension (W), the starting coordinates can be set to w = 0, indicating starting from the first pixel for processing. For example, the starting coordinates can be expressed as (n = 0, d = 0, h = 0, w = 0). It can be understood that the setting of the starting coordinates needs to be aligned with the memory layout to ensure the continuity of subsequent coordinate increment operations.
[0125] The coordinate increment data (slope) can be how much to increment the coordinates in each dimension. The absolute value of the coordinate increment data needs to be less than or equal to the maximum size of the corresponding dimension. For example, if the image width W = 100, then the value of w_slope is less than or equal to 100 to prevent coordinate out-of-bounds.
[0126] Specifically, the current pen instruction carries the coordinate increment data, data block size, and starting coordinates required for moving data. Further, the coordinate increment data is pre-calculated. Determine the sizes of each dimension in the multi-dimensional data, and use the data block size and the sizes of each dimension in the multi-dimensional data for pre-calculation to obtain the coordinate increment data in each dimension. After the current pen instruction is completed, the starting position of the next pen instruction of the current pen instruction can be located through the coordinate increment data carried in the current pen instruction, reducing the resource consumption caused by real-time calculation.
[0127] S120. In the initial operation cycle, add the coordinate increment data of each dimension to the starting coordinates of the corresponding dimension to obtain the initial updated coordinates of each dimension and the initial cross-dimension enable signal in the initial operation cycle.
[0128] Among them, the initial operation cycle can be the clock cycle for the first execution of coordinate calculation in the hardware pipeline, which can be denoted as cycle0. The initial updated coordinates can refer to the intermediate calculation results obtained by adding the coordinate increments corresponding to each dimension to the starting coordinates of each dimension. This intermediate calculation result can be temporarily stored in the register bank for subsequent judgment of whether it causes a jump from the current dimension to the adjacent higher dimension.
[0129] Specifically, within the initial operation cycle (cycle0), the starting coordinates on each dimension and the coordinate increment data are synchronously added to obtain the initial updated coordinates for each dimension. The addition operation of the coordinate increment data and the starting coordinates on each dimension is implemented by a hardware adder. In the initial operation cycle, the adder for the width dimension performs the operation of w_coord_b + w_slope. For example, if the starting coordinate w = 50 and w_slope = 10, after the addition operation, the initial updated coordinate for the width dimension is equal to 60.
[0130] Exemplarily, for two-dimensional data, the starting coordinates are (w = 0, h = 0), the coordinate increment data w_slope in the width dimension is equal to 1, and the coordinate increment data h_slope in the height dimension is equal to 1. The operations of w + 1 and h + 1 are respectively executed in the initial operation cycle to generate the initial updated coordinates for the width dimension and the height dimension respectively.
[0131] In this embodiment, the initial cross-dimension enable signal within the initial operation cycle is used to indicate whether the initial updated coordinates cause a cross-dimension in the adjacent higher dimension of each dimension. The initial cross-dimension enable signal is generated based on the initial updated coordinates of each dimension. Among them, the initial cross-dimension enable signal is generated by a comparator. If the initial updated coordinate of the current dimension exceeds the size of the current dimension, the initial cross-dimension enable signal can be pulled high, and the jump from the current dimension to the adjacent higher dimension is triggered by this initial cross-dimension enable signal.
[0132] The adjacent higher dimension can refer to the next dimension with a higher logical level than the current dimension in the multi-dimensional data. In four-dimensional data (NDHW), the adjacent higher dimension of the width (W) is the height (H), the adjacent higher dimension of the height is the depth (D), and the adjacent higher dimension of the depth is the batch (N). For example, when the width coordinate overflows, the coordinate needs to be incremented in the height dimension.
[0133] In this embodiment, within the initial operation cycle, for the initially updated coordinates of each dimension, it is determined whether the initially updated coordinates of the current dimension exceed the size of the current dimension, and an initial cross-dimension enable signal is generated. The initial cross-dimension enable signal indicates whether the coordinate update of the current dimension causes the coordinate adjustment of the adjacent higher dimension. Further, since it is indicated by the initial cross-dimension enable signal whether cross-dimension coordinate updates are triggered after adding coordinate increments in each dimension, the coordinate calculation result of each dimension in the initial operation cycle is determined according to the initial cross-dimension enable signal. The initial coordinate calculation result may include the calculation results obtained by calculating the starting coordinates of each dimension respectively after the initial operation cycle. Exemplarily, the starting coordinates may be denoted as (w_coord_b, h_coord_b, d_coord_b, n_coord_b), and the initial coordinate calculation result may be denoted as (w_coord_c0, h_coord_c0, d_coord_c0, n_coord_c0).
[0134] S130. Perform coordinate calculation and address conversion based on the initially updated coordinates and the initial cross-dimension enable signal to obtain the physical address of the multi-dimensional data.
[0135] Among them, the hardware accelerator may include a coordinate calculation unit and an address conversion unit. The physical address may refer to the actual storage location of the multi-dimensional data in the memory. Specifically, since the initial cross-dimension enable signal within the initial operation cycle indicates whether the initially updated coordinates of the current dimension cause the adjustment of the adjacent higher dimension, therefore, within the initial operation cycle, the initially updated coordinates are calculated according to the initial cross-dimension enable signal to obtain the initial coordinate calculation result of the initial operation cycle. In the subsequent operation cycles of the initial operation cycle, the initial coordinate calculation result and the initial cross-dimension enable signal may be used as the inputs of the subsequent operation cycles. In the subsequent operation cycles, the initial coordinate calculation result is calculated according to the initial cross-dimension enable signal to obtain the subsequent coordinate calculation result of the subsequent operation cycles. The initial coordinate calculation result and the subsequent coordinate calculation results are summarized to obtain the logical coordinates of the multi-dimensional data. The logical coordinates of the multi-dimensional data are converted into a physical memory address by the address conversion unit to obtain the physical address of the multi-dimensional data. Exemplarily, the hardware accelerator is used to map the logical coordinates of a four-dimensional tensor (batch N, depth D, height H, width W) to the linear physical address of the HBM, and the physical address can be determined by linearization calculation. It should be noted that the subsequent operation cycles of the initial operation cycle may be one operation cycle or two operation cycles.
[0136] In the above embodiments, first, the starting coordinates and coordinate increment data of the multi-dimensional data in each dimension are determined; secondly, within the initial operation cycle, the coordinate increment data of each dimension is added to the starting coordinate of the corresponding dimension in advance to obtain the initial updated coordinates of each dimension, and an initial cross-dimension enabling signal within the initial operation cycle is further generated. Finally, based on the initial updated coordinates and the initial cross-dimension enabling signal, coordinate calculation and address conversion are performed to obtain the physical address of the multi-dimensional data. Compared with the calculation method of accumulating by dimension in the related art, the adder resources utilized in the calculation process of this embodiment are less, so that the area occupied by the adder can be reduced, the generation of the physical address can be achieved with less hardware overhead, and further the logic resource overhead within the initial operation cycle can be reduced.
[0137] In some embodiments, in step S130, performing coordinate calculation and address conversion based on the initial updated coordinates and the initial cross-dimension enabling signal to obtain the physical address of the multi-dimensional data includes:
[0138] S132. Within the initial operation cycle, perform coordinate calculation based on the initial cross-dimension enabling signal and the initial updated coordinates to obtain the initial coordinate calculation result of each dimension within the initial operation cycle.
[0139] Specifically, the initial operation cycle is the first clock cycle for performing coordinate calculation in the hardware pipeline. Within the initial operation cycle, the corresponding coordinate increment data is added to the starting coordinate of each dimension to obtain the initial updated coordinates of each dimension. Further, within the initial operation cycle, the initial updated coordinates of each dimension are calculated through the initial cross-dimension enabling signal to obtain the initial coordinate calculation result of each dimension.
[0140] In some embodiments, the size of the multi-dimensional data in each dimension is denoted as the dimension size. Performing coordinate calculation based on the initial cross-dimension enabling signal and the initial updated coordinates to obtain the initial coordinate calculation result of each dimension within the initial operation cycle includes: If the initial cross-dimension enabling signal indicates that the initial updated coordinates cause cross-dimension, subtract the dimension size of the corresponding dimension from the initial updated coordinates to obtain the initial coordinate calculation result. If the initial cross-dimension enabling signal indicates that the initial updated coordinates do not cause cross-dimension, use the initial updated coordinates as the initial coordinate calculation result.
[0141] In this embodiment, the dimensional size may be the maximum size of each dimension in the multi-dimensional data. Exemplarily, each dimension is regarded as the current dimension, and the initial updated coordinate of the current dimension is obtained by adding the starting coordinate of the current dimension and the coordinate increment data of the current dimension. In the case where the initial updated coordinate of the current dimension exceeds the dimensional size of the current dimension, the generated initial cross-dimensional enable signal indicates that cross-dimension is required, and the initial coordinate calculation result on the current dimension is equal to the initial updated coordinate minus the dimensional size on the current dimension. In the case where the initial updated coordinate of the current dimension does not exceed the dimensional size of the current dimension, the generated initial cross-dimensional enable signal indicates that cross-dimension is not required, and the initial coordinate calculation result on the current dimension is equal to the initial updated coordinate.
[0142] S134. In the first operation cycle after the initial operation cycle, perform coordinate calculation on the initial coordinate calculation result based on the initial cross-dimensional enable signal to obtain the first coordinate calculation result of the first operation cycle.
[0143] Among them, the first operation cycle may be the second calculation stage following the initial operation cycle in the hardware pipeline. Specifically, in the initial operation cycle, an initial cross-dimensional enable signal and an initial coordinate calculation result are obtained. The first operation cycle receives the initial cross-dimensional enable signal and the initial coordinate calculation result, and determines whether to accumulate the dimension step length in the corresponding dimension according to the initial cross-dimensional enable signal to obtain the first coordinate calculation result in the corresponding dimension.
[0144] S136. Perform coordinate calculation and address conversion according to the initial coordinate calculation result and the first coordinate calculation result to obtain the physical address of the multi-dimensional data.
[0145] Among them, the initial coordinate calculation result is the intermediate calculation result after the end of the initial operation cycle. The first coordinate calculation result is the intermediate calculation result after the end of the first operation cycle. The initial coordinate calculation result and the first coordinate calculation result are used for subsequent address conversion. Specifically, determine the logical coordinates of the multi-dimensional data based on the initial coordinate calculation result and the first coordinate calculation result, and the logical coordinates of the multi-dimensional data can be converted into the physical address of the multi-dimensional data according to the physical address calculation formula.
[0146] In the above embodiment, the coordinate updates of each dimension are synchronously processed in the initial operation cycle, and the parallel calculation optimization of coordinate self-increment is realized in a hardware manner to simplify the hardware area. Further, coordinate calculation and address conversion are performed based on the output of the initial operation cycle after the initial operation cycle to realize the conversion from logical coordinates to physical addresses, optimize the hardware timing performance to a certain extent, and improve the data transfer efficiency.
[0147] In some embodiments, a first cross-dimension enabling signal is generated during a first operation cycle. The first cross-dimension enabling signal is used to indicate whether the first updated coordinates in the first operation cycle cause cross-dimension on the adjacent higher dimension of the corresponding dimension. For any dimension calculated during the first operation cycle, the first updated coordinates of the any dimension in the first operation cycle can be equal to the initial coordinate calculation result of the any dimension plus the dimension step of the any dimension. In step S136, coordinate calculation and address conversion are performed according to the initial coordinate calculation result and the first coordinate calculation result to obtain the physical address of the multi-dimensional data, including:
[0148] S1362. During a second operation cycle after the first operation cycle, coordinate calculation is performed on the first coordinate calculation result based on the first cross-dimension enabling signal to obtain the second coordinate calculation result of the second operation cycle.
[0149] Wherein, the second operation cycle can be the third calculation stage following the first operation cycle in the hardware pipeline. The second coordinate calculation result can refer to the intermediate calculation result on the corresponding dimension after the end of the second operation cycle. Specifically, a first cross-dimension enabling signal and a first coordinate calculation result are obtained during the first operation cycle, and the second operation cycle receives the first cross-dimension enabling signal and the first coordinate calculation result, and determines whether to accumulate the dimension step on the corresponding dimension according to the first cross-dimension enabling signal to obtain the second coordinate calculation result on the corresponding dimension.
[0150] S1364. During the second operation cycle, summarization and address conversion are performed based on the initial coordinate calculation result, the first coordinate calculation result, and the second coordinate calculation result to obtain the physical address of the multi-dimensional data.
[0151] Wherein, the initial coordinate calculation result, the first coordinate calculation result, and the second coordinate calculation result are used for subsequent address conversion. Specifically, the logical coordinates of the multi-dimensional data are determined based on the initial coordinate calculation result, the first coordinate calculation result, and the second coordinate calculation result, and linearization calculation is performed according to the physical address calculation formula to convert the logical coordinates of the multi-dimensional data into the physical address of the multi-dimensional data.
[0152] In the above embodiments, the coordinate updates of each dimension are synchronously processed in the initial operation cycle, and the parallel calculation optimization of coordinate self-increment is realized in a hardware manner to simplify the hardware area. Further, coordinate calculation is performed based on the output of the initial operation cycle in the first operation cycle, and coordinate calculation is performed based on the output of the first operation cycle in the second operation cycle, and then address conversion is performed based on the initial coordinate calculation result, the first coordinate calculation result, and the second coordinate calculation result to realize the conversion from logical coordinates to physical addresses, optimizing the hardware timing performance to a certain extent and improving the data transfer efficiency.
[0153] In some embodiments, the sizes of multi-dimensional data in each dimension are denoted as dimension sizes.
[0154] Adding the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinates of each dimension and the initial cross-dimension enabling signal within the initial operation period, including: adding the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinates of each dimension; using the initial updated coordinates of each dimension and the dimension size on the corresponding dimension to perform cross-dimension judgment on adjacent higher dimensions to generate the initial cross-dimension enabling signal.
[0155] Specifically, each dimension is regarded as the current dimension, and the coordinate increment data of the current dimension is added to the starting coordinate of the current dimension to obtain the initial updated coordinates of the current dimension. Further, comparing the initial updated coordinates of the current dimension and the dimension size on the current dimension, and generating the initial cross-dimension enabling signal according to the comparison result to indicate whether the initial updated coordinates of the current dimension cause cross-dimension on the adjacent higher dimension of the current dimension, that is, determining whether to trigger the logical operation of updating the coordinates of the adjacent higher dimension. For example, if the initial updated coordinates of the current dimension are greater than the dimension size on the current dimension, the initial updated coordinates of the current dimension cause cross-dimension on the adjacent higher dimension. If the initial updated coordinates of the current dimension are not greater than the dimension size on the current dimension, the initial updated coordinates of the current dimension do not cause cross-dimension on the adjacent higher dimension.
[0156] Exemplarily, the current dimension is the width, the width coordinate is denoted as W, the adjacent higher dimension of the current dimension is the height H. The size of the current dimension is denoted as W0. The coordinate increment data of the current dimension is denoted as W_slope. The initial updated coordinates of the current dimension are equal to W + W_slope, comparing the size of W + W_slope and W0, and correspondingly generating the initial cross-dimension enabling signal. For example, the initial cross-dimension enabling signal is denoted as w_cross_c0. If W + W_slope is greater than W0, the initial cross-dimension enabling signal is pulled high, that is, w_cross_c0 = 1. Otherwise, the initial cross-dimension enabling signal remains invalid, that is, w_cross_c0 = 0.
[0157] In the above embodiments, by comparing the initial updated coordinates of each dimension and the dimension size on the corresponding dimension, cross-dimension judgment is performed on adjacent higher dimensions, and the cross-dimension judgment result is converted into a signal executable by hardware to coordinate multi-dimensional coordinate updates, thereby realizing the dynamic triggering of cross-dimension operations and ensuring data continuity.
[0158] In some embodiments, the lowest dimension among the various dimensions of the multi-dimensional data is denoted as the first lowest dimension. The step size in each dimension of the multi-dimensional data is denoted as the dimension step size. Coordinate calculation is performed on the initial coordinate calculation result based on the initial cross-dimension enabling signal to obtain the first coordinate calculation result of the first operation cycle, including: determining whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross-dimension enabling signal to obtain the first coordinate calculation result of the first target dimension within the first operation cycle.
[0159] Among them, the first target dimension is any dimension other than the first lowest dimension among the various dimensions. The first lowest dimension may refer to the dimension with the lowest logical level in the multi-dimensional data, that is, the innermost dimension where the data is continuously stored in memory. In four-dimensional data (NDHW), the first lowest dimension is usually the width dimension (W), indicating that the data is stored continuously by rows. For example, the coordinate update frequency of the first lowest dimension is the highest. The first target dimension may be the height dimension (H), may be the depth dimension (D), or may also be the batch dimension (N). The dimension step size can be used to describe the physical address interval of cross-dimension data access. The step size in each dimension can be the depth step size or the height step size. The first operation cycle (cycle1) after the initial operation cycle may refer to the second calculation stage in the hardware pipeline following the initial operation cycle (cycle0), which is used to process the coordinate update of the first target dimension. The input of the first operation cycle is the output of the initial operation cycle, forming a pipeline processing chain. The first coordinate calculation result includes the calculation result obtained after the calculation of the initial coordinate calculation result of the first target dimension in the first operation cycle.
[0160] Specifically, the first target dimension may be the height dimension, and its height coordinate is denoted as H. In the initial operation cycle, since the initial cross-dimension enabling signal of the width dimension indicates whether the initial updated coordinate of the width dimension can cause cross-dimension in the height dimension, in the first operation cycle after the initial operation cycle, it is determined whether to accumulate the dimension step size h_stride of the height dimension on the initial coordinate calculation result of the height dimension based on the initial cross-dimension enabling signal of the width dimension to obtain the first coordinate calculation result of the height dimension within the first operation cycle.
[0161] The first target dimension may be the depth dimension, and its depth coordinate is denoted as D. In the initial operation cycle, since the initial cross-dimension enabling signal of the height dimension indicates whether the initial updated coordinate of the height dimension can cause cross-dimension in the depth dimension, in the first operation cycle after the initial operation cycle, it is determined whether to accumulate the dimension step size d_stride of the depth dimension on the initial coordinate calculation result of the depth dimension based on the initial cross-dimension enabling signal of the height dimension to obtain the first coordinate calculation result of the depth dimension within the first operation cycle.
[0162] The first target dimension can be the batch dimension, and its batch coordinate is denoted as N. In the initial operation cycle, since the initial cross-dimension enabling signal of the depth dimension indicates whether the initial updated coordinate of the depth dimension can cause cross-dimension in the batch dimension, in the first operation cycle after the initial operation cycle, based on the initial cross-dimension enabling signal of the depth dimension, it is determined whether to accumulate the batch dimension step size on the initial coordinate calculation result of the batch dimension, so as to obtain the first coordinate calculation result of the batch dimension in the first operation cycle.
[0163] In the above implementation manner, in the first operation cycle after the initial operation cycle, by determining whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross-dimension enabling signal, the signal-driven mechanism is used to dynamically coordinate the multi-dimensional coordinate update, reduce the number of adders used, and thus reduce the occupied area of the subtractors.
[0164] In some embodiments, the size of the multi-dimensional data in each dimension is denoted as the dimension size. Please refer to Figure 2 , determining whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross-dimension enabling signal, and obtaining the first coordinate calculation result of the first target dimension in the first operation cycle, including:
[0165] S210. If the initial cross-dimension enabling signal indicates that the initial updated coordinate causes cross-dimension, accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension to obtain the first updated coordinate of the first target dimension.
[0166] S220. Generate the first cross-dimension enabling signal in the first operation cycle based on the first updated coordinate to determine the first coordinate calculation result.
[0167] Among them, the first updated coordinate is an intermediate calculation result of accumulating the dimension step size on the initial coordinate calculation result in the first target dimension in the first operation cycle. The first cross-dimension enabling signal is used to indicate whether the first updated coordinate causes cross-dimension in the adjacent higher dimension of the first target dimension. The first cross-dimension enabling signal is generated by a comparator. If the first updated coordinate of the first target dimension exceeds the size of the first target dimension, the first cross-dimension enabling signal can be pulled high, and the jump from the first target dimension to the adjacent higher dimension is triggered through this first cross-dimension enabling signal.
[0168] Specifically, if the initial cross - dimensional enabling signal indicates that the coordinates after the initial update cause cross - dimension, within the first operation cycle, for the initial coordinate calculation result of the first target dimension, the dimension step of the corresponding dimension is accumulated on the initial coordinate calculation result of the first target dimension to obtain the first updated coordinate of the first target dimension. Then, it is continuously determined whether the first updated coordinate of the first target dimension exceeds the size of the first target dimension, and a first cross - dimensional enabling signal is generated. The first cross - dimensional enabling signal is used to indicate whether the coordinate update of the first target dimension causes the coordinate adjustment of the adjacent higher dimension. Further, since it is indicated by the first cross - dimensional enabling signal whether cross - dimensional coordinate update is caused after increasing the corresponding dimension step in the first target dimension, the first coordinate calculation result of the first target dimension in the first operation cycle is determined according to the first cross - dimensional enabling signal.
[0169] Exemplarily, the first coordinate calculation result includes the height updated coordinate h_coord_c1 within the first operation cycle. The first cross - dimensional enabling signal includes the height enabling signal h_cross_c1. The first target dimension is the height dimension, the initial coordinate calculation result of the height dimension in the initial operation cycle is h_coord_c0, the dimension step on the height dimension is h_stride, and the dimension size on the height dimension is tensor_h. If the height enabling signal h_cross_c1 is high, the height updated coordinate h_coord_c1 within the first operation cycle is equal to h_coord_c0 + h_stride - tensor_h. If the height enabling signal h_cross_c1 is low, the height updated coordinate h_coord_c1 within the first operation cycle is equal to h_coord_c0 + h_stride.
[0170] Exemplarily, the first coordinate calculation result includes the depth updated coordinate d_coord_c1 within the first operation cycle. The first cross - dimensional enabling signal includes the depth enabling signal d_cross_c1. The first target dimension is the depth dimension, the initial coordinate calculation result of the depth dimension in the initial operation cycle is d_coord_c0, the dimension step on the depth dimension is d_stride, and the dimension size on the depth dimension is tensor_d. If the depth enabling signal d_cross_c1 is high, the depth updated coordinate d_coord_c1 within the first operation cycle is equal to d_coord_c0 + d_stride - tensor_d. If the depth enabling signal d_cross_c1 is low, the depth updated coordinate d_coord_c1 within the first operation cycle is equal to d_coord_c0 + d_stride.
[0171] In some embodiments, determining whether to accumulate the dimension step of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross-dimension enabling signal to obtain the first coordinate calculation result of the first target dimension in the first operation period may include: if the initial cross-dimension enabling signal indicates that the coordinate after the initial update does not cause cross-dimension, keeping the initial coordinate calculation result of the first target dimension unchanged and using it as the first coordinate calculation result.
[0172] Specifically, if the initial cross-dimension enabling signal indicates that the coordinate after the initial update does not cause cross-dimension, then the coordinate value of the first target dimension does not need to be updated. Therefore, the first coordinate calculation result of the first target dimension adopts the initial coordinate calculation result of the first target dimension.
[0173] In some embodiments, generating the first cross-dimension enabling signal in the first operation period based on the first updated coordinate may include: performing a cross-dimension judgment on the adjacent higher dimension of the first target dimension by using the first updated coordinate of the first target dimension and the dimension size on the first target dimension to generate the first cross-dimension enabling signal.
[0174] Specifically, comparing the first updated coordinate of the first target dimension and the dimension size on the first target dimension, and generating the first cross-dimension enabling signal according to the comparison result to indicate whether the first updated coordinate of the first target dimension causes cross-dimension on the adjacent higher dimension of the first target dimension, that is, determining whether to trigger the logical operation of updating the adjacent higher dimension coordinate. For example, if the first updated coordinate of the first target dimension is greater than the dimension size on the first target dimension, the first updated coordinate of the first target dimension causes cross-dimension on the adjacent higher dimension. If the first updated coordinate of the first target dimension is not greater than the dimension size on the first target dimension, the first updated coordinate of the first target dimension does not cause cross-dimension on the adjacent higher dimension.
[0175] Exemplarily, the first target dimension is height, the height coordinate is denoted as H1, the adjacent higher dimension of the first target dimension is depth D. The dimension size of the first target dimension is denoted as H0. The height step of the first target dimension is denoted as H_stride. The first updated coordinate of the first target dimension is equal to H1 + H_stride. Comparing the size of H1 + H_stride and H0, and correspondingly generating the first cross-dimension enabling signal. For example, the first cross-dimension enabling signal is denoted as h_cross_c1. If H1 + H_stride is greater than H0, pulling up the first cross-dimension enabling signal, that is, h_cross_c1 = 1. Otherwise, the first cross-dimension enabling signal remains invalid, that is, h_cross_c1 = 0.
[0176] In the above embodiments, by comparing the first updated coordinates of the first target dimension and the dimension size on the first target dimension, cross-dimension judgment is performed on adjacent high dimensions, and the cross-dimension judgment result is converted into a first cross-dimension enable signal that can be executed by hardware to coordinate multi-dimensional coordinate updates, thereby realizing the dynamic trigger of cross-dimension operations and ensuring data continuity.
[0177] In some embodiments, the lowest dimension among the multiple first target dimensions is denoted as the second lowest dimension. The method may further include: in the second operation cycle after the first operation cycle, determining whether to accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal to obtain the second coordinate calculation result of the second target dimension in the second operation cycle.
[0178] Wherein, the second target dimension is any dimension among the multiple first target dimensions except the second lowest dimension. If the data storage format of four-dimensional data adopts NDHW. The first lowest dimension is usually the width dimension (W), indicating that the data is stored continuously by rows. The second lowest dimension is usually the height dimension (H). The second target dimension is the depth dimension (D) or the batch dimension (N). The second operation cycle (cycle2) after the first operation cycle may refer to the third calculation stage in the hardware pipeline that follows the first operation cycle (cycle1) and is used to process the coordinate update of the second target dimension. The input of the second operation cycle is the output of the first operation cycle, forming a pipeline processing chain. The second coordinate calculation result includes the calculation result obtained after the calculation of the first coordinate calculation result of the second target dimension in the second operation cycle.
[0179] Specifically, if the second target dimension is the depth dimension, if the first cross-dimension enable signal of the height dimension indicates that the first updated coordinates of the height dimension can cause cross-dimension in the depth dimension, therefore, it is determined whether to accumulate the dimension step d_stride of the depth dimension on the first coordinate calculation result of the depth dimension based on the first cross-dimension enable signal of the height dimension to obtain the second coordinate calculation result of the depth dimension in the second operation cycle.
[0180] If the second target dimension is the batch dimension, if the first cross-dimension enable signal of the depth dimension indicates that the first updated coordinates of the depth dimension can cause cross-dimension in the batch dimension, therefore, it is determined whether to accumulate 1 on the first coordinate calculation result of the batch dimension based on the first cross-dimension enable signal of the depth dimension to obtain the second coordinate calculation result of the batch dimension in the second operation cycle.
[0181] In the above embodiments, in the second operation period after the first operation period, by determining whether to accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal, the multi-dimensional coordinate update is dynamically coordinated by a signal driving mechanism, reducing the number of adders used, thereby reducing the occupied area of subtractors.
[0182] In some embodiments, refer to Figure 3 , determining whether to accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal, and obtaining the second coordinate calculation result of the second target dimension in the second operation period, including:
[0183] S310. If the first cross-dimension enable signal indicates that the first updated coordinate causes cross-dimension, accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension to obtain the second updated coordinate of the second target dimension.
[0184] S320. Generate a second cross-dimension enable signal in the second operation period based on the second updated coordinate to determine the second coordinate calculation result.
[0185] Among them, the second cross-dimension enable signal is used to indicate whether the second updated coordinate causes cross-dimension on the adjacent higher dimension of the second target dimension.
[0186] Among them, the second updated coordinate is an intermediate calculation result of accumulating the dimension step on the first coordinate calculation result of the second target dimension in the second operation period. The second cross-dimension enable signal is used to indicate whether the second updated coordinate causes cross-dimension on the adjacent higher dimension of the second target dimension. The second cross-dimension enable signal is generated by a comparator. If the second updated coordinate of the second target dimension exceeds the size of the second target dimension, the second cross-dimension enable signal can be pulled high, and a jump from the second target dimension to its adjacent higher dimension is triggered by this second cross-dimension enable signal.
[0187] Specifically, if the first cross-dimension enable signal indicates that the first updated coordinate causes cross-dimension, in the second operation period, for the first coordinate calculation result of the second target dimension, accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension to obtain the second updated coordinate of the second target dimension. Continue to determine whether the second updated coordinate of the second target dimension exceeds the size of the second target dimension, and generate a second cross-dimension enable signal. The second cross-dimension enable signal is used to indicate whether the coordinate update of the second target dimension causes the coordinate adjustment of its adjacent higher dimension. Further, since it is indicated by the second cross-dimension enable signal whether cross-dimension coordinate update is caused after increasing the corresponding dimension step on the second target dimension, the second coordinate calculation result of the second target dimension in the second operation period is determined according to the second cross-dimension enable signal.
[0188] Exemplarily, the second coordinate calculation result includes the depth-updated coordinate d_coord_c2 in the second operation period and the batch-updated coordinate n_coord_c2 in the second operation period. The second cross-dimension enable signal includes the depth enable signal d_cross_c2. The second target dimension is the depth dimension, and the first coordinate calculation result of the depth dimension in the first operation period is d_coord_c1. The first coordinate calculation result of the batch dimension in the first operation period is n_coord_c1. The dimension step size on the depth dimension is d_stride. The dimension size on the depth dimension is tensor_d. If the depth enable signal d_cross_c2 is high, the depth-updated coordinate d_coord_c2 in the second operation period is equal to d_coord_c1 + d_stride - tensor_d. Further, the batch-updated coordinate n_coord_c2 in the second operation period is equal to n_coord_c1 + 1.
[0189] If the depth enable signal d_cross_c2 is low, the depth-updated coordinate d_coord_c2 in the second operation period is equal to d_coord_c1 + d_stride. Further, the batch-updated coordinate n_coord_c2 in the second operation period is equal to n_coord_c1.
[0190] In some embodiments, determining whether to accumulate the dimension step size of the corresponding dimension on the first coordinate calculation result of the second target dimension to obtain the second coordinate calculation result of the second target dimension in the second operation period according to the first cross-dimension enable signal may include: if the first cross-dimension enable signal indicates that the first updated coordinate does not cause cross-dimension, keeping the first coordinate calculation result of the second target dimension unchanged and using it as the second coordinate calculation result.
[0191] Specifically, if the first cross-dimension enable signal indicates that the first updated coordinate does not cause cross-dimension, the coordinate value of the second target dimension does not need to be updated. Therefore, the second coordinate calculation result of the second target dimension uses the first coordinate calculation result of the second target dimension.
[0192] In some embodiments, generating the second cross-dimension enable signal in the second operation period based on the second updated coordinate may include: performing a cross-dimension judgment on the second updated coordinate of the second target dimension and the dimension size on the second target dimension in the adjacent higher dimension of the second target dimension to generate the second cross-dimension enable signal.
[0193] Specifically, compare the second updated coordinate of the second target dimension and the dimension size on the second target dimension, and generate a second cross-dimension enabling signal according to the comparison result to indicate whether the second updated coordinate of the second target dimension causes a cross-dimension on the adjacent higher dimension of the second target dimension, that is, determine the logical operation of whether to trigger the update of the coordinates of its adjacent higher dimension. For example, if the second updated coordinate of the second target dimension is greater than the dimension size on the second target dimension, the second updated coordinate of the second target dimension causes a cross-dimension on its adjacent higher dimension. If the second updated coordinate of the second target dimension is not greater than the dimension size on the second target dimension, the second updated coordinate of the second target dimension does not cause a cross-dimension on the adjacent higher dimension.
[0194] Exemplarily, the second target dimension is depth, the depth coordinate is denoted as D1, the adjacent higher dimension of the second target dimension is batch. The size of the second target dimension is denoted as D0. The depth stride of the second target dimension is denoted as D_stride. The second updated coordinate of the second target dimension is equal to D1 + D_stride. Compare the size of D1 + D_stride and D0, and correspondingly generate a second cross-dimension enabling signal. For example, the second cross-dimension enabling signal is denoted as d_cross_c2. If D1 + D_stride is greater than D0, raise the second cross-dimension enabling signal, that is, d_cross_c2 = 1. Otherwise, the second cross-dimension enabling signal remains invalid, that is, d_cross_c2 = 0.
[0195] In the above embodiments, by comparing the second updated coordinate of the second target dimension and the dimension size on the second target dimension, cross-dimension judgment is performed on the adjacent higher dimension, and the cross-dimension judgment result is converted into a second cross-dimension enabling signal executable by hardware to coordinate the update of multi-dimensional coordinates, thereby realizing the dynamic trigger of cross-dimension operations and ensuring data continuity.
[0196] Please refer to Figure 4 , in the embodiments of the present application, there is also provided a physical address generation device, and the generation device includes:
[0197] A coordinate data determination module 410, configured to determine the starting coordinates and coordinate increment data of multi-dimensional data in each dimension;
[0198] A coordinate signal determination module 420, configured to, in the initial operation period, add the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinate of each dimension and the initial cross-dimension enabling signal in the initial operation period; wherein, the initial cross-dimension enabling signal is used to indicate whether the initial updated coordinate causes a cross-dimension on the adjacent higher dimension of each dimension;
[0199] The physical address generation module 430 is configured to perform coordinate calculation and address conversion based on the initially updated coordinates and the initial cross-dimension enable signal to obtain the physical address of the multi-dimensional data.
[0200] In an embodiment of the present application, a processor is further provided. The processor includes a logic circuit and a power supply circuit. The power supply circuit is configured to supply power to the logic circuit, and the logic circuit is configured to execute the steps of the method in any one of the above embodiments.
[0201] In an embodiment of the present application, a chip is further provided. The chip includes the processor in the above embodiment.
[0202] The further function descriptions of the above respective modules are the same as those in the corresponding above embodiments and will not be elaborated herein. The implementation device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0203] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the sake of convenience of description, when describing the above devices, they are divided into various units according to functions and described separately. Of course, when implementing the present application, the functions of the respective units can be implemented in one or more software and / or hardware.
[0204] Those skilled in the art should understand that the embodiments of the present application can be provided as a method or a device. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0205] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems) according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0206] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0207] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the related content.
[0208] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
[0209] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for generating a physical address, characterized in that, Applied to a hardware accelerator, the method includes: Determine the starting coordinates and coordinate increment data of the multi-dimensional data in each dimension; In an initial operation cycle, add the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinates of each dimension and the initial cross-dimension enable signal in the initial operation cycle; wherein, the initial cross-dimension enable signal is used to indicate whether the initial updated coordinates cause cross-dimension in the adjacent higher dimensions of each dimension; Based on the initial updated coordinates and the initial cross-dimension enable signal, perform coordinate calculation and address conversion to obtain the physical address of the multi-dimensional data.
2. The method according to claim 1, characterized in that, The performing coordinate calculation and address conversion based on the initial updated coordinates and the initial cross-dimension enable signal to obtain the physical address of the multi-dimensional data includes: In the initial operation cycle, perform coordinate calculation based on the initial cross-dimension enable signal and the initial updated coordinates to obtain the initial coordinate calculation results of each dimension in the initial operation cycle; In the first operation cycle after the initial operation cycle, perform coordinate calculation on the initial coordinate calculation results based on the initial cross-dimension enable signal to obtain the first coordinate calculation results of the first operation cycle; Perform coordinate calculation and address conversion according to the initial coordinate calculation results and the first coordinate calculation results to obtain the physical address of the multi-dimensional data.
3. The method according to claim 2, wherein A first cross-dimension enable signal is generated in the first operation cycle; the performing coordinate calculation and address conversion according to the initial coordinate calculation results and the first coordinate calculation results to obtain the physical address of the multi-dimensional data includes: In the second operation cycle after the first operation cycle, perform coordinate calculation on the first coordinate calculation results based on the first cross-dimension enable signal to obtain the second coordinate calculation results of the second operation cycle; In the second operation cycle, perform summarization and address conversion based on the initial coordinate calculation results, the first coordinate calculation results, and the second coordinate calculation results to obtain the physical address of the multi-dimensional data.
4. The method according to claim 2, characterized in that, The size of the multi-dimensional data in each dimension is denoted as the dimension size; the performing coordinate calculation based on the initial cross-dimension enable signal and the initial updated coordinates to obtain the initial coordinate calculation results of each dimension in the initial operation cycle includes: If the initial cross-dimension enable signal indicates that the initial updated coordinates cause cross-dimension, subtract the dimension size of the corresponding dimension from the initial updated coordinates to obtain the initial coordinate calculation results; If the initial cross-dimension enable signal indicates that the initial updated coordinates do not cause cross-dimension, use the initial updated coordinates as the initial coordinate calculation results.
5. The method according to claim 1, characterized in that The size of the multi-dimensional data in each dimension is denoted as the dimension size; the adding the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinates of each dimension and the initial cross-dimension enable signal in the initial operation cycle includes: Increment the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinate of each dimension; Use the initial updated coordinate of each dimension and the dimension size on the corresponding dimension to perform cross-dimension judgment on the adjacent higher dimension to generate the initial cross-dimension enable signal.
6. The method according to claim 2, characterized in that, The lowest dimension among all dimensions of the multi-dimensional data is denoted as the first lowest dimension; the step size of the multi-dimensional data on each dimension is denoted as the dimension step size; the coordinate calculation of the initial coordinate calculation result based on the initial cross-dimension enable signal to obtain the first coordinate calculation result of the first operation period includes: Determine whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross-dimension enable signal to obtain the first coordinate calculation result of the first target dimension within the first operation period; wherein, the first target dimension is any dimension other than the first lowest dimension among all dimensions.
7. The method according to claim 6, wherein The size of the multi-dimensional data on each dimension is denoted as the dimension size; the determining whether to accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension according to the initial cross-dimension enable signal to obtain the first coordinate calculation result of the first target dimension within the first operation period includes: If the initial cross-dimension enable signal indicates that the initial updated coordinate causes cross-dimension, accumulate the dimension step size of the corresponding dimension on the initial coordinate calculation result of the first target dimension to obtain the first updated coordinate of the first target dimension; Generate the first cross-dimension enable signal within the first operation period based on the first updated coordinate to determine the first coordinate calculation result; wherein, the first cross-dimension enable signal is used to indicate whether the first updated coordinate causes cross-dimension on the adjacent higher dimension of the first target dimension.
8. The method according to claim 7, wherein The generating the first cross-dimension enable signal within the first operation period based on the first updated coordinate includes: Use the first updated coordinate of the first target dimension and the dimension size on the first target dimension to perform cross-dimension judgment on the adjacent higher dimension of the first target dimension to generate the first cross-dimension enable signal.
9. The method according to claim 7, characterized in that, The lowest dimension among multiple first target dimensions is denoted as the second lowest dimension; the method further includes: In the second operation period after the first operation period, determine whether to accumulate the dimension step size of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal to obtain the second coordinate calculation result of the second target dimension within the second operation period; wherein, the second target dimension is any dimension other than the second lowest dimension among multiple first target dimensions.
10. The method according to claim 9, wherein The determining whether to accumulate the dimension step size of the corresponding dimension on the first coordinate calculation result of the second target dimension according to the first cross-dimension enable signal to obtain the second coordinate calculation result of the second target dimension within the second operation period includes: If the first cross - dimensional enabling signal indicates that the first updated coordinate causes cross - dimension, accumulate the dimension step of the corresponding dimension on the first coordinate calculation result of the second target dimension to obtain the second updated coordinate of the second target dimension; Generate a second cross - dimensional enabling signal within the second operation period based on the second updated coordinate to determine the second coordinate calculation result; wherein, the second cross - dimensional enabling signal is used to indicate whether the second updated coordinate causes cross - dimension on the adjacent higher dimension of the second target dimension.
11. The method according to claim 10, wherein The generating the second cross - dimensional enabling signal within the second operation period based on the second updated coordinate includes: Perform cross - dimension judgment on the adjacent higher dimension of the second target dimension by using the second updated coordinate of the second target dimension and the dimension size on the second target dimension to generate the second cross - dimensional enabling signal.
12. A physical address generation device, characterized in that Applied to a hardware accelerator, the device includes: A coordinate data determination module, configured to determine the starting coordinates and coordinate increment data of multi - dimensional data in each dimension; A coordinate signal determination module, configured to, within an initial operation period, add the coordinate increment data of each dimension to the starting coordinate of the corresponding dimension to obtain the initial updated coordinate of each dimension and the initial cross - dimensional enabling signal within the initial operation period; wherein, the initial cross - dimensional enabling signal is used to indicate whether the initial updated coordinate causes cross - dimension on the adjacent higher dimension of each dimension; A physical address generation module, configured to perform coordinate calculation and address conversion based on the initial updated coordinate and the initial cross - dimensional enabling signal to obtain the physical address of the multi - dimensional data.
13. A processor, characterized in that, The processor includes a logic circuit and a power supply circuit, the power supply circuit is used to supply power to the logic circuit, and the logic circuit is used to execute the steps of the method according to any one of claims 1 to 11.
14. A chip, characterized in that, The chip includes the processor according to claim 13.