Data Loading Method, Data Storage Method, Processor, Electronic Device, and Medium

Through pixel-by-pixel point-by-pixel data loading and NDHWC format storage, the problem of low data access efficiency in parallel processors is solved, and more efficient data handling and computing performance improvement is achieved.

CN120123264BActive Publication Date: 2025-07-15SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607240.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-15
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, data loading and storage methods are inefficient in parallel processors, especially in convolutional operations, which cannot meet the needs of fast and efficient data access, which affects computing performance.

Method used

A data loading method is adopted to load the pending tensor from the original tensor in memory into the cache area by pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pixel point-by-pix

Benefits of technology

It improves the data memory access bandwidth and hardware utilization of computing units, reduces memory access time, and improves the overall performance of the processor, especially in convolutional operations, which significantly improves the computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123264B_ABST
    Figure CN120123264B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a medium. The data loading method is used to load a tensor to be processed from an original tensor in a memory into a buffer area, and includes: determining a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor, where the plurality of requests are used to sequentially obtain data from the original tensor with the starting coordinates as the starting point until the number of obtained data is equal to the number of pixels included in the tensor to be processed, and data loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a medium. Background Art

[0002] A tensor is a data structure of a multi-dimensional array. Tensor operations are widely used in processors such as parallel processors. For example, in the field of deep learning, the dimensions of the input data, the intermediate data processed during the deep learning process, and the output data are flexible and not exact. Therefore, an elastic data form is required to describe various types of data, and thus the concept of tensors is generated. In the field of deep learning, all data to be operated on can be stored and exist in the form of tensors. If the data is not in the form of tensors, it needs to be first converted into the data structure form of tensors. As an example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor itself can have any number of dimensions. For example, it can be represented as a 5-dimensional array.

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processing devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of parallel processors need to process a large amount of data, and this data is generally stored in the storage components of parallel processors. For example, the storage component can be a high-speed memory. Through data loading instructions, this data can be loaded from the storage component to the buffer for calculation, and through data storage instructions, the data in the buffer can be stored in the memory.

[0004] How to provide a fast and efficient data loading / storing method is crucial for the computing performance of the device. Summary of the Invention

[0005] Embodiments of the present disclosure provide a data loading method, a data storage method, a processor, an electronic device, and a medium, which are used to provide a fast and efficient data loading / storing method, improve the memory access bandwidth, and increase the hardware utilization rate of computing devices.

[0006] According to a first aspect of the present disclosure, there is provided a data loading method for loading a tensor to be processed from an original tensor in memory into a buffer. The original tensor has a data storage format of N(C / x)DHW(xC) in memory, and the tensor to be processed has a data storage format of NDHWC when stored in the buffer, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is regarded as a pixel, and the data accumulates step by step to higher dimensions. The channel number dimension does not count the number of pixels. The data loading method includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into the buffer to load the tensor to be processed into the buffer. When the size relationship between the tensor to be processed and the original tensor in the first dimension causes the tensor to be processed to not be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And when the data loaded by the request all belongs to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0007] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. When the first dimension is the C dimension or the first dimension is the N dimension and the size of the tensor to be processed in the C dimension is greater than x, the second dimension is the C dimension. When the first dimension is the N dimension and the size of the tensor to be processed in the C dimension is equal to x, the second dimension is the overall dimension higher than the N dimension.

[0008] According to some embodiments of the present disclosure, when the size of the tensor to be processed in the first dimension is not equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension is not equal to the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0009] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension, and it is determined to adopt a division of requests in the H dimension.

[0010] According to some embodiments of the present disclosure, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

[0011] According to some embodiments of the present disclosure, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

[0012] According to some embodiments of the present disclosure, the tensor to be processed is used for a convolution operation in a computing unit within a processor.

[0013] According to some embodiments of the present disclosure, each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data to the buffer, and the number of pixels to be acquired by the request. The request initial coordinate corresponding to the next request is determined based on the request initial coordinate corresponding to the previous request. The request initial coordinate corresponding to each request is used to determine the data read address, data write address, and the number of pixels to be acquired by the request. Among them, when sequentially determining the request initial coordinate corresponding to the request, the coordinate value of the request initial coordinate is incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0014] According to some embodiments of the present disclosure, multiple requests for loading a tensor to be processed are determined in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, including: determining a first request among the multiple requests and an initial state when the first request enters a state machine based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, where the first request indicates the number of pixels to be acquired by the request; and determining each request after the first request among the multiple requests using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed.

[0015] According to some embodiments of the present disclosure, each request includes a data read address for indicating a starting position for reading data from memory, a data write address for indicating a starting position for writing data to a buffer, and the number of pixels to be acquired by the request. Among them, determining a first request among the multiple requests and an initial state when the first request enters a state machine based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0016] According to some embodiments of the present disclosure, determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0017] According to some embodiments of the present disclosure, based on an initial state, in combination with the number of pixels to be acquired in a first request and the number of pixels included in a tensor to be processed, a state machine is utilized to determine each request among a plurality of requests after the first request, including: based on the initial state, the state machine is utilized to determine a request initial coordinate corresponding to a second request after the first request and the number of pixels to be acquired in the second request; according to the shape and size of an original tensor, the request initial coordinate corresponding to the second request, and the data storage format of the original tensor, a data read address of the second request is determined; according to the number of pixels to be acquired in the second request, a data write address of the second request is determined; the state machine is updated based on information related to the second request, and subsequent requests among each request are sequentially determined based on the updated state machine.

[0018] According to some embodiments of the present disclosure, for a plurality of requests for loading a tensor to be processed, determining a data read address of an nth request among the plurality of requests includes:

[0019] Calculating a data read address Addr1_n of the nth request according to the following formula:

[0020] Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0021] c_coord * tensor_d * tensor_h * tensor_w +

[0022] d_coord * tensor_h * tensor_w * xC +

[0023] h_coord * tensor_w * xC +

[0024] w_coord * xC

[0025] Wherein, u_addr_base represents the storage address in memory of a pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the original tensor in five dimensions respectively.

[0026] According to some embodiments of the present disclosure, for multiple requests for loading tensors to be processed, determining the data write address of the nth request among the multiple requests includes: for m requests with the same coordinate value range of the requested data in the C dimension, calculating the data write address Addr_2_i of the ith request sent among the m requests according to the following formula:

[0027] Addr_2_i = b_addr + req_size_1*copy_c + req_size_2*copy_c +…+ req_size_i-1*copy_c

[0028] wherein, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2,..., req_size_i-1 represent the number of pixels to be acquired by the respective previous i-1 requests, and copy_c is the size of the tensor to be processed in the C dimension; wherein, in response to the data of each request among the m requests including the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer, and in response to the data of each request among the m requests including the tth data element to the (t + x)th data element of the original tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula:

[0029] b_addr = b_addr_base+ (t / x - 1)*x

[0030] t, m, and i are positive integers.

[0031] According to some embodiments of the present disclosure, sequentially sending multiple requests and sequentially writing the data corresponding to each request into the buffer to load the tensor to be processed into the buffer includes: for any request, in response to all the data loaded by any request being located in the memory, sending any request to the memory; in response to all the data loaded by any request not being located in the memory, converting any request into writing multiple predetermined values into the buffer, wherein the number of the multiple predetermined values is determined by the number of pixels to be acquired by any request.

[0032] According to some embodiments of the present disclosure, the data loading method further includes: obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, where the boundary values are used to define the boundary range for acquiring data from the original tensor, and for the dimension with the boundary value set, the data loaded by each request belongs to the valid data range defined by the original tensor and the boundary values or does not belong to the valid data range.

[0033] According to a second aspect of the present disclosure, a data loading method is provided, including: receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in memory into a buffer, wherein the data storage format of the original tensor in memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is a pixel, and accumulates gradually to higher dimensions, where the channel number dimension does not calculate the number of pixels; and after parsing the data loading instruction, using an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into the buffer to load the tensor to be processed into the buffer, where in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as or adjacent to the second dimension.

[0034] According to a third aspect of the present disclosure, a data storage method is provided for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory. The data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in memory is NDHWC. Here, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the number of channels dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is regarded as a pixel, and it accumulates step by step to higher dimensions. The number of pixels is not calculated for the number of channels dimension. The data storage method includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine multiple requests for loading the second tensor. The multiple requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the multiple requests to write the data corresponding to each request into memory in sequence to store the second tensor in memory. When loading the second tensor due to the size relationship between the second tensor and the first tensor in the first dimension, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension. The data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And when the data loaded by the request all belongs to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer. Here, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

[0035] According to a fourth aspect of the present disclosure, there is provided a data storage method, including: receiving a data storage instruction indicating to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, where the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the channel number of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not calculate the number of pixels; and after parsing the data storage instruction, using an execution unit to execute the data storage instruction, where using the execution unit to execute the data storage instruction includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into memory to store the second tensor in the memory, where in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor, and in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0036] According to a fifth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the original tensor is stored in the memory in a data storage format of N(C / x)DHW(xC), and the tensor to be processed is stored in the buffer in a data storage format of NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not count the number of pixels; and the execution unit is configured to: execute the data loading instruction, wherein the execution unit executes the data loading instruction, including: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to write the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer, where in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, the requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0037] According to a sixth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory. The first tensor has a data storage format of N(C / x)DHW(xC) in the buffer, and the second tensor is stored in memory in a data storage format of NDHWC. Here, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions. The channel number dimension does not count the number of pixels; and the execution unit is configured to: execute the data storage instruction. When the execution unit executes the data storage instruction, it includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to write the data corresponding to each request into memory in sequence to store the second tensor into memory. When the size relationship between the second tensor and the first tensor in the first dimension causes that when loading the second tensor, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, the requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And when the data loaded by the request all belongs to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer. Here, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

[0038] According to a seventh aspect of the present disclosure, an electronic device is provided, including a processor and a memory connected to the processor. The processor includes a buffer. The processor is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor, they implement the data loading method according to the embodiments of the present disclosure to load a tensor to be processed from an original tensor in the memory into the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain a second tensor based on a first tensor in the buffer and write the second tensor into the memory.

[0039] According to an eighth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the data loading method according to an embodiment of the present disclosure or implement the data storage method according to an embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 FIG. shows a schematic structural diagram of a General-Purpose computing on Graphics Processing Unit (GPGPU);

[0042] Figure 2 FIG. shows a schematic structure of a tensor;

[0043] Figure 3A FIG. shows a schematic diagram of the data storage format of NDHWC;

[0044] Figure 3B FIG. shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32;

[0045] Figure 4 FIG. shows a schematic diagram of obtaining a tensor in the related art;

[0046] Figure 5 FIG. shows a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure;

[0047] Figure 6 FIG. shows a schematic diagram of continuously obtaining a tensor to be processed according to an embodiment of the present disclosure;

[0048] Figure 7A FIG. shows a schematic diagram of continuously obtaining a tensor to be processed in a continuous manner for the data storage format of NDHWC according to an embodiment of the present disclosure;

[0049] Figure 7B FIG. shows a schematic diagram of continuously obtaining a tensor to be processed in a continuous manner for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure;

[0050] Figure 8AShows a schematic diagram of setting boundary values for the original tensor according to an embodiment of the present disclosure;

[0051] Figure 8B Shows a schematic diagram of continuous data acquisition in the case of setting boundaries for the original tensor according to an embodiment of the present disclosure;

[0052] Figure 8C Shows a schematic diagram of continuously obtaining the tensor to be processed in a continuous manner according to the set step size according to an embodiment of the present disclosure;

[0053] Figure 9A Shows a schematic diagram of loading the original tensor of N(C / x)DHW(xC) into the tensor to be processed stored in NDHWC in a continuous manner according to an embodiment of the present disclosure;

[0054] Figure 9B Shows a schematic diagram of the states of the state machine provided according to some embodiments of the present disclosure;

[0055] Figure 9C Shows according to an embodiment of the present disclosure for Figure 9A The state transition schematic diagram of the example according to the PerH division method;

[0056] Figure 10 Shows a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure;

[0057] Figure 11 Shows a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure;

[0058] Figure 12 Shows a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure;

[0059] Figure 13 Shows a schematic block diagram of a processor according to some embodiments of the present disclosure;

[0060] Figure 14 Shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure;

[0061] Figure 15 Shows a block diagram of an example computing device implementing some embodiments of the present disclosure; and

[0062] Figure 16 Shows a schematic block diagram of a computer-readable storage medium according to some embodiments of the present disclosure. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0064] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before such a term cover the elements or items listed after such a term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, such relative positional relationships may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of some known functions and known components are omitted in the present disclosure.

[0065] Figure 1 A schematic structural diagram of a GPGPU is shown. As Figure 1 shown, a GPGPU is actually an array of programmable multi-processors. For example, the programmable multi-processors can be Streaming Processor Clusters (SPCs), for example, including Figure 1 the shown Streaming Processor Cluster 1,..., Streaming Processor Cluster M, where M is a positive integer. In a general-purpose graphics processor, one Streaming Processor Cluster processes one computing task, or multiple Streaming Processor Clusters process one computing task. Data sharing among multiple Streaming Processor Clusters is performed through a global cache or High Bandwidth Memory (HBM).

[0066] As Figure 1 shown, taking Streaming Processor Cluster 1 as an example, one Streaming Processor Cluster can include multiple Compute Units (CUs), for example Figure 1The computing units 1, 2, ..., K in it, where K is a positive integer. Each computing unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit may include multiple cores (also known as computing cores or computing kernels), and each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The computing core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in it) and shared memory, which are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit. In addition, the buffer can be understood as being used to share data among the computing units within the streaming processor cluster.

[0067] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processing unit (or parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in it). All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0068] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in it) schedules and allocates the thread bundles so that the multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, the multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. The multiple threads in each thread bundle will execute the same instructions. The memory execution instructions will be issued to the shared memory in the computing unit or further issued to the intermediate-level cache or global cache or high-bandwidth memory for read / write operations, etc.

[0069] As Figure 1 shown, general computing operations, such as computing operations on matrices in the field of artificial intelligence, usually require a large amount of data. These data are usually stored in a memory, such as in a high-bandwidth memory HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing results, data needs to be stored in the memory (Store operation). The storage method of data in the memory will affect the memory access bandwidth, and thus affect the hardware utilization rate of the computing unit.

[0070] For example, general computing operations include General Matrix Multiplication (GEMM). As an example, the data for general matrix multiplication is represented as two 5-dimensional arrays, such as the matrix multiplication calculation of two tensors A and tensor B. In addition, general computing operations also include convolution operations, which are manifested as data dot products. It can be understood that in the field of artificial intelligence, there are other computing operations, which will not be listed one by one here. The data involved in these calculations is usually embodied in the form of tensors.

[0071] For example, for a certain tensor A in the buffer, its shape dimensions can be represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor data in 5 dimensions, and a1, a2, a3, a4, a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, N represents the batch dimension, which is the number of data samples grabbed in one training. The D dimension represents the depth dimension, the H dimension represents the height dimension of the input data, the W dimension represents the width dimension of the input data, and the C dimension represents the number of channels dimension. For example, taking tensor A as an example, a1 can be the N dimension size, a2 can be the D dimension size, a3 can be the H dimension size, a4 can be the W dimension size, and a5 can be the C dimension size. Of course, the present disclosure does not make specific limitations on this.

[0072] As an example, Figure 2 shows a schematic structure of a tensor. In Figure 2 the shown tensor, a1 is the N dimension size and equals 1, a2 is the D dimension size and equals 1, a3 is the H dimension size and equals 5, a4 is the W dimension size and equals 4, and a5 is the C dimension size and equals 64. For example, Figure 2 the pixel elements of the tensor in

[0073] The placement of tensors in the memory (such as memory or buffer) can have various formats, called data storage formats (layout). The data storage format is used to indicate the storage order and dimension arrangement of tensors in the storage component. The following uses the Figure 2 shown tensor to describe different data storage formats.

[0074] In the related art, the data storage format can include NDHWC, also known as the Linear mode. Figure 3A shows a schematic diagram of the NDHWC data storage format.

[0075] For example, for the NDHWC linear mode, as Figure 3A shown, from the first channel (a5 = 0,Figure 3A the first element of c0 in Figure 3A starting from the element 0 in Figure 3A the first element of c1 in Figure 3A the element 20 in Figure 3A and so on until the first elements of all channels are laid out, for example, until the first element of the 64th channel (a5 = 63, Figure 3A the element 1260 in c63), after which the second element of the first channel (a5 = 0, Figure 3A the second element of c0 in Figure 3A the element 1 in Figure 3A is stored, then the second element of the second channel (a5 = 1, Figure 3A the second element of c1 in

[0076] In the related art, the data storage format may also include N(C / x)DHW(xC), also known as the Interleave mode, where x can be set to 8, 16, 32, etc. as needed.

[0077] The N(C / x)DHW(xC) data storage format is similar to the NDHWC data storage format, but there is a key difference. In the layout of N(C / x)DHW(xC), a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.

[0078] Figure 3B shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32.

[0079] As Figure 3B shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 3B c0 in Figure 3B to a5 = 31 (

[0080] c31 in

[0081] It can be understood that in the related art and possible future developments, the data storage format of tensors is not limited to the above-described two data storage formats of N(C / x)DHW(xC) and NDHWC. Further, for multiple tensors that can be stored in the memory and cache during the calculation process, generally the storage space of the memory is much larger than that of the cache, but it is farther from the calculation unit, and the data transfer efficiency is lower than that of the cache.

[0082] In the related art, during the calculation process of a processing device, a large amount of calculation data will be generated, for example, in the form of tensors, which can be temporarily stored in the cache. For example, the cache here can refer to Figure 1 the cache (buffer) in the streaming processor cluster shown in, and further, these data can also be transferred from the cache or directly stored in the memory. For example, the memory can be Figure 1 the high-bandwidth memory HBM shown in. The storage form of tensors in the memory and cache can be, for example, any one of the above-described N(C / x)DHW(xC) and NDHWC data storage formats. Thus, during the calculation process, according to factors such as technical requirements and the respective storage characteristics of the cache and memory, a large amount of data transfer processes need to be performed between the two. For example, tensors in the cache are stored in the memory through storage instructions, or tensors in the memory are loaded into the cache through load instructions.

[0083] It can be understood that in this article, the data storage process of storing tensors in the cache in the memory and the data loading process of loading tensors in the memory into the cache can be implemented in a similar manner. Therefore, for the sake of convenience in description, in some embodiments or examples, only the data loading process is described as an example, and those skilled in the art can apply it similarly to the data storage process. For the differences between the two, separate descriptions will be given.

[0084] As an example, Figure 4 shows a schematic diagram of obtaining tensors in the related art. As Figure 4 shown, tensor A is stored in the memory, and it can be placed in the memory in the above-mentioned N(C / x)DHW(xC) or NDHWC data storage format. That is, tensor A is a 5D array. In Figure 4 only three dimensions of W, H, and C are schematically shown, and schematically, on the left side of Figure 4 a dimension coordinate system of tensor A is shown, where they are the W dimension, H dimension, and C dimension respectively. During the data loading process, all or part of the data in tensor A can be loaded into the cache through a load instruction. For example, Figure 4The tensor B in it is loaded into the buffer as a whole. Specifically, the loading instruction can indicate the first starting point of the tensor A to be loaded (C = 0, W = 0, H = 0), the second starting point of the tensor B in the coordinate system of the tensor A (C = 0, W = 3, H = 0), and the dimensions of the tensor B in each dimension. Schematically, in Figure 4 In the 3D schematic diagram shown, the tensor B is a cuboid determined by the above second starting point and the dimensions in each dimension. Through the information about the second starting point and the dimensions in each dimension of the tensor to be loaded, the memory can load the tensor B into the buffer.

[0085] The whole-piece data loading method adopted in the above related technologies can be applied to the general matrix multiplication GEMM calculation process moment in general computing operations. However, the above whole-piece data loading method is not applicable to the convolution operation, whose calculation feature is data dot product. The above whole-piece data loading method will limit the calculation efficiency of this kind of per-pixel point type convolution operation, increase the data access time, and reduce the overall performance of the processor, which limits the further development space of efficient and general-purpose processors.

[0086] Furthermore, in the memory, the original tensor is continuously stored in the memory in the data storage format as described above. The shape of the original tensor is represented as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the dimensions of the original tensor in these 5 dimensions and are all positive integers. When extracting or storing a partial tensor in the original tensor, since it is a partial tensor inside the original tensor, it may not be possible to continuously extract in each dimension during extraction, and it is impossible to accurately obtain the amount of data and the data position when loading or storing data to optimally implement data loading or storing, which greatly reduces the data bandwidth when loading data and reduces the hardware calculation efficiency.

[0087] In view of the above technical problems in the related art, the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. In the present disclosure, first, tensors are no longer fetched and transferred in a whole block, but are fetched and transferred sequentially in units of pixel points to be applicable to computational operations such as convolution operations. Further, for the operation of loading a tensor to be processed from the original tensor in the memory to the buffer, multiple requests for loading the tensor to be processed are determined according to the number of pixels included in the tensor to be loaded, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, so as to use the determined multiple requests to sequentially fetch data from the original tensor starting from the starting coordinates until the number of fetched data is equal to the number of pixels included in the tensor to be processed, thereby improving the data transfer efficiency for the data loading operation, reducing the memory access time, and improving the overall performance of the processor.

[0088] In the data loading method provided by at least one embodiment of the present disclosure, requests are divided according to whether tensors can be continuously loaded in the first dimension. The data for each request to load belongs to the data range of the original tensor or does not belong to the data range of the original tensor. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.

[0089] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0090] The data loading method provided by at least one embodiment of the present disclosure is used to load a tensor to be processed from the original tensor in the memory to the buffer. The memory according to the embodiments of the present disclosure may be, for example, a high-bandwidth memory (HBM), and the buffer is, for example, a buffer in a streaming processor cluster, which is not limited herein.

[0091] In the method according to the embodiments of the present disclosure, the data storage format of the original tensor in the memory is different from the data storage format of the tensor to be processed in the buffer. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. For example, the storage component refers to the above-mentioned HBM or buffer. That is to say, in the data loading method according to the embodiments of the present disclosure, the placement manner of the original tensor in the memory is different from the placement manner of the tensor to be loaded in the buffer next. Specifically, the original tensor from which data is to be fetched is placed in the memory in an N(C / x)DHW(xC) interleaved manner, and after being loaded into the buffer, it is changed to be placed in an NDHWC linear manner. The data loading / storage method proposed by the present disclosure is applicable to the situation where the data placement manner is changed during the data transfer process.

[0092] The original tensor in memory is a 5D tensor, and its shape dimensions are represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in five dimensions and are all positive integers. The five dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. Similarly, the shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in five dimensions and are all positive integers.

[0093] For the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension.

[0094] For the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension. Generally, x can be set to an integer multiple of 4.

[0095] Regarding the characteristics of the above two data storage formats, reference can be made to the description above in combination with Figure 3A - Figure 3B which will not be repeated here.

[0096] Furthermore, for the 5D data of the original tensor, the data determined by the height dimension and the width dimension represents a pixel (or, it can also be called an element), and it accumulates step by step to higher dimensions. Among them, the number of pixels is not calculated for the number of channels dimension. As an example, as Figure 4 shown, the values of each W dimension and H dimension can determine a pixel. For example, W = 0 and H = 0 correspond to the first pixel (or element) in tensor A, and W = 1 and H = 0 correspond to the second pixel in tensor A. In the tensor, the C dimension does not affect the number of pixels.

[0097] Figure 5 is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 5 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S101 - S103.

[0098] In step S101, obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Next, in step S102, combine the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine multiple requests for loading the tensor to be processed, where the multiple requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed. In step S103, sequentially send the multiple requests and sequentially write the data corresponding to each request into the buffer to load the tensor to be processed into the buffer. According to an embodiment of the present disclosure, the tensor to be processed may be used for a convolution operation in a computing unit within a processor. Herein, the number of data may also be expressed as the number of pixels.

[0099] In the data loading method according to an embodiment of the present disclosure, in order to obtain the tensor to be processed from the original tensor in the memory, it is necessary to indicate the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. For example, the starting coordinates may be the second starting point (C = 0, W = 3, H = 0) described in Figure 4 Furthermore, in the data loading method according to an embodiment of the present disclosure, it is also necessary to indicate the number of pixels included in the tensor to be processed (e.g., expressed as copy_pixel_num), that is, the total number of pixels in the tensor that is desired to be loaded into the buffer. This continuous data copying method is different from the block-based data loading method adopted in the related art above (where it is necessary to indicate the dimensions of the tensor to be processed in each dimension). According to the data loading method of an embodiment of the present disclosure, the range of the tensor to be processed is determined by the number of pixels to be obtained. Further, instead of obtaining the data in a block from the original tensor, the data is sequentially obtained from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed.

[0100] Next, first describe the implementation process of implementing the above continuous data loading based on copy_pixel_num and the starting coordinates, and then describe the implementation process of determining multiple requests for loading the tensor to be processed later.

[0101] As an example, Figure 6 shows a schematic diagram of obtaining the tensor to be processed in a continuous manner according to an embodiment of the present disclosure. In Figure 6 tensor A represents the original tensor in the memory. In Figure 6 each pixel point is shown as a square, and tensor A covers all the squares, that is, both the white squares and the gray-shaded squares. The starting point of the original tensor is represented as (C = 0, W = 0, H = 0), for example, to indicate the position of the original tensor in the memory. Next, in Figure 6Among them, tensor B represents the tensor to be processed. Its starting point is represented as (C = 0, W = 3, H = 0), that is, data is obtained starting from the 3rd pixel in the first row of tensor A. Then, in Figure 6 FIG. Figure 6 schematically shows the process of obtaining the tensor to be processed in a sequential (or also called continuous) manner according to an embodiment of the present disclosure. That is, as shown by the dashed arrow, starting from the starting point (C = 0, W = 3, H = 0), data is obtained pixel by pixel from tensor A until the number of obtained pixels reaches copy_pixel_num. In Figure 6 In the example of FIG. Figure 6 , only the three dimensions of W, H, and C are shown. It can be understood that if the data in these 3 dimensions still does not reach copy_pixel_num, data of higher dimensions can be further obtained, which is not limited here. In addition, as Figure 6 shown in FIG. Figure 6 , the C dimension itself does not affect the number of pixels. Assume that Figure 6 the data storage format of tensor A in memory in FIG. Figure 6 is NDHWC. In Figure 6 the example shown in FIG. Figure 6 , the total number of pixels of the tensor to be processed is copy_pixel_num = 27.

[0102] In an embodiment according to the present disclosure, obtaining data from the original tensor in sequence with the starting coordinate as the starting point includes: excluding the channel number dimension from the 5 dimensions of the original tensor, and in the order of dimensions from low to high, using the starting coordinate as the starting point for data acquisition, and sequentially obtaining data until the number of obtained pixels reaches copy_pixel_num. The reason for excluding the channel number dimension from the 5 dimensions is that for a tensor, the channel number dimension does not count the number of pixels. That is to say, S102 may include obtaining pixel data from the original tensor with the starting coordinate as the starting point for data acquisition in the order of the width dimension, height dimension, depth dimension, and batch dimension until the number of obtained pixels reaches copy_pixel_num.

[0103] Comparing Figure 4 and Figure 6 the two ways of obtaining the tensor to be processed shown in FIG. Figure 6 , the data loading method provided by the embodiment of the present disclosure can achieve data loading pixel by pixel, rather than Figure 4The acquisition of the monolithic form in [the above context] is such a data loading method that is more conducive to the calculation process such as convolution operations and helps improve the operation efficiency. Thus, based on the method provided in the embodiments of the present disclosure, for the data to be subjected to convolution operations next, the processor can, for example, indicate in the form of an instruction to fetch this part of the data from the memory to the buffer according to the above continuous data loading method for use in convolution operations. It can be understood that the above processor can reasonably use either the continuous data loading method or the monolithic data loading method according to the type of operation to be performed or the data processing characteristics, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or the buffer can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0104] According to some embodiments of the present disclosure, there are multiple tensors stored in the memory, and the data loading method can further include: obtaining indication information about the storage location of the original tensor in the memory. As an example, this indication information can include the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. It can be understood that the storage space of the memory is generally much larger than that of the buffer, and various tensor data generated during the calculation process can be stored therein. These tensor data can be placed in the memory in any one of the above N(C / x)DHW(xC) or NDHWC data storage formats. In the embodiments according to the present disclosure, in order for the memory to know the specific location of the tensor to be acquired in the memory, indication information about the storage location of the original tensor in the memory can also be obtained. This indication information includes the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. As an example, in the case where the original tensor is in the N(C / x)DHW(xC) data storage format, tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5 respectively represent the values of the original tensor in the 5 dimensions of N, C / x, D, H, and W.

[0105] Regarding the influence of the data storage format on the loading process of the tensor to be processed ( "tensor B"), it will be described below in combination with Figure 7A - Figure 7B which will be described. In combination with Figure 7A and Figure 7B In the example shown, the data storage method of the tensor to be processed in the buffer is the same as the data placement method of the original tensor in the memory.

[0106] Figure 7AShows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to the data storage format of NDHWC for the embodiments of the present disclosure. For Figure 7A the tensor shown, its data storage format in memory is in the form of NDHWC, or is arranged in the NDHWC manner.

[0107] In Figure 7A the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 8 (the C dimension is not shown in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Taking the N dimension as an example, N_dim = 2 corresponds to Figure 7A N0 and N1 in. Based on the above parameters, the total number of pixel points included in the original tensor can be obtained as 64.

[0108] Next, in Figure 7A the example of, the starting point coordinates of the tensor to be processed in the original tensor are expressed as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixel points included in the tensor to be processed copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7A the pixel points where W = 2 and H = 2 are located under the N = 0 and D = 0 dimensions shown in. Taking this pixel point as the starting point, according to the data arrangement method, in the order from low dimension to high dimension (that is, in the order of W, H, D, N, as shown by the data acquisition schematic arrow in Figure 7A ), data is sequentially acquired until the number of acquired data is equal to the number of pixel points included in the tensor to be processed. In Figure 7A the example shown, starting from the determined starting point, data is acquired pixel by pixel in the order of the arrow, so as to load this part of the data in the original tensor in memory into the buffer area as the tensor to be processed for subsequent operations such as convolution calculations.

[0109] Figure 7B Shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to the data storage format of N(C / x)DHW(xC) for the embodiments of the present disclosure. For Figure 7B the tensor shown, its data storage format in memory is in the form of N(C / x)DHW(xC), or is arranged in the N(C / x)DHW(xC) manner. In Figure 7BIn the example, x = 8, that is, every 8 C channels form a group and are bound to the W dimension. Compared with Figure 7A the data storage format of NDHWC shown in

[0110] In Figure 7B the shown example, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 16 (shown as two 8*C dimensions in the figure), the size of the W dimension is 4 (W_dim = 4), the size of the H dimension is 4 (H_dim = 4), the size of the D dimension is 2 (D_dim = 2), and the size of the N dimension is 2 (N_dim = 2). Thus, the total number of pixel points included in the original tensor can be obtained as 64.

[0111] Next, in Figure 7B the example, the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixel points included in the tensor to be processed copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7B the pixel points where W = 2 and H = 2 under N = 0 and D = 0 and the first 8*C dimension shown in Figure 7B . Starting from this pixel point, in the data arrangement manner, in the order from low dimension to high dimension (that is, in the order of W, H, D, C, N, as shown by the data acquisition indication arrow in Figure 7A ), data is sequentially acquired until the number of acquired data is equal to the number of pixel points included in the tensor to be processed. Compared with Figure 7B the acquisition order shown in Figure 7A and Figure 7B the example, since the C dimension has a higher rank, in Figure 7B the example, first, data of W, H, D, and the first 8*C dimension is acquired according to the starting point, and then, data of the second 8*C dimension is acquired in the order of W, H, D, and finally, data of the N dimension is acquired. It can be understood that in

[0112] In Figure 7BIn the example shown, starting from the determined starting coordinates, data is obtained pixel by pixel in the order of the arrows, for loading this part of the data in the original tensor in the memory into the buffer as the tensor to be processed, for subsequent operations such as convolution calculations.

[0113] In some embodiments according to the present disclosure, boundary values can also be defined for the original tensor in at least a part of five dimensions respectively. As an example, left and right boundary values can be set for, for example, the C dimension, the W dimension, the H dimension, and the D dimension, such as represented as (L - tensor_C, R - tensor_C), (L - tensor_W, R - tensor_W), (L - tensor_H, R - tensor_H), (L - tensor_D, R - tensor_D). The above - mentioned boundary values are used to define the boundary range for obtaining data from the original tensor, where the boundary values are arbitrary values compared to the size values of the original tensor in this dimension.

[0114] In combination Figure 7A and Figure 7B In the method described, the range of the original tensor (i.e., the boundary values) is not defined, that is, the tensor to be processed is obtained from the complete data of the original tensor. In the method according to the embodiments of the present disclosure, it is also proposed that the range for obtaining the tensor to be processed from the original tensor can be delimited by setting boundary values. Further, in the implementation process, the boundary values can be set as arbitrary values compared to the size values of the original tensor in this dimension. That is to say, the boundary values can exceed the range of the original tensor itself.

[0115] As an example, Figure 8A shows a schematic diagram of setting boundary values for the original tensor according to an embodiment of the present disclosure. As Figure 8A shown, the rectangular box represents the range covered by the original tensor, which can be any dimension in the original tensor, such as the C dimension, the W dimension, the H dimension, or the D dimension. Generally, boundary values are not set for the N dimension. In Figure 8A 's six sub - pictures, the relationships between the left boundary value and the right boundary value ( Figure 8A shown as "L" and "R" in Figure 8A ) and the size values of the original tensor in this dimension are respectively shown.

[0116] According to the embodiments of the present disclosure, when setting boundary values, sequentially obtaining data from the original tensor includes: for the data part of the original tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the original tensor defined by the boundary values. For example, referring to Figure 7A and Figure 7BThe described order. Comparatively, for the data part of the original tensor not covered by the range defined by the boundary values, it is represented as invalid data. For the invalid data, the buffer directly fills the to-be-processed tensor with a predetermined value, where the predetermined value is equal to 0.

[0117] As an example, in Figure 8A the first sub-picture, both the left boundary value L and the right boundary value R are on the left side of the original tensor, that is, all the data to be obtained in this dimension is invalid data. In this case, for the invalid data, this part of the data can be automatically filled, for example, by sending a zero-padding instruction to the buffer. Again, for example, in Figure 8A the second sub-picture, the left boundary value L is on the left side of the left boundary of the original tensor data range, and the right boundary value R is on the left side of the right boundary of the original tensor, that is, for the to-be-processed tensor to be obtained, part of it is invalid data and part of it is valid data in the original tensor. Schematically, in Figure 8A the data corresponding to the slanted shaded part is represented as valid data, and the rest is invalid data. In this case, for the valid data, for example, refer to Figure 7A and Figure 7B for the described order, while for the invalid data, this part of the data can be automatically filled, for example, by sending a zero-padding instruction to the buffer.

[0118] In the method according to the embodiments of the present disclosure, by setting boundary values for the original tensor, the range from which the to-be-processed tensor is to be taken can be further delimited. In practical applications, this implementation method can adapt to the characteristics of operations such as convolution. For example, it is beneficial to reduce the amount of calculation and greatly improve the flexibility of data. As an example, assume that the original tensor corresponds to an intermediate tensor for feature extraction of an entire input picture, and the input picture includes a specific target, such as an object to be recognized, and the object does not cover the entire picture, that is, the picture includes a background part. In this case, by setting boundary values, the range of the to-be-processed tensor to be obtained can be limited to the part of the original tensor corresponding to the specific target, so as to reduce the amount of calculation of subsequent operations such as convolution and improve the processing efficiency.

[0119] As an example, Figure 8B shows a schematic diagram of continuous data acquisition in the case where boundaries are set for the original tensor. Among them, compared with the situation shown in Figure 6 , Figure 8B can be understood as boundary values (bound) are respectively set for the C dimension, W dimension, and H dimension of the tensor A in Figure 6 , and the boundaries are all within the size ranges of the tensor A in the C dimension, W dimension, and H dimension, which is equivalent to the boundary situation shown in the fourth sub-picture in Figure 8A . Specifically, Figure 8BThe outer box as a whole in corresponds to tensor A (corresponding to the tensor A in Figure 6 ). After setting the boundary therein, according to the data loading method of the embodiments of the present disclosure, data will be sequentially acquired starting from the starting point coordinates within the set boundary range until the number of acquired data is equal to the number of pixels included in the tensor to be processed. In Figure 8B , the data part composed of squares corresponds to the data range framed by the boundary values set for the C dimension, W dimension, and H dimension, and data acquisition is sequentially performed from within the data range framed by the boundary. Data located outside the data range framed by the boundary can be regarded as invalid data. Regarding the process of sequentially acquiring data in Figure 8B , reference can be made to the description in conjunction with Figure 6 , and details will not be elaborated here.

[0120] In some embodiments according to the present disclosure, a step value for acquiring data can also be set. As an example, the set data step value is used to specify the step (Stride) for sequentially acquiring data from the original tensor. Among them, for the data of the original tensor, starting from the starting coordinates, sequentially acquiring data from the original tensor includes: for the data of the original tensor, starting from the starting coordinates, sequentially acquiring data from the original tensor according to the data step value.

[0121] Figure 8C shows a schematic diagram of sequentially acquiring the tensor to be processed in a continuous manner according to the set step in the embodiments of the present disclosure. Among them, the step is equal to 2, that is, one data is taken every other pixel point, that is, only the pixels shown in the shaded part are sequentially acquired. The step being equal to 2 means skipping one pixel. Similarly, when the step is equal to 3, it means skipping two pixels, and so on. It can be understood that when the set step is equal to 1, it corresponds to the data loading method of Figure 7A - Figure 7B showing pixel-by-pixel. In practical applications, by setting the step, the computational amount of subsequent operations such as convolution operations on the data can be further reduced, the processing efficiency can be improved, and in addition, the flexibility of data loading can be further enhanced.

[0122] In the data loading method provided by the embodiments of the present disclosure, a new data acquisition mode different from the whole-block data acquisition method is provided, that is, pixel-by-pixel data loading can be realized according to the starting point and the number of pixels to be acquired (as shown in Figure 6 ), rather than Figure 4The acquisition of the monolithic form in [the above context] is such that this data loading method is more conducive to computational processes such as convolution operations, which is beneficial to improving the computational efficiency. Specifically, the tensor is no longer acquired and transported in a monolithic form, but is acquired and transported sequentially in units of pixel points to be applicable to computational operations such as convolution operations, improving the data transportation efficiency for such operations, reducing the memory access time, and enhancing the overall performance of the processor.

[0123] The above combination Figure 6 、 Figure 7A - Figure 7B 、 Figure 8A - Figure 8C has described in detail the implementation process in the data loading method according to the embodiments of the present disclosure regarding sequentially acquiring data from the original tensor starting from the starting coordinates until the number of acquired data is equal to the number of pixels (copy_pixel_num) included in the tensor to be processed.

[0124] Next, the implementation process regarding Figure 5 determining multiple requests for loading the tensor to be processed by combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor in step S102 shown in [the above context] will be described.

[0125] In an embodiment according to the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as or adjacent to the second dimension. As an example, when continuous loading cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, requests are divided in the second dimension.

[0126] For example, taking a certain element in the original tensor as the origin of the coordinate system to determine a coordinate system. For example, referring to Figure 6 the embodiment of [the above context], taking the upper left vertex of the original tensor as the origin of the coordinate system. Of course, the present disclosure is not limited thereto.

[0127] The starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor are, for example, Figure 6 the coordinates of the upper left vertex of the tensor to be processed in the embodiment of [the above context]. The coordinate values of the starting coordinates of the tensor to be processed are the minimum coordinate values among the coordinate values of all elements in the tensor to be processed.

[0128] For example, the first dimension is the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension, and the first dimension has priority over the second dimension during loading. Specifically, if the first dimension is the W dimension, the second dimension is the H dimension; if the first dimension is the H dimension, the second dimension is the D dimension; if the first dimension is the D dimension, the second dimension is the C dimension.

[0129] In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension not being equal to x, the second dimension is the C dimension.

[0130] In response to the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension. This can be understood as a special case of the requested partitioning. The overall dimension higher than the N dimension refers to the entire tensor. For example, it is expressed as requesting partitioning in the Per1 manner, that is, the entire tensor to be processed can be loaded into the buffer through one request. In this case, there is only one request for loading the tensor to be processed.

[0131] In the present disclosure, the original tensor is stored in the memory in the data storage format of N(C / x)DHW(xC). The data storage format of the original tensor in the memory is different from the data storage format (NDHWC) of the tensor to be processed in the buffer. In the memory, the C dimension of the original tensor is the lowest dimension, followed by the W dimension, then the H dimension, then the D dimension, and the highest dimension is the N dimension. When the tensor to be processed is stored in the buffer, the W dimension is the lowest dimension, followed by the H dimension, then the D dimension, then the C dimension, and the highest dimension is the N dimension. And for the W dimension, continuous xC is substantially considered.

[0132] For example, in response to the size copy_t of the tensor to be processed in the first dimension not being equal to the size tensor_t of the original tensor in the first dimension, and / or the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0133] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension, and it is determined to adopt partitioning of the request in the H dimension. For ease of description, this way of partitioning the request can be called the request partitioning rule according to the H dimension (PerH).

[0134] Taking the second coordinate value as 0 as an example, in the case where the first dimension is the W dimension and the second dimension is the H dimension, when any of the following conditions is satisfied, it is determined that continuous loading cannot be performed in the W dimension:

[0135] (1) The starting coordinate of the tensor to be processed in the W dimension is not equal to 0

[0136] (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension

[0137] (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension

[0138] At this time, it can be understood that the data loaded by each request belongs to the same row (the same H dimension), and the data loaded by different requests is located in different rows, that is, the requests are divided in the H dimension, and the requests are split by row (H dimension).

[0139] According to some embodiments of the present disclosure, in the case where the first dimension is the H dimension and the second dimension is the D dimension, in response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that continuous loading cannot be performed on the tensor to be processed in the H dimension, and it is determined that the requests are divided in the D dimension. For ease of description, this way of dividing requests can be referred to as the request division rule according to the D dimension (PerD).

[0140] Assuming that the first dimension is the H dimension and the second dimension is the D dimension, when any of the following conditions is satisfied, it is determined that continuous loading cannot be performed in the H dimension:

[0141] (1) The starting coordinate of the tensor to be processed in the H dimension is not equal to 0

[0142] (2) The size copy_h of the tensor to be processed in the H dimension is greater than the size tensor_h of the original tensor in the H dimension

[0143] (3) The size copy_h of the tensor to be processed in the H dimension is less than the size tensor_h of the original tensor in the H dimension

[0144] At this time, it can be understood that the coordinates of the data to be loaded by each request are different in the W dimension and the H dimension, but the coordinates in the D dimension and the N dimension are the same. Therefore, the requests are divided in the D dimension.

[0145] According to some embodiments of the present disclosure, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt the partitioning of requests in the C dimension. For ease of description, this way of partitioning requests can be referred to as the request partitioning rule according to the C dimension (PerC).

[0146] Assume that the first dimension is the D dimension and the second dimension is the C dimension. When any of the following conditions is satisfied, it is determined that continuous loading cannot be performed in the D dimension:

[0147] (1) The starting coordinate of the tensor to be processed in the D dimension is not equal to 0

[0148] (2) The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension

[0149] (3) The size copy_d of the tensor to be processed in the D dimension is less than the size tensor_d of the original tensor in the D dimension

[0150] And when the size of the tensor to be processed in the C dimension is greater than x (denoted as copy_c greater than xC), then partitioning of requests is performed in the C dimension. Specifically, considering the particularity of the N(C / x)DHW(xC) interleaving pattern, partitioning is performed in the C dimension with every x data elements as a group. This is because for the N(C / x)DHW(xC) data storage format, every x data elements are continuously stored in the channel number dimension in the storage component.

[0151] Assume that the first dimension is the C dimension and the second dimension is also the C dimension at this time. When any of the following conditions is satisfied, it is determined that continuous loading cannot be performed in the C dimension:

[0152] (1) The starting coordinate of the tensor to be processed in the C dimension is not equal to 0

[0153] (2) The size copy_c of the tensor to be processed in the C dimension is greater than the size tensor_c of the original tensor in the C dimension

[0154] (3) The size copy_c of the tensor to be processed in the C dimension is less than the size tensor_c of the original tensor in the C dimension

[0155] At this time, the requests are still divided in the C dimension. Specifically, considering the particularity of the N(C / x)DHW(xC) interleaving pattern, the requests are divided into groups of every x data elements in the C dimension. This is because for the N(C / x)DHW(xC) format, every x data elements are continuously stored in the channel number dimension of the storage component. For the convenience of description, this way of dividing requests can be called the request division rule according to the C dimension (PerC).

[0156] Assume that the first dimension is the N dimension. Determine that continuous loading cannot be performed in the N dimension when any of the following conditions is satisfied:

[0157] (1) The starting coordinate of the tensor to be processed in the N dimension is not equal to 0

[0158] (2) The size copy_n of the tensor to be processed in the N dimension is greater than the size tensor_n of the original tensor in the N dimension

[0159] (3) The size copy_n of the tensor to be processed in the N dimension is less than the size tensor_n of the original tensor in the N dimension

[0160] If the size of the tensor to be processed in the C dimension is greater than x, at this time the second dimension is the C dimension, and the requests are still divided in the C dimension. Specifically, considering the particularity of the N(C / x)DHW(xC) interleaving pattern, the requests are divided into groups of every x data elements in the C dimension. This is because for the N(C / x)DHW(xC) format, every x data elements are continuously stored in the channel number dimension of the storage component.

[0161] If the size of the tensor to be processed in the C dimension is equal to x, at this time the second dimension is the overall dimension higher than the N dimension. For the tensor data located in the storage component, in essence, a request can be sent to the memory to load the data in the memory at this time, and other requests can be used to fill the data not located in the memory, such as the above invalid data. For the convenience of description, this way of dividing requests can be called the request division according to the whole of the tensor to be processed, and can be expressed as Per1.

[0162] In addition, since the data in the N(C / x)DHW(xC) format needs to be written into the buffer in the NDHWC format, this is different from the data storage format of the original tensor itself, and the dimension arrangement orders of the two data storage formats are also different. Therefore, the update logic of the request initial coordinate during request splitting is different from the dimension arrangement order of the data storage format itself.

[0163] Each request includes a data read address for indicating the starting position to read data from memory, a data write address for indicating the starting position to write data to the buffer, and the number of pixels to be acquired by the request. The request initial coordinates corresponding to the next request are determined based on the request initial coordinates corresponding to the previous request. The request initial coordinates corresponding to each request are used to determine the data read address, data write address, and the number of pixels to be acquired by the request.

[0164] As described later, except for the first dimension, the coordinate values of other dimensions higher than the first dimension are updated based on the coordinate values of the dimensions lower than that other dimension. Since the data of the original tensor needs to be written to the buffer in the NDHWC storage format, for the convenience of hardware implementation, the coordinate values of the request initial coordinates are incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and the coordinate value of the C dimension is incremented by x when updated. That is to say, when updating, the N dimension is updated prior to the C dimension. However, since the N dimension is actually the highest dimension of the tensor to be processed, and the N dimension will be given priority when writing to the buffer, it is necessary to reserve in advance the data positions that should be written to the buffer but have not been written yet. When writing this data subsequently, write the data to the reserved data positions, thereby ensuring that the tensor to be processed in the buffer can be arranged in the dimension order of NDHWC.

[0165] For example, taking the first dimension as the W dimension as an example to illustrate the coordinate update dimension order. The coordinate value h_coord of the request initial coordinates corresponding to the next request in the H dimension is incremented by 1, and when reaching the boundary of the tensor to be processed in the H dimension, it returns to the coordinate initial value h_coord_b; the coordinate value d_coord of the request initial coordinates corresponding to the next request in the D dimension is incremented by 1 when h_coord returns to h_coord_b. If it does not reach the boundary of the tensor to be processed in the D dimension, it remains unchanged, and when reaching the boundary of the tensor to be processed in the D dimension, it returns to the coordinate initial value d_coord_b; the coordinate value n_coord of the request initial coordinates corresponding to the next request in the N dimension is incremented by 1 when d_coord returns to d_coord_b. If it does not reach the boundary of the tensor to be processed in the N dimension, it remains unchanged, and when reaching the boundary of the tensor to be processed in the N dimension, it returns to the coordinate initial value n_coord_b; the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension is incremented by x when n_coord returns to the coordinate initial value n_coord_b, and remains unchanged in other cases.

[0166] The calculation method of the data write address also varies considering the need to reserve data positions in advance. For specific details, reference can be made to the relevant descriptions later.

[0167] The object to be loaded by the present disclosure, i.e., the tensor to be processed, is split into multiple different requests according to a certain rule and sent sequentially, and the tensor data is written into the buffer area in order, thereby efficiently loading the tensor to be processed into the buffer area. A principle during splitting is that the data loaded by each request is either all in memory or all not in memory. This is because when the requested data is in memory, a request is sent to memory, and when the requested data is not in memory, the request can be converted into an operation such as writing a predetermined value to the buffer area by hardware. This splitting method can send corresponding loading requests to different hardware more reasonably.

[0168] In addition, during splitting, if the size relationship between the tensor to be processed and the original tensor in the first dimension causes the tensor to be processed to not be continuously loadable in the first dimension but to be continuously loadable in dimensions lower than the first dimension, then the requests are further divided with the first dimension as the boundary, so that the requests can be split more reasonably, the continuously stored data in memory can be retained as much as possible, the number of requests can be reduced, the tensor to be processed can be loaded efficiently, the bandwidth when loading or storing data from memory can be increased significantly, the performance can be improved, and the efficiency of the hardware computing unit can be improved.

[0169] In an implementation solution where boundary values are set for the original tensor, the data loading method according to the embodiments of the present disclosure may further include: obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, and the boundary values are used to define the boundary range for obtaining data from the original tensor. In some implementation manners, the boundary values may be arbitrary values compared with the size values of the original tensor in this dimension. As an example, the boundary value set for the W dimension may be any of the Figure 8A shown situations. Further, for the dimension with boundary values set, the data loaded by each request either all belongs to the data boundary of the original tensor itself and the valid data range defined by the boundary values, or all does not belong to the valid data range defined by both the data boundary of the tensor itself and the set boundary values.

[0170] As Figure 8A shown, for the first sub - figure, both the left and right boundary values are located on the left side of the data range of the original tensor, that is, both are outside the data range of the original tensor, thereby indicating that all the data to be obtained for this dimension is invalid (for example, represented as out of bound, oob). For Figure 8A the second sub - figure in, the left boundary is located on the left side of the data range of the original tensor, and the right boundary is located within the data range of the original tensor. Thus, only the part of the data that is both within the data range of the original tensor and to the left of the right boundary R is valid data (in Figure 8AAmong them, the data corresponding to the diagonal shaded part is represented as valid data), and the remaining data to be obtained are all invalid data. According to the embodiments of the present disclosure, for the dimension with boundary values set, the rule for dividing requests is that the data loaded by each request is either all valid data or all invalid data (i.e., oob), that is, valid data and invalid data cannot be loaded by the same request.

[0171] The following specifically describes the determination method for multiple requests used to load the tensor to be processed.

[0172] According to the embodiments of the present disclosure, for the multiple requests determined according to step S102, in the actual implementation process of the processor, it can be implemented by setting a state machine. For the state machine, multiple parameters required to determine the above requests can be set, and according to the specific values of the parameters and the initial parameters of the original tensor and the tensor to be processed, the specific tensor to be processed in the original tensor is loaded into the buffer area. The following will describe the implementation scheme related to determining multiple requests for loading the tensor to be processed in combination with the state machine.

[0173] According to some embodiments of the present disclosure, in step S102, in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determining multiple requests for loading the tensor to be processed includes: based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determining the first request among the multiple requests and the initial state when the first request enters the state machine, where the first request indicates the number of pixels to be acquired by the request; and based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed, using the state machine to determine each request after the first request among the multiple requests.

[0174] As an example, for the task of loading the tensor to be processed in the original tensor in memory into the buffer area, the first request can be determined first based on the initial parameters. For example, based on the number of pixels copy_pixel_num included in the tensor to be processed, the data storage format of the original tensor (N(C / x)DHW(xC)), and the starting coordinates of the tensor to be processed (the starting coordinates indicate the position of the first pixel of the tensor to be processed in the coordinate system of the original tensor, for example, represented as its positions in 5 dimensions, c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b), determining the first request among the multiple requests and the initial state when the first request enters the state machine, where the first request indicates the number of pixels to be acquired by the request (req_size_1).

[0175] Next, in combination with Figure 9ADescribe how to determine the partitioning method of the request and the first request. Specifically, Figure 9A A schematic diagram is shown of loading the original tensor of N(C / x)DHW(xC) in a continuous manner as a tensor to be processed stored in NDHWC according to an embodiment of the present disclosure. That is, in Figure 9A the example of, the original tensor is placed in memory in the N(C / x)DHW(xC) data storage format, and the tensor to be processed is loaded into the buffer area in a continuous manner from it and the tensor to be processed is placed in the buffer area in the NDHWC data storage format, so that during the data loading process, the change in the data placement method is realized. It is precisely because of this change that Figure 9A the data acquisition order shown in is different from Figure 7B the order shown in, in Figure 7B the example of, it is a schematic diagram of loading the original tensor of N(C / x)DHW(xC) in a continuous manner as a tensor to be processed stored in N(C / x)DHW(xC), that is, in Figure 7B the example of, there is no involvement in the change of the data storage format.

[0176] In Figure 9A the example of, copy_pixel_num = 40, and the starting coordinates are respectively w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0. As Figure 9A shown, for the above starting coordinates, the starting pixel corresponding to the starting coordinates in the original tensor, that is, start obtaining data from this pixel. Then, for the N(C / x)DHW(xC) data storage format, the W dimension is the lowest dimension. Thus, taking the W dimension as the first dimension, determine whether it can be continuously loaded in the W dimension. In response to the size relationship between the tensor to be processed and the original tensor in the W dimension making it impossible to continuously load the tensor to be processed in the W dimension, then adopt the above-mentioned PerH request partitioning method, that is, partition the request one by one for each H. That is, each row of pixels is loaded into the buffer area by one request. According to the judgment described above regarding whether it can be continuously loaded in the W dimension: In response to the size (copy_w) of the tensor to be processed in the W dimension not being equal to the size (tensor_w) of the original tensor in the W dimension, and / or the first coordinate value (w_coord_b) of the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value (for example, 0) of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension. Summarized as follows:

[0177] Determine that the data to be obtained cannot be continuously loaded in the W dimension when any of the following conditions is satisfied:

[0178] (1) The starting coordinate w_coord_b of the tensor to be processed in the W dimension is not equal to 0;

[0179] (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension;

[0180] (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension;

[0181] After determining the partitioning method, for example, in the PerH manner, the first request can be correspondingly determined. Refer to Figure 9A , the first request is used to load the row of pixels where the starting pixel corresponding to the starting coordinate is located. Its initial state is equal to the starting coordinate of the tensor to be processed, and the number of pixels to be acquired by this first request is req_size_1 = 2 because the subsequent requests are partitioned in a row-by-row (PerH) manner. The specific acquisition order is as shown by the arrows in Figure 9A until the number of pixels to be acquired is equal to the copy_pixel_num = 40 of the tensor to be processed.

[0182] Specifically, in the example of Figure 9A , compared with the example of Figure 7B , the difference lies in the difference in the loading priority for the C dimension. Specifically, in the example of Figure 9A , in order to achieve a change in the data placement method during the data loading process, during the actual loading process, the C dimension is taken as the highest priority. Refer to the loading order shown by the arrows in Figure 9A , first load the H dimension of the first 8 * C dimension, then the D dimension, then the N dimension. After the N dimension (N1) corresponding to the first 8 * C dimension is loaded, jump back to the second 8 * C dimension and continue to load data in the order of the H dimension, D dimension, and N dimension until the number of pixels acquired is equal to the number of pixels included in the tensor to be processed, that is, copy_pixel_num = 40.

[0183] According to an embodiment of the present disclosure, after determining a request, it is further necessary to determine a data read address for the starting position of reading data from the memory by the request, and a data write address for indicating the starting position of writing data to the buffer. Among them, based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, determining the first request among multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0184] According to some embodiments of the present disclosure, the data storage format of the original tensor is N(C / x)DHW(xC). For multiple requests for loading the tensor to be processed, determining the data read address of the nth request among multiple requests includes:

[0185] Calculating the data read address Addr1_n of the nth request according to the following formula:

[0186] Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0187] c_coord * tensor_d * tensor_h * tensor_w +

[0188] d_coord * tensor_h * tensor_w * xC +

[0189] h_coord * tensor_w * xC +

[0190] w_coord * xC

[0191] Among them, u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in 5 dimensions respectively.

[0192] According to some embodiments of the present disclosure, for multiple requests for loading tensors to be processed, determining the data write address of the nth request among the multiple requests includes:

[0193] For m requests with the same coordinate value range of the data of the request in the C dimension, for example, the data with the same coordinate value range is split into m requests for sending, for example, the coordinate values of the m requests are all from 0 to x - 1, or from x to 2*x - 1, etc., calculate the data write address Addr_2_i of the ith request sent among the m requests according to the following formula:

[0194] Addr_2_i = b_addr + req_size_1*copy_c + req_size_2*copy_c +…+ req_size_i-1*copy_c

[0195] Among them, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2,..., req_size_i-1 represent the number of pixels to be acquired by the previous i - 1 requests respectively, and copy_c is the size of the tensor to be processed in the C dimension.

[0196] In response to that the data of each request among the m requests includes the first x data elements of the original tensor in the C dimension, for example, the coordinate values are from 0 to x - 1, b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer.

[0197] In response to that the data of each request among the m requests includes the tth data element to the (t + x)th data element of the original tensor in the C dimension, t is greater than x, for example, t is equal to 2*x or 3*x, etc., b_addr is calculated according to the following formula:

[0198] b_addr = b_addr_base+ (t / x - 1)*x

[0199] Among them, t, m, and i are positive integers.

[0200] For the first sent request, the corresponding request initial coordinate is the starting coordinate of the tensor to be processed. Thus, referring to the above formula, the data read address of the first request can be calculated. The starting address of the tensor to be processed written into the buffer is used as the data write address of the first request.

[0201] The number of pixels to be acquired by the first request can be determined according to parameters such as the determined request partitioning method, the starting coordinate of the tensor to be processed, and the size of the original tensor. For example, if the sum of the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is less than the second coordinate value (i.e., the coordinate value of the starting coordinate of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the shape size copy_t of the tensor to be processed in the first dimension. For example, if the sum of the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value (i.e., the coordinate value of the starting coordinate of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the absolute value of the first coordinate value t_coord_b, and this data length represents the number of pixels to be acquired by this request.

[0202] Considering that the state machine has the advantages of a clear logical structure, being easy to maintain and expand, and being particularly suitable for processing logical scenarios with multiple conditions and multiple branches, in the present disclosure, the state machine is adopted to automatically update the request initial coordinates and the loaded data lengths corresponding to each request, avoiding complex conditional nesting, with clear logic, easy to maintain, and strong scalability.

[0203] According to some embodiments of the present disclosure, based on the starting coordinate of the tensor to be processed, determining the initial state of the first request entering the state machine includes: in response to the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0204] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired in the first request and the number of pixels included in the tensor to be processed, a state machine is used to determine each request after the first request among multiple requests, including: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired in the second request; determining the data read address of the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; determining the data write address of the second request according to the number of pixels to be acquired in the second request; updating the state machine according to the information related to the second request, and sequentially determining subsequent requests in each request based on the updated state machine.

[0205] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the data to be loaded by the current request is located in the memory. At this time, the data read address and data write address of the current request can be determined with reference to the above calculation formula, and then the current request is sent to the memory to load the corresponding data into the buffer.

[0206] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the data to be loaded by the current request is not located in the memory. At this time, in response to all the data to be loaded by the current request not being located in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values into the buffer, where the number of multiple predetermined values is determined by the length of the data to be loaded specified by the current request. For example, the predetermined value is 0.

[0207] Figure 9B The state diagram of the state machine provided according to some embodiments of the present disclosure is shown.

[0208] For example, the states of the state machine include a first state s0, a second state s1, and a third state s2, and the initial state of entering the state machine is determined by the starting coordinates (c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b) of the tensor to be processed.

[0209] For example, in response to the first coordinate value t_coord_b of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that the initial state is the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that the initial state is the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, it is determined that the initial state is the third state.

[0210] Reference Figure 9B , the data in the tensor_t direction represents the data in the first dimension. The data with the coordinate in the t dimension less than the second coordinate value and greater than the third coordinate value is not in memory or outside the range of the original tensor data ( Figure 9B the part shown by the diagonal shading, and this part of the data can be called invalid data), and the data with the size of the t dimension between the second coordinate value and the third coordinate value ( Figure 9B the white part in) is stored in memory, which is the actual size of the original tensor data in the first dimension.

[0211] The state only switches to itself and adjacent states. For example Figure 9B in, the first state s0 can jump to the first state s0 or the second state s1, the second state s1 can jump to the first state s0, the second state s1 and the third state s2, and the third state s2 can jump to the first state s0, the second state s1 and the third state s2.

[0212] Figure 9B The 6 cases (① to ⑥) in show the state transition changes that the tensor to be processed experiences for different sizes in the first dimension.

[0213] For example, Figure 9B in case ①, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the second coordinate value. The state machine always loops and jumps in the first state s0 on the left.

[0214] For example, Figure 9B in case ②, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value and less than the third coordinate value. The state machine loops and jumps between the first state s0 and the second state s1.

[0215] For example, Figure 9B in case ③, the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value. The state machine always loops and jumps in the second state s1.

[0216] For example, Figure 9B in case ④, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1 and the third state s2.

[0217] For example, Figure 9BIn case ⑤, the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps between the second state s1 and the third state s2.

[0218] For example, Figure 9B In case ⑥, the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops and jumps in the third state s2.

[0219] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and combine with the trigger condition to determine the state that the second sent request enters. Then, based on the state that the second sent request enters, determine the request initial coordinates corresponding to the third sent request and the data length loaded by the second sent request, and combine with the trigger condition to determine the state that the third sent request enters, and so on.

[0220] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request in each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state. After determining the request initial coordinates corresponding to the current request, it can be judged whether the data loaded by the current request is in the memory according to the request initial coordinates.

[0221] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the data to be loaded by the current request is in the memory. It can be understood that for the data range that all belongs to the original tensor, it will all be within the memory range (this part of the data is actually stored in the memory), and for the data range that all does not belong to the original tensor, it will all not be within the memory range and is not the actually stored data. That is, the data exceeding the original tensor range can be understood as invalid data. For example, this part of the invalid data can be directly filled with zeros in the buffer area. At this time, the data read address and data write address of the current request can be determined with reference to the above formula, and then the current request is sent to the memory to load the corresponding data into the buffer area.

[0222] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the data to be loaded by the current request is not in the memory. At this time, in response to all the data to be loaded by the current request not being in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values into the buffer area, where the number of the multiple predetermined values is determined by the data length specified by the current request. For example, the predetermined value is 0.

[0223] The process of using the state machine to determine the request initial coordinates and the data length loaded by each request will be specifically described below.

[0224] Figure 9C shows a state transition schematic diagram according to the PerH partitioning method for an example of Figure 9A For the PerH partitioning method, the condition is that continuous loading cannot be performed in the W dimension. At this time, the adjacent two rows of pixels to be acquired are not continuous in memory, and split requests for each row of pixels are required to obtain data.

[0225] In Figure 9C on the left side, only three states s0, s1, and s2 are shown for the W and H dimensions. Specifically, s0 can correspond to the first state, s1 can correspond to the second state, and s2 can correspond to the third state. In addition, in Figure 9C the state transition schematic diagram on the right side, there is also an idle state, which is used to represent the idle state before data loading or after all the tensors to be processed have been loaded. Table 1 below shows the meanings corresponding to each state:

[0226] Table 1

[0227]

[0228] That is to say, in the example of Figure 9C the middle blank box can correspond to the size of the original tensor in the W dimension (tensor_w). Its left boundary can correspond to, for example, Figure 9B the second coordinate value in Figure 9C which is equal to 0, and its right boundary can correspond to, for example,

[0229] Figure 9C the third coordinate value in

[0230] Table 2

[0231]

[0232] It should be noted that in the examples of Table 2 and subsequent tables above, the case of setting boundary values for the original tensor is further considered. Among them, right_bound represents the right boundary value set for the W dimension, down_bound represents the boundary value in the H dimension, and hind_bound represents the boundary value in the D dimension. The setting of boundary values and strides can refer to the description in the above text in combination with Figure 8A - Figure 8C the description. In addition, for the case where no boundary value is set, it can be understood that each boundary is equal to the size of the original tensor in that dimension. In addition, it can be understood that in other embodiments according to the present disclosure, boundary values and strides can also be set for the C dimension, for example, and there is no limitation here. Regarding the left boundary left_bound and right boundary right_bound for the W dimension, the upper boundary up_bound and lower boundary down_bound for the H dimension, and the front boundary front_bound and rear boundary hind_bound for the D dimension, reference can be made to Figure 6 the cube shown in

[0233] In each table in this article, the parameter remain_copy_p represents the number of remaining pixels that have not been acquired. For example, before the first request, remain_copy_p is equal to the number of pixels copy_pixel_num corresponding to the tensor to be processed. After the first request, remain_copy_p is equal to copy_pixel_num - req_size_1, where req_size_1 is the number of pixels to be acquired in the first request. The parameters Is_last_c and Is_not_last_c are respectively used to indicate whether the currently acquired data is for the last 8*C dimension. For example, referring to Figure 9A the schematic diagram in

[0234] Next, the information to be updated in each state and how to update it will be introduced.

[0235] First, in the s0 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table:

[0236] Table 3

[0237]

[0238] The parameter update rules in Table 3 are described below. Here, stride_x, stride_y, and stride_z respectively represent the stride values in the W dimension, H dimension, and D dimension. As an implementation, the strides can all be set to 1. In addition, in the examples of Tables 3 - 5, x = 8 is used for illustration. For example, it is expressed as 8C, that is, the size of the C dimension bound to the W dimension is 8 * C.

[0239] Specifically, in the solutions of Table 3 and Tables 4 - 5, it is first necessary to determine whether to jump within the same 8C or between different 8Cs. Taking Figure 9A as an example, jumping within the same 8C can be, for example, to obtain data within D0 - D1 of the first 8 * C for the N0 dimension. Jumping between different 8Cs can be, for example, to jump from the N0 dimension to the first 8 * C of the N1 dimension to obtain data. That is, for some implementation manners of the present disclosure, it is necessary to distinguish whether to jump in the 8 * C dimension. This is because in the data loading method of the present application, the to - be - processed tensor is loaded from the original tensor with the data storage format of N(C / x)DHW(xC), and the to - be - processed tensor is placed in the buffer area in the NDHWC data storage format. Due to the change in the data storage format, the priority of the C dimension has changed, so that in the process of loading data for multiple requests based on partitioning, it is necessary to reserve in advance this part of the request address for the C dimension. The specific implementation process is reflected in the state update and jump conditions of the state machine.

[0240] According to Table 3, first, for the W dimension, (1) for the case of jumping within the same 8C, if jumping from state s0 to state s0, the coordinate w_coord of the W dimension is updated to left_bound, where left_bound represents the left - hand boundary value set for the W dimension. If jumping from state s0 to state s1, the coordinate w_coord of the W dimension is updated to 0; (2) for the case of jumping between different 8Cs, the coordinate w_coord of the W dimension is directly updated to the starting coordinate w_coord_b of the W dimension.

[0241] The coordinate update logic for the other dimensions in Table 3 is the same as that of the W dimension, which is described separately as follows.

[0242] For the H dimension, (1) For the case of jumps within the same 8C, if the jump is from state s0 to state s0, when h_coord + stride_y >= down_bound, the coordinate h_coord of the H dimension is updated to h_coord + stride_y - copy_h; otherwise (h_coord + stride_y < down_bound), the coordinate h_coord of the H dimension is updated to h_coord + stride_y. Otherwise (the jump is from state s0 to a state other than s0), the coordinate h_coord of the H dimension remains unchanged. (2) For the case of jumps between different 8Cs, the coordinate h_coord of the H dimension is directly updated to the starting coordinate h_coord_b of the H dimension.

[0243] For the D dimension, (1) For the case of jumps within the same 8C, if the jump is from state s0 to state s0 and h_coord + stride_y >= down_bound, when d_coord + stride_z >= hind_bound, the coordinate d_coord of the D dimension is updated to d_coord + stride_z - copy_d, where hind_bound represents the rear boundary value set for the D dimension; otherwise (d_coord + stride_z < hind_bound), the coordinate d_coord of the D dimension is updated to d_coord + stride_z. Otherwise (the jump is from state s0 to a state other than s0), the coordinate d_coord of the D dimension remains unchanged. (2) For the case of jumps between different 8Cs, the coordinate d_coord of the D dimension is directly updated to the starting coordinate d_coord_b of the D dimension.

[0244] For the C dimension, if it is a jump between different 8Cs, the coordinate c_coord of the C dimension is updated to c_coord_b + 8; otherwise, the coordinate c_coord of the C dimension remains unchanged. The coordinate update logic of the C dimension is slightly different from that of the W, H, and D dimensions because, as described above, the placement method after data loading has changed from the original N(C / x)DHW(xC) to NDHWC, which makes the data in the C dimension have the lowest priority after loading.

[0245] For the N dimension, (1) for the case of jumping within the same 8C, if jumping from state s0 to state s0, at this time h_coord + stride_y >= down_bound and d_coord + stride_z >= hind_bound, the coordinate of the N dimension is updated to n_coord + 1; otherwise, the coordinate n_coord of the N dimension remains unchanged. (2) For the case of jumping between different 8Cs, the coordinate n_coord of the N dimension is directly updated to the starting coordinate n_coord_b of the N dimension.

[0246] In addition to the above dimension coordinate parameters, the state machine can also maintain some parameters required for the state update process and the calculation of the request address. For example, the parameter remain_copy_p in Table 3, which represents the remaining number of pixels to be acquired. (1) For the case of jumping within the same 8C, if jumping from state s0 to state s0, then remain_copy_p is updated to remain_copy_p - (right_bound - w_coord), and if jumping from state s0 to state s1, then remain_copy_p is updated to remain_copy_p + w_coord. (2) For the case of jumping between different 8Cs, remain_copy_p is updated to copy_pixel_num.

[0247] Second, in state s1, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table:

[0248] Table 4

[0249]

[0250] Third, in state s2, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table:

[0251] Table 5

[0252]

[0253] For Table 4 and Table 5, the state update process of each parameter can refer to the description for Table 3 and will not be elaborated here. Based on the above update logic, the update results of Table 4 and Table 5 can be derived analogously.

[0254] The above describes the state transition according to the PerH partitioning method during the process of loading the original tensor stored in N(C / x)DHW(xC) into the tensor to be processed in the NDHWC data storage format, which involves states s0, s1, and s2, as well as the updated outputs of the state machine in these three states. It can be understood that the principles and update logics of the state machine for other partitioning methods (e.g., PerD, PerC, etc.) are similar to those described above and will not be described here.

[0255] Of course, as the coordinates continue to increase, when the number of acquired data is equal to the number of pixels included in the tensor to be processed, it is determined that the state machine can end the coordinate determination for each request, and the state machine can enter the idle state, which indicates that the data for the current pen instruction has been acquired.

[0256] The specific state transition conditions of the state machine and the updated coordinate content can be changed and set according to actual needs, and the logic is similar to the state machine logic described above, and no further examples will be given here.

[0257] In the above embodiment, the requests are partitioned according to whether the tensor can be continuously loaded in the first dimension. Each request is used to load data that is either all in memory or all not in memory. The partitioning of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.

[0258] According to some embodiments of the present disclosure, another data loading method is provided. Figure 10 The schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure is shown. As Figure 10 shown, the data loading method according to the embodiment of the present disclosure includes step S201 and step S202.

[0259] In step S201: Receive a data loading instruction indicating to execute loading the tensor to be processed from the original tensor in the memory to the buffer area.

[0260] According to the embodiment of the present disclosure, the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer area is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is regarded as a pixel, and it accumulates step by step to higher dimensions, where the channel number dimension does not calculate the number of pixels.

[0261] After parsing the data loading instruction in step S202, the execution unit executes the data loading instruction.

[0262] As Figure 10 shown, wherein step S202 uses the execution unit to execute the data loading instruction, including:

[0263] S2021: Obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor;

[0264] S2022: Combine the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine multiple requests for loading the tensor to be processed, wherein the multiple requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and

[0265] S2023: Send multiple requests in sequence, and write the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer.

[0266] For the related descriptions of the tensor to be processed and the original tensor and the specific implementation process of steps S2021 - S2023, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be described again.

[0267] As an example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction is implemented in the form of a Load instruction.

[0268] According to some embodiments of the present disclosure, a data storage method is further provided for obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into the memory. It can be understood that the data storage method according to the embodiments of the present disclosure can be understood as the reverse process of the data loading method described above, that is, moving data from the buffer to the memory. The implementation principle according to the embodiments of the present disclosure is similar to the above data loading method, and the repeated parts will not be described again, and only the different parts will be described in detail.

[0269] In the data storage method according to an embodiment of the present disclosure, the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step in the higher dimensions, where the channel number dimension does not calculate the number of pixels.

[0270] As an example, the first tensor may refer to the tensor arranged in the data storage format of N(C / x)DHW(xC) stored in the buffer. For the second tensor obtained therefrom, it is arranged in the data storage format of NDHWC in the memory. In terms of understanding the implementation principle, the first tensor can be correspondingly understood as the original tensor in the data loading method described above, that is, part or all of the data is obtained from the first tensor and transferred and stored in the memory. This part of the tensor obtained from the first tensor is represented as the second tensor, which can be correspondingly understood as the tensor to be processed in the data loading method described above.

[0271] Figure 11 The schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown, as Figure 11 shown, the data storage method includes steps S301 - S303.

[0272] In step S301, obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor. Then, in step S302, combine the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor. In step S303, sequentially send the plurality of requests and sequentially write the data corresponding to each request into the memory to store the second tensor in the memory.

[0273] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request belongs to the data range of the first tensor or does not belong to the data range of the first tensor. Moreover, in response to the data loaded by the request belonging to the data range of the first tensor, the data loaded by the request is from the first tensor and is continuously stored in the buffer area, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0274] In the data storage method according to an embodiment of the present disclosure, the manner of sequentially obtaining data from the first tensor can be referred to and described in combination with Figure 6 the description shown. Compared with the block-based data storage method in the related art, the data storage method provided by the embodiment of the present disclosure can achieve pixel-by-pixel data storage, rather than Figure 4 the block-based acquisition in . This data storage method is more beneficial to calculation processes such as convolution operations and is conducive to improving the operation efficiency. It can be understood that, for example, a processor can reasonably use, according to the type of operation to be performed or data processing characteristics, whether to use a continuous data loading method or a block-based data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer area can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0275] In actual applications, for example, in the convolution operation in general computing operations, its operation process is multi-layered. For example, after the first layer of processing is completed, the data result needs to be stored for use in the next layer of processing. As described above, for the data that needs to perform convolution operations, adopting the sequential data acquisition method provided by the embodiment of the present disclosure (which can refer to the description combined with Figure 5 、 Figure 6 、 Figure 7A 、 Figure 7B 、 Figure 8A 、 Figure 8B and Figure 8C above) is more conducive to improving the calculation efficiency. As an example, the 1x3 weight data can only obtain data in the form of a sliding window, so it can only be obtained row by row, unless the shape of the graph to be obtained is very regular, exactly a cube and no operations such as padding need to be performed in the next layer.

[0276] According to some embodiments of the present disclosure, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the second tensor in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the second tensor in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

[0277] According to some embodiments of the present disclosure, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the first dimension, it is determined that the second tensor cannot be continuously loaded in the first dimension.

[0278] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the second tensor in the W dimension not being equal to the size of the first tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the W dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the W dimension, it is determined that the second tensor cannot be continuously loaded in the W dimension, and it is determined to adopt a division of requests in the H dimension.

[0279] According to some embodiments of the present disclosure, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the second tensor in the H dimension not being equal to the size of the first tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the H dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the H dimension, it is determined that the second tensor cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

[0280] According to some embodiments of the present disclosure, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the second tensor in the D dimension not being equal to the size of the first tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the D dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the D dimension, it is determined that the second tensor cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

[0281] According to some embodiments of the present disclosure, each request includes a data read address for indicating a starting position to read data from a buffer, a data write address for indicating a starting position to write data to a memory, and the number of pixels to be acquired by the request. The request initial coordinates corresponding to the next request are determined based on the request initial coordinates corresponding to the previous request. The request initial coordinates corresponding to each request are used to determine the data read address, the data write address, and the number of pixels to be acquired by the request. Wherein, when sequentially determining the request initial coordinates corresponding to each request, the coordinate values of the request initial coordinates are incrementally updated in the order of the W dimension, the H dimension, the D dimension, the N dimension, and the C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0282] According to some embodiments of the present disclosure, in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, multiple requests for storing the second tensor are determined, including: determining the first request among the multiple requests and the initial state of the first request entering the state machine based on the number of pixels included in the second tensor and the starting coordinates of the second tensor, wherein the number of pixels to be acquired by the request is indicated in the first request; and determining each request after the first request among the multiple requests by using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the second tensor.

[0283] According to some embodiments of the present disclosure, each request includes a data read address for indicating a starting position to read data from a buffer, a data write address for indicating a starting position to write data to a memory, and the number of pixels to be acquired by the request. Wherein, determining the first request among the multiple requests and the initial state of the first request entering the state machine based on the number of pixels included in the second tensor and the starting coordinates of the second tensor includes: determining the initial state of the first request entering the state machine based on the starting coordinates of the second tensor; taking the starting coordinates of the second tensor as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape size of the first tensor, the request initial coordinates corresponding to the first request, and the data storage format of the second tensor; determining the starting address for writing the second tensor in the memory as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the second tensor, the data storage format of the first tensor, the starting coordinates of the second tensor, and the shape size of the first tensor.

[0284] According to some embodiments of the present disclosure, determining an initial state in a state machine for a first request based on a starting coordinate of a second tensor includes: in response to a first coordinate value of the starting coordinate of the second tensor in a first dimension being less than a second coordinate value of the starting coordinate of a first tensor in the first dimension, determining the initial state as a first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, determining the initial state as a second state, where a difference between the second coordinate value and the third coordinate value is equal to a size of the first tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as a third state.

[0285] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the second tensor, using a state machine to determine each request among a plurality of requests after the first request includes: based on the initial state, using the state machine to determine a request initial coordinate corresponding to a second request after the first request and the number of pixels to be acquired by the second request; according to a shape size of the first tensor, the request initial coordinate corresponding to the second request, and a data storage format of the first tensor, determining a data read address of the second request; according to the number of pixels to be acquired by the second request, determining a data write address of the second request; updating the state machine based on information related to the second request, and sequentially determining subsequent requests among each request based on the updated state machine.

[0286] According to some embodiments of the present disclosure, for a plurality of requests for storing a second tensor, determining a data read address of an nth request among the plurality of requests includes:

[0287] Calculating a data read address Addr1_n of the nth request according to the following formula:

[0288] Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0289] c_coord * tensor_d * tensor_h * tensor_w +

[0290] d_coord * tensor_h * tensor_w * xC +

[0291] h_coord * tensor_w * xC +

[0292] w_coord * xC

[0293] Among them, u_addr_base represents the storage address in the buffer of the pixel at the starting coordinate position of the first tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the first tensor in 5 dimensions respectively.

[0294] According to some embodiments of the present disclosure, for multiple requests for storing a second tensor, determining the data write address of the nth request among the multiple requests includes:

[0295] For m requests with the same coordinate value range of the data of the requests in the C dimension, calculate the data write address Addr_2_i of the ith request sent among the m requests according to the following formula:

[0296] Addr_2_i = b_addr + req_size_1*copy_c + req_size_2*copy_c +…+ req_size_i-1*copy_c

[0297] Among them, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2,..., req_size_i-1 represent the number of pixels to be obtained by each of the previous i-1 requests, and copy_c is the size of the second tensor in the C dimension;

[0298] Among them, in response to the data of each request among the m requests including the first x data elements of the first tensor in the C dimension, b_addr is the starting address b_addr_base for writing the second tensor into the memory,

[0299] In response to the data of each request among the m requests including the tth data element to the (t + x)th data element of the first tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula:

[0300] b_addr = b_addr_base+ (t / x - 1)*x

[0301] t, m, and i are positive integers.

[0302] According to some embodiments of the present disclosure, sequentially sending multiple requests and sequentially writing the data corresponding to each request into the memory to store the second tensor in the memory includes: for any request, in response to all the data corresponding to any request being located in the buffer, sending any request to the buffer; in response to all the data corresponding to any request not being located in the buffer, then abandoning the request.

[0303] The data storage method according to an embodiment of the present disclosure may further include: obtaining boundary values for at least a part of five dimensions of the first tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the first tensor. For example, the boundary values can be any values compared to the size values of the first tensor in this dimension. As an example, the boundary value set for the W dimension can be any of the situations shown in Figure 8A the figure. Further, for the dimension with boundary values set, all the data loaded by each request either belongs to the data boundary of the first tensor itself and within the valid data range defined by the boundary values, or does not belong to the valid data range defined by both the data boundary of the tensor itself and the set boundary values.

[0304] It can be understood that the data storage method according to an embodiment of the present disclosure can achieve similar technical effects to the data loading method according to an embodiment of the present disclosure.

[0305] According to some embodiments of the present disclosure, another data storage method is provided. Figure 12 The schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown. As Figure 12 shown, the data storage method according to an embodiment of the present disclosure includes step S401 and step S402.

[0306] In step S401: Receive a data storage instruction indicating to execute obtaining a second tensor based on the first tensor in the buffer area and writing the second tensor into the memory.

[0307] According to an embodiment of the present disclosure, the data storage format of the first tensor in the buffer area is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and accumulates gradually to higher dimensions, where the channel number dimension does not calculate the number of pixels.

[0308] In step S402: After parsing the data storage instruction, use the execution unit to execute the data storage instruction.

[0309] As Figure 12 shown, where step S402 uses the execution unit to execute the data storage instruction, including:

[0310] S4021: Obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor;

[0311] S4022: Determine a plurality of requests for loading the second tensor in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and

[0312] S4023: Sequentially send the plurality of requests and sequentially write the data corresponding to each request into the memory to store the second tensor into the memory.

[0313] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, divide the requests in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0314] For the related descriptions of the first tensor and the second tensor and the specific implementation processes of steps S4021 - S4023, reference can be made to the related descriptions of the foregoing data storage method, and repeated parts will not be elaborated.

[0315] As an example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction is implemented in the form of a Store instruction.

[0316] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit, where the instruction parsing unit is configured to: receive and parse a data loading instruction, where the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not calculate the number of pixels; and the execution unit is configured to: execute the data loading instruction.

[0317] According to an embodiment of the present disclosure, an execution unit executes a data loading instruction, including: obtaining the number of pixels included in a tensor to be processed and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into a buffer area to load the tensor to be processed into the buffer area.

[0318] According to an embodiment of the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as or adjacent to the second dimension.

[0319] For example, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

[0320] For example, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that continuous loading cannot be performed on the tensor to be processed in the first dimension.

[0321] For example, in the case where the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinates of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinates of the original tensor in the W dimension, it is determined that continuous loading cannot be performed on the tensor to be processed in the W dimension, and it is determined to perform division of requests in the H dimension.

[0322] For example, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the tensor to be processed in the H dimension being not equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension being not equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be continuously loaded in the H dimension, and it is determined to adopt partitioning of requests in the D dimension.

[0323] For example, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the tensor to be processed in the D dimension being not equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension being not equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension, and when the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt partitioning of requests in the C dimension.

[0324] For example, each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data to the buffer, and the number of pixels to be acquired by this request. The initial coordinates of the request corresponding to the next request are determined according to the initial coordinates of the request corresponding to the previous request. The initial coordinates of each request are used to determine the data read address, the data write address, and the number of pixels to be acquired by the request. Among them, when successively determining the initial coordinates of each request corresponding to the requests, the coordinate values of the initial coordinates of the requests are incrementally updated in the order of the W dimension, the H dimension, the D dimension, the N dimension, and the C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0325] For example, in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, a plurality of requests for loading the tensor to be processed are determined, including: based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determining the first request among the plurality of requests and the initial state when the first request enters the state machine, where the first request indicates the number of pixels to be acquired by this request; and based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed, using the state machine to determine each request among the plurality of requests after the first request.

[0326] For example, each request includes a data read address for indicating the starting position to read data from the memory, a data write address for indicating the starting position to write data to the buffer, and the number of pixels to be acquired by the request. Among them, based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determining the first request among multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address to write the tensor to be processed in the buffer as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0327] For example, determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0328] For example, based on the initial state, combining the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed, using the state machine to determine each request after the first request among multiple requests includes: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired by the second request; determining the data read address of the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; determining the data write address of the second request according to the number of pixels to be acquired by the second request; updating the state machine according to the information related to the second request, and sequentially determining the subsequent requests among each request based on the updated state machine.

[0329] For example, for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among multiple requests includes:

[0330] Calculating the data read address Addr1_n of the nth request according to the following formula:

[0331] Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0332] c_coord * tensor_d * tensor_h * tensor_w +

[0333] d_coord * tensor_h * tensor_w * xC +

[0334] h_coord * tensor_w * xC +

[0335] w_coord * xC

[0336] Among them, u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in 5 dimensions respectively.

[0337] For example, for multiple requests for loading tensors to be processed, determining the data write address of the nth request among the multiple requests includes: for m requests with the same coordinate value range of the data of the requests in the C dimension, calculating the data write address Addr_2_i sent for the ith request among the m requests according to the following formula:

[0338] Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i - 1 * copy_c

[0339] Among them, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2,..., req_size_i - 1 represent the number of pixels to be obtained by each of the previous i - 1 requests, and copy_c is the size of the tensor to be processed in the C dimension; among them, in response to each request among the m requests, the data includes the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base written to the tensor to be processed in the buffer, in response to each request among the m requests, the data includes the tth data element to the (t + x)th data element of the original tensor in the C dimension, t > x, and b_addr is calculated according to the following formula:

[0340] b_addr = b_addr_base+ (t / x-1)*x

[0341] t, m, and i are positive integers.

[0342] For example, multiple requests are sent in sequence, and the data corresponding to each request is written into the buffer in order to load the tensor to be processed into the buffer, including: for any one of the requests, in response to all the data loaded by any one of the requests being located entirely in the memory, sending any one of the requests to the memory; in response to all the data loaded by any one of the requests not being located entirely in the memory, converting any one of the requests into writing multiple predetermined values into the buffer, where the number of the multiple predetermined values is determined by the number of pixels to be obtained by any one of the requests.

[0343] The data storage method according to an embodiment of the present disclosure may further include: obtaining boundary values for at least some of the five dimensions of the first tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the first tensor, and for the dimensions with boundary values set, the data loaded by each request belongs to the effective data range defined by the original tensor and the boundary values or does not belong to the effective data range.

[0344] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit, where the instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into the memory, where the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, the channel number of xC is bound to the width dimension, and in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as one pixel, and is accumulated step by step to the higher dimension, where the channel number dimension does not calculate the number of pixels; and the execution unit is configured to: execute the data storage instruction.

[0345] According to an embodiment of the present disclosure, an execution unit executes a data storage instruction, including: obtaining the number of pixels included in a second tensor and the starting coordinates of the second tensor in a coordinate system determined by a first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into memory to store the second tensor in memory.

[0346] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0347] For example, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the second tensor in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the second tensor in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

[0348] For example, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the first coordinate value of the starting coordinates of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinates of the first tensor in the first dimension, it is determined that continuous loading of the second tensor cannot be performed in the first dimension.

[0349] For example, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the second tensor in the W dimension not being equal to the size of the first tensor in the W dimension, and / or the first coordinate value of the starting coordinates of the second tensor in the W dimension not being equal to the second coordinate value of the starting coordinates of the first tensor in the W dimension, it is determined that continuous loading of the second tensor cannot be performed in the W dimension, and it is determined to perform division of requests in the H dimension.

[0350] For example, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the second tensor in the H dimension not being equal to the size of the first tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the H dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the H dimension, it is determined that the second tensor cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

[0351] For example, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the second tensor in the D dimension not being equal to the size of the first tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the D dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the D dimension, it is determined that the second tensor cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

[0352] For example, each request includes a data read address for indicating the starting position of reading data from the buffer, a data write address for indicating the starting position of writing data to the memory, and the number of pixels to be acquired by the request. The initial coordinate of the request corresponding to the next request is determined based on the initial coordinate of the request corresponding to the previous request. The initial coordinate of each request is used to determine the data read address, the data write address, and the number of pixels to be acquired by the request. Among them, when successively determining the initial coordinates of the requests corresponding to each request, the coordinate values of the initial coordinates of the requests are incrementally updated in the order of the W dimension, the H dimension, the D dimension, the N dimension, and the C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0353] For example, in combination with the number of pixels included in the second tensor and the starting coordinate of the second tensor in the coordinate system determined by the first tensor, multiple requests for storing the second tensor are determined, including: based on the number of pixels included in the second tensor and the starting coordinate of the second tensor, determining the first request among the multiple requests and the initial state when the first request enters the state machine, where the first request indicates the number of pixels to be acquired by the request; and based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the second tensor, using the state machine to determine each request among the multiple requests after the first request.

[0354] For example, each request includes a data read address for indicating the starting position to read data from the buffer, a data write address for indicating the starting position to write data to the memory, and the number of pixels to be acquired by the request. Among them, based on the number of pixels included in the second tensor and the starting coordinates of the second tensor, determining the first request among the multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the second tensor; taking the starting coordinates of the second tensor as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the first tensor, the request initial coordinates corresponding to the first request, and the data storage format of the second tensor; determining the starting address to write the second tensor in the memory as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the second tensor, the data storage format of the first tensor, the starting coordinates of the second tensor, and the shape and size of the first tensor.

[0355] For example, determining the initial state when the first request enters the state machine based on the starting coordinates of the second tensor includes: in response to the first coordinate value of the starting coordinates of the second tensor in the first dimension being less than the second coordinate value of the starting coordinates of the first tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the first tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0356] For example, based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the second tensor, using the state machine to determine each request after the first request among the multiple requests includes: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired by the second request; determining the data read address of the second request according to the shape and size of the first tensor, the request initial coordinates corresponding to the second request, and the data storage format of the first tensor; determining the data write address of the second request according to the number of pixels to be acquired by the second request; updating the state machine according to the information related to the second request, and sequentially determining the subsequent requests in each request based on the updated state machine.

[0357] For example, for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes:

[0358] Calculating the data read address Addr1_n of the nth request according to the following formula:

[0359] Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0360] c_coord * tensor_d * tensor_h * tensor_w +

[0361] d_coord * tensor_h * tensor_w * xC +

[0362] h_coord * tensor_w * xC +

[0363] w_coord * xC

[0364] Wherein, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the first tensor in five dimensions respectively.

[0365] For example, for multiple requests for storing the second tensor, determining the data write address of the nth request among the multiple requests includes:

[0366] For m requests with the same coordinate value range of the data of the requests in the C dimension, calculate the data write address Addr_2_i of the ith request sent among the m requests according to the following formula:

[0367] Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i - 1 * copy_c

[0368] Wherein, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2, ..., req_size_i - 1 represent the number of pixels to be acquired by the respective previous i - 1 requests, and copy_c is the size of the second tensor in the C dimension;

[0369] Wherein, in response to the data of each of the m requests including the first x data elements of the first tensor in the C dimension, b_addr is the starting address b_addr_base for writing the second tensor into the memory,

[0370] The data corresponding to each of the m requests includes the t-th to the t+x-th data elements of the first tensor in the C dimension, where t is greater than x, and b_addr is calculated according to the following formula:

[0371] b_addr = b_addr_base+ (t / x-1)*x

[0372] t, m, and i are positive integers.

[0373] For example, multiple requests are sent in sequence, and the data corresponding to each request is written into the memory in order to store the second tensor in the memory, including: for any request, in response to the data corresponding to any request being entirely located in the buffer, sending any request to the buffer; in response to the data corresponding to any request not being entirely located in the buffer, abandoning the request.

[0374] As an example, Figure 13 shows a schematic block diagram of a processor according to some embodiments of the present disclosure. As Figure 13 shown, the processor 1000 may include an instruction parsing unit 1010 and an execution unit 1020. It can be understood that the processor 1000 may be implemented to execute the data loading method according to the embodiments of the present disclosure to load the tensor to be processed from the original tensor in the memory to the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain the second tensor based on the first tensor in the buffer and write the second tensor into the memory.

[0375] Regarding the specific implementation processes of the data storage method and the data loading method, reference can be made to the above description and will not be repeated here. The processor provided by at least one embodiment of the present disclosure can achieve technical effects similar to those of the aforementioned data loading method / data storage method, and the repeated parts will not be elaborated.

[0376] According to some embodiments of the present disclosure, an electronic device is further provided. Figure 14 shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure. As Figure 14 shown, the electronic device 2000 may include a processor 2010 and a memory 2020 connected to the processor 2010. In addition, the processor 2010 may further include a buffer. According to the embodiments of the present disclosure, the memory 2020 may be implemented in the form of a high-bandwidth memory HBM, which is not limited thereto. Specifically, according to the embodiments of the present disclosure, the processor 2010 is configured to run computer-executable instructions, which, when run by the processor 2010, implement the data loading method according to the embodiments of the present disclosure to load the tensor to be processed from the original tensor in the memory to the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain the second tensor based on the first tensor in the buffer and write the second tensor into the memory.

[0377] The processor 2010 can perform various actions and processes according to a program stored in a non-transitory memory or the like. Specifically, the processor 2010 can refer to a processor chip capable of performing parallel computing. For example, it can be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural network Processing Unit), a DPU (Deep learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In addition, the processor 2010 can also be implemented as other conventional types of processors, which are not limited herein.

[0378] Regarding the specific implementation processes of the data storage method and the data loading method, reference can be made to the above description, which will not be repeated here. The processor provided in at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0379] Figure 15 A block diagram showing an example computing device implementing some embodiments of the present disclosure is shown. As Figure 15 shown, the computing device 3000 is, for example, suitable for implementing the data loading method or the data storage method provided by the embodiments of the present disclosure. It should be noted that Figure 15 the components of the computing device 3000 shown are exemplary and not restrictive. According to actual application needs, the computing device 3000 may also have other components.

[0380] As Figure 15 shown, the computing device 3000 may include a processing device 3010 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.

[0381] For example, when the computer-readable instructions are executed by the processing device 3010, one or more steps in the data loading method according to any of the above embodiments, or one or more steps in the data storage method according to any of the above embodiments may be performed. It should be noted that for a detailed description of the processing procedure of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing procedure of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0382] For example, the processing device 3010, the read-only memory (ROM) 3020, and the random access memory (RAM) 3030 are connected to each other via the bus 3040. The input / output (I / O) interface 3050 is also connected to the bus 3040.

[0383] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, the random access memory (RAM) 3030 and / or a cache, etc. For example, the computer-readable instructions may be loaded from the storage device 3080 into the random access memory (RAM) 3030 to execute the computer-readable instructions. The non-volatile memory may include, for example, the read-only memory (ROM) 3020, a hard disk, an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage medium, as well as various data used and / or generated by the application programs, etc.

[0384] Generally, the following devices may be connected to the I / O interface 3050: an input device 3060 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 3070 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 3080 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 3090. The communication device 3090 may allow the computing device 3000 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 15A computing device 3000 with various devices is shown. However, it should be understood that it is not required to implement or have all the shown devices, and the computing device 3000 may alternatively implement or have more or fewer devices. For example, the processing device 3010 may control other components in the computing device 3000 to perform desired functions. The processing device 3010 may be a Central Processing Unit (CPU), a Tensor Processing Unit (TPU), or a Graphics Processing Unit (GPU), etc., which has data processing capabilities and / or program execution capabilities. The GPU may be directly integrated into a System on Chip (SOC), directly integrated onto the motherboard, or built into the North Bridge chip of the motherboard.

[0385] According to some embodiments of the present disclosure, a non-transitory computer-readable storage medium is also provided, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they implement the data loading method according to the embodiments of the present disclosure, or implement the data storage method according to the embodiments of the present disclosure.

[0386] Figure 16 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 16 shown, the computer-readable storage medium 4000 may be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 4010 may be non-temporarily stored on the storage medium 4000. For example, when the computer-readable instructions 4010 are executed by a processor, one or more steps of the data loading method described in any of the above embodiments may be executed, or one or more steps of the data storage method described in any of the above embodiments may be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0387] As an example, the storage medium 4000 may be applied to the electronic device 2000 and / or the computing device 3000. For example, the storage medium 4000 may be implemented as the storage device 3080 in the computing device 3000.

[0388] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0389] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases.

[0390] The functions described above in this article can be at least partially performed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on. The above description is only the preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0391] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0392] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

[0393] The following points also need to be noted regarding the present disclosure:

[0394] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.

[0395] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.

[0396] The above are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the appended claims.

Claims

1. A data loading method for loading a tensor to be processed from an original tensor in memory into a buffer, where The data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is a pixel, and it accumulates to higher dimensions level by level. The channel number dimension does not calculate the number of pixels. The data loading method includes: Obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and Sequentially sending the plurality of requests and sequentially writing the data corresponding to each request into the buffer to load the tensor to be processed into the buffer. Among them, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory. Among them, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

2. The method according to claim 1, wherein In response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being greater than x, the second dimension is the C dimension, and In response to the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

3. The method according to claim 2, wherein In response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that continuous loading cannot be performed on the tensor to be processed in the first dimension.

4. The method according to claim 1, wherein When the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension, and it is determined to adopt a division of requests in the H dimension.

5. The method according to claim 1, wherein When the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

6. The method according to claim 1, wherein When the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

7. The method according to claim 1, wherein The tensor to be processed is used for convolution operations in the computing units within the processor.

8. The method according to claim 1, wherein Each request includes a data read address for indicating the starting position to read data from the memory, a data write address for indicating the starting position to write data to the buffer, and the number of pixels to be acquired by the request. Determine the request initial coordinates corresponding to the next request based on the request initial coordinates corresponding to the previous request. The request initial coordinates corresponding to each request are used to determine the data read address, data write address, and the number of pixels to be acquired by the request. Among them, when successively determining the request initial coordinates corresponding to each request, the coordinate values of the request initial coordinates are incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

9. The method according to claim 1, wherein Combined with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determine multiple requests for loading the tensor to be processed, including: Based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determine the first request among the multiple requests and the initial state when the first request enters the state machine, where the first request indicates the number of pixels to be acquired by the request; and Based on the initial state, combined with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed, use the state machine to determine each request among the multiple requests after the first request.

10. The method according to claim 9, wherein, Each request includes a data read address for indicating a starting position to read data from the memory, a data write address for indicating a starting position to write data to the buffer, and the number of pixels that the request is to obtain. Among them, determining the first request among the multiple requests and the initial state of the first request entering the state machine based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed includes: Determining the initial state of the first request entering the state machine based on the starting coordinates of the tensor to be processed; Taking the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; Determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; Determining the starting address of writing the tensor to be processed in the buffer as the data write address of the first request; and Determining the number of pixels that the first request is to obtain according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

11. The method according to claim 10, wherein, Determining the initial state of the first request entering the state machine based on the starting coordinates of the tensor to be processed includes: In response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; In response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and In response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

12. The method according to claim 9, wherein, Based on the initial state, combining the number of pixels that the first request is to obtain and the number of pixels included in the tensor to be processed, using the state machine to determine each request after the first request among the multiple requests includes: Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels that the second request is to obtain; Determining the data read address of the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; Determining the data write address of the second request according to the number of pixels that the second request is to obtain; Updating the state machine according to the information related to the second request, and sequentially determining subsequent requests among the respective requests based on the updated state machine.

13. The method according to claim 10, wherein, For multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Wherein, u_addr_base represents the storage address in the memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in 5 dimensions respectively.

14. The method according to claim 10, wherein, For multiple requests for loading the tensor to be processed, determining the data write address of the nth request among the multiple requests includes: For m requests with the same coordinate value range of the data of the requests in the C dimension, calculate the data write address Addr_2_i of the ith request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i - 1 * copy_c Wherein, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2,..., req_size_i - 1 represent the number of pixels to be acquired by the previous i - 1 requests respectively, and copy_c is the size of the tensor to be processed in the C dimension; Wherein, in response to the data of each of the m requests including the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer, In response to the data of each of the m requests including the tth data element to the (t + x)th data element of the original tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1) * x t, m, and i are positive integers.

15. The method according to claim 1, wherein, Sequentially sending the multiple requests and sequentially writing the data corresponding to each request into the buffer to load the tensor to be processed into the buffer includes: For any one request, in response to all the data loaded by the any one request being located in the memory, sending the any one request to the memory; In response to all the data loaded by the any one request not being located in the memory, converting the any one request into writing multiple predetermined values into the buffer, wherein the number of the multiple predetermined values is determined by the number of pixels to be acquired by the any one request.

16. The method according to claim 1 further includes: Obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the original tensor, wherein, for the dimensions with boundary values set, the data loaded by each request either belongs to the valid data range defined by the original tensor and the boundary values or does not belong to the valid data range.

17. A data loading method includes: Receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in a memory into a buffer area, where the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer area is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the channel number of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is a pixel, and is accumulated to higher dimensions step by step, where the channel number dimension does not calculate the number of pixels; and After parsing the data loading instruction, using an execution unit to execute the data loading instruction, wherein using the execution unit to execute the data loading instruction includes: Obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and Sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the buffer area to load the tensor to be processed into the buffer area, wherein, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, the data loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the data loaded by the request belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, wherein the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

18. A data storage method for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, where The data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions. The channel number dimension does not count the number of pixels. The data storage method includes: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and Sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the memory to store the second tensor in the memory. Wherein, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, the requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer. Wherein, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

19. A data storage method, including: Receiving a data storage instruction indicating to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into the memory, where the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions. The channel number dimension does not count the number of pixels; and After parsing the data storage instruction, using an execution unit to execute the data storage instruction. Wherein, using the execution unit to execute the data storage instruction includes: Obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Combine the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, wherein the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and Sequentially send the plurality of requests, and sequentially write the data corresponding to each request into the memory to store the second tensor to the memory, wherein, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor, and in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer; wherein the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

20. A processor, comprising an instruction parsing unit and an execution unit, wherein, The instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the data loading instruction instructs to load a tensor to be processed from an original tensor in a memory to a buffer, wherein the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension, wherein in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension represents a pixel, and is accumulated step by step to higher dimensions, and the number of pixels is not calculated for the channel number dimension; and The execution unit is configured to: execute the data loading instruction, wherein the execution unit executes the data loading instruction, including: Obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Combine the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer. Among them, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension. The requests are divided in the second dimension. The data loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor. And in response to the data loaded by the request belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory. Among them, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

21. A processor includes an instruction parsing unit and an execution unit. Among them, The instruction parsing unit is configured to: receive and parse a data storage instruction. Among them, the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into the memory. Among them, the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC. Among them, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the number of channels dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among them, in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension represents a pixel, and is accumulated step by step to higher dimensions. Among them, the number of channels dimension does not calculate the number of pixels; and The execution unit is configured to: execute the data storage instruction. Among them, when the execution unit executes the data storage instruction, it includes: Obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Combine the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor. Among them, the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request into the memory in sequence to store the second tensor in the memory. Among them, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, a division of requests is performed in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer. Among them, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

22. An electronic device, comprising a processor and a memory connected to the processor, wherein, The processor includes a buffer, among which, The processor is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor, it implements the data loading method according to any one of claims 1-17 to load the tensor to be processed from the original tensor in the memory to the buffer, or implements the data storage method according to claim 18 or 19 to obtain a second tensor based on the first tensor in the buffer and write the second tensor to the memory.

23. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by the processor, it implements the data loading method according to any one of claims 1-17, or implements the data storage method according to claim 18 or 19.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN116822612A

  • Data processing method and device, processor, electronic equipment and storage medium

    CN119312003A