Data Loading Method, Data Storage Method, Processor, Electronic Device, and Medium
By adopting the request division method of NDHWC data storage format and state machine optimization in parallel processors, the problem of low loading and storage efficiency of tensor data is solved, and memory bandwidth and hardware performance are improved.
Patent Information
- Application Number
- CN202510607298.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In prior art, in parallel processors, the loading and storage efficiency of tensor data is low, resulting in low hardware utilization of computing units and inability to effectively utilize memory bandwidth.
Using the NDHWC data storage format, when loading the tensor to be processed into the cache, multiple requests are determined by obtaining its shape size and starting coordinates, combining the original tensor, and request division is performed on the second dimension when the first dimension cannot be continuously loaded, ensuring that the sub-data of each request is stored continuously or discontinuously in memory, and the initial coordinates and data length of the request are optimized using the state machine.
It improves data memory access bandwidth and efficiency, improves hardware utilization of computing units, and improves hardware performance.
Smart Images

Figure CN120123265B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] A tensor is a multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. For example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor can have any number of dimensions. Tensor operations are widely used in processors such as parallel processors.
[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processor devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of a parallel processor require a large amount of data, and this data is generally stored in the storage components of the parallel processor. For example, the storage component can be a memory. The data can be extracted from the storage component to the buffer for calculation through a data loading instruction, and the data in the buffer can be stored in the memory through a data storage instruction. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a data loading method for loading a tensor to be processed into a buffer. The tensor to be processed is stored in the buffer in an NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data loading method includes: obtaining the shape size of the tensor to be processed and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor, where the data storage format of the original tensor is N(C / x)DHW(xC) and x is a positive integer greater than 1; determining, in combination with the original tensor, the shape size of the tensor to be processed, and the starting coordinates of the tensor to be processed, a plurality of requests for loading the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer to load the tensor to be processed into the buffer; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, partitioning the requests in the second dimension, the sub-data loaded by each request is all located in the memory or all not located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.
[0005] For example, in the data loading method provided by at least one embodiment of the present disclosure, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension takes precedence over the second dimension during loading; in response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being greater than x, the second dimension is the C dimension; in response to the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is the N dimension.
[0006] For example, in the data loading method provided by at least one embodiment of the present disclosure, each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data to the buffer, and the length of the data loaded by the request. The initial coordinates of the request corresponding to the next request are determined based on the initial coordinates of the request corresponding to the previous request. The initial coordinates of each request are used to determine the data read address, the data write address, and the length of the data loaded by the request. Among them, when successively determining the initial coordinates of each request corresponding to the request, the coordinate values of the initial coordinates of the request are incrementally updated in the order of the W dimension, the H dimension, the D dimension, the N dimension, and the C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.
[0007] For example, in the data loading method provided by at least one embodiment of the present disclosure, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.
[0008] For example, in the data loading method provided by at least one embodiment of the present disclosure, combining the shape and size of the original tensor, the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for loading the tensor to be processed, including: based on the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed, determining the first request sent among the plurality of requests and the initial state when the first request enters the state machine; combining the initial state, and based on the shape and size of the original tensor, using the state machine to determine each request other than the first request among the plurality of requests.
[0009] For example, in the data loading method provided by at least one embodiment of the present disclosure, each request includes a data read address for indicating the starting position to read data from the memory, a data write address for indicating the starting position to write data to the buffer, and the length of the data loaded by the request. Based on the starting coordinates of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed, determining the first request sent among the multiple requests and the initial state of the first request entering the state machine includes: determining the initial state of the first request entering the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor and the request initial coordinates corresponding to the first request; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and determining the length of the data loaded by the first request according to the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed.
[0010] For example, in the data loading method provided by at least one embodiment of the present disclosure, determining the initial state of the first request entering the state machine based on the starting coordinates of the tensor to be processed includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state to be the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state to be the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state to be the third state.
[0011] For example, in the data loading method provided by at least one embodiment of the present disclosure, in combination with the initial state, determining each request other than the first request among the multiple requests by using the state machine based on the shape and size of the original tensor includes: determining the request initial coordinates corresponding to each request and the length of the data loaded by each request by using the state machine based on the initial state; determining the data read address of each request according to the shape and size of the original tensor and the request initial coordinates corresponding to each request; and determining the data write address of each request according to the length of the data loaded by each request.
[0012] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine includes a first state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: in response to the state of the current request being the first state and satisfying a first condition, determining that the next request of the current request enters the first state, where in response to the current request being the first request, the state of the current request is the initial state; in response to the next request entering the first state, determining that the coordinate value of the request initial coordinates corresponding to the next request in the first dimension is a first coordinate value, and determining that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining that the data length loaded by the current request is the size of the tensor to be processed in the first dimension; where the first condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first condition includes that the sum of the fourth coordinate value of the request initial coordinates corresponding to the current request in the first dimension and the remaining size is less than the second coordinate value and the fourth coordinate value is less than the second coordinate value, the remaining size is the number of remaining tensor data in the first dimension of the tensor to be processed that has not been requested to be loaded, and the first coordinate value is the coordinate value of the starting coordinate of the tensor to be processed in the first dimension.
[0013] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine further includes a second state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests further includes: in response to the state of the current request being the first state and satisfying a second condition, determining that the next request of the current request enters the second state, in response to the next request entering the second state, determining that the coordinate value of the request initial coordinates corresponding to the next request in the first dimension is the second coordinate value, and determining that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions except the first dimension remain unchanged; determining the data length loaded by the current request based on the first coordinate value and the second coordinate value; where the second condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the second coordinate value, or the second condition includes that the sum of the fourth coordinate value and the remaining size is greater than the second coordinate value.
[0014] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine includes a first state and a second state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths of the data loaded by the respective requests includes: in response to the state of the current request being the second state and satisfying a third condition, determining that the next request of the current request enters the first state, wherein, in response to the current request being the first request, the state of the current request is the initial state; in response to the next request entering the first state, determining that the coordinate value of the request initial coordinates corresponding to the next request in the first dimension is a first coordinate value, and determining that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions other than the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining the data length loaded by the current request based on the first coordinate value and the size of the tensor to be processed in the first dimension; wherein, the third condition includes that the first coordinate value is less than a second coordinate value and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than a third coordinate value, or, the third condition includes that the sum of the fourth coordinate value of the request initial coordinates corresponding to the current request in the first dimension and the remaining size is less than the third coordinate value and the fourth coordinate value is less than the second coordinate value, the remaining size is the number of remaining tensor data in the first dimension of the tensor to be processed that has not been requested to be loaded, the first coordinate value is the coordinate value of the starting coordinate of the tensor to be processed in the first dimension, the second coordinate value is the coordinate value of the starting coordinate of the original tensor in the first dimension, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension.
[0015] For example, in the data loading method provided by at least one embodiment of the present disclosure, based on the initial state, the state machine is used to determine the request initial coordinates corresponding to the respective requests and the data lengths to be loaded for the respective requests, and further includes: in response to the state of the current request being the second state and satisfying the fourth condition, determining that the next request enters the second state; in response to the next request entering the second state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining that the data length to be loaded for the current request is the size of the tensor to be processed in the first dimension; where the fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value, or the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension.
[0016] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine further includes a third state. Based on the initial state, the state machine is used to determine the request initial coordinates corresponding to the respective requests and the data lengths to be loaded for the respective requests, and further includes: in response to the state of the current request being the second state and satisfying the fifth condition, determining that the next request enters the third state; in response to the next request entering the third state, determining that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the third coordinate value, and determining that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged; determining the data length to be loaded for the current request based on the coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the third coordinate value; where the fifth condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the third coordinate value, or the fifth condition includes that the sum of the fourth coordinate value and the remaining size is greater than the third coordinate value.
[0017] For example, in the data loading method provided by at least one embodiment of the present disclosure, the state machine includes a first state, a second state, and a third state. Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: in response to the state of the current request being the third state, determining the state entered by the next request of the current request according to the first coordinate value, wherein in response to the current request being the first request, the state of the current request is the initial state; determining that the coordinate value of the request initial coordinates corresponding to the next request in the first dimension is the first coordinate value, and determining that the coordinate values of the request initial coordinates corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; determining that the data length loaded by the current request is determined based on the fourth coordinate value and the third coordinate value of the request initial coordinates corresponding to the current request in the first dimension, wherein the first coordinate value is the coordinate value of the starting coordinate of the tensor to be processed in the first dimension, and the third coordinate value is the sum of the second coordinate value of the starting coordinate of the original tensor in the first dimension and the size of the original tensor in the first dimension.
[0018] For example, in the data loading method provided by at least one embodiment of the present disclosure, according to the shape and size of the original tensor and the request initial coordinates corresponding to the respective requests, determining the data read addresses of the respective requests includes: calculating the data read address Addr_1 of the nth request according to the following formula: Addr_1 = u_addr_base +
[0019] ((((n_coord × tensor_d + d_coord) × tensor_h + h_coord) × tensor_w + w_coord) × tensor_c + c_coord)
[0020] wherein, u_addr_base represents the storage address in the memory of the element at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the original tensor in a total of 5 dimensions of NDHWC.
[0021] For example, in the data loading method provided by at least one embodiment of the present disclosure, determining the data writing addresses of the respective requests according to the data lengths loaded by the respective requests includes: for m requests with the same coordinate value range of the sub-data of the requests in the C dimension, calculating the data writing address Addr_2_i of the i-th request sent among the m requests according to the following formula:
[0022] Addr_2_i = b_addr + req_size_0×copy_c + req_size_1×copy_c +…+ req_size_i-1×copy_c
[0023] where b_addr represents the data writing address of the first request sent among the m requests, req_size_0, req_size_1,..., req_size_i-1 represent the data lengths loaded by the respective first i-1 requests, and copy_c is the size of the to-be-processed tensor in the C dimension;
[0024] where, in response to the sub-data of each request among the m requests including the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base for writing the to-be-processed tensor into the buffer,
[0025] in response to the sub-data of each request among the m requests including the t-th data element to the (t + x)-th data element of the original tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula:
[0026] b_addr = b_addr_base+ (t / x - 1)×x
[0027] t, m, and i are positive integers.
[0028] For example, in the data loading method provided by at least one embodiment of the present disclosure, sequentially sending the multiple requests and sequentially writing the sub-data returned by each request into the buffer to load the to-be-processed tensor into the buffer includes: for any one of the requests, in response to all of the sub-data loaded by the any one request being located in the memory, sending the any one request to the memory; in response to all of the sub-data loaded by the any one request not being located in the memory, converting the any one request into writing a plurality of predetermined values into the buffer, where the number of the plurality of predetermined values is determined by the data length specified by the any one request.
[0029] At least one embodiment of the present disclosure provides a data loading method, including: receiving a data loading instruction instructing to load a tensor to be processed into a buffer, where the data loading instruction includes the shape and size of the tensor to be processed as an input parameter, and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor, and the tensor to be processed is stored in the buffer in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension, and the data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1; after parsing the data loading instruction, using an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: combining the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for loading the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer to load the tensor to be processed into the buffer; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, dividing the requests in the second dimension, where the sub-data loaded by each request are all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request are from the original tensor and are continuously stored in the memory, where the first dimension is the same as or adjacent to the second dimension.
[0030] At least one embodiment of the present disclosure provides a data storage method for writing a tensor to be processed in a buffer into a memory in an NDHWC data storage format. The data storage format of the tensor to be processed in the buffer is N(C / x)DHW(xC), where x is a positive integer greater than 1. The data storage method includes: obtaining the shape and size of the tensor to be processed and the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor, where the data storage format of the original tensor is NDHWC; determining a plurality of requests for storing the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed; sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in sequence to store the tensor to be processed; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when storing the tensor to be processed, continuous storage cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, requests are divided in the second dimension, and the sub-data to be written by each request all belong to the data range of the original tensor or all do not belong to the data range of the original tensor. And in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belong to the original tensor and are continuously stored in the memory, where the first dimension is the same as or adjacent to the second dimension.
[0031] For example, in the data storage method provided by at least one embodiment of the present disclosure, the request initial coordinates corresponding to each request are used to determine the data storage address of the sub-data to be stored by the request in the memory. Sequentially sending at least some of the plurality of requests and writing the data to be written indicated by each request into the memory in sequence to store the tensor to be processed includes: for any one request, in response to the first request coordinate value of the request initial coordinates corresponding to the any one request in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determining not to send the any one request; in response to the first request coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining to send the any one request; where the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
[0032] At least one embodiment of the present disclosure provides a processor, including an instruction parsing unit and an execution unit. Wherein, the instruction parsing unit is configured to receive and parse a data loading instruction, and the data loading instruction includes the shape and size of the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters. The tensor to be processed is stored in a buffer area in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1. After the instruction parsing unit parses the data loading instruction, the execution unit executes the data loading instruction. When the execution unit executes the data loading instruction, the following operations are included: determining a plurality of requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed; sequentially sending the plurality of requests, and writing the sub-data returned by each request into the buffer area in sequence to load the tensor to be processed into the buffer area. Wherein, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, requests are divided in the second dimension, and the sub-data loaded by each request is either all located in the memory or none of them is located in the memory. And in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.
[0033] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, where the computer-executable instructions, when run by the processor, implement the data loading method according to any embodiment of the present disclosure, or the data storage method according to any embodiment of the present disclosure.
[0034] At least one embodiment of the present disclosure provides a non-transient computer-readable storage medium, where the non-transient computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the data loading method according to any embodiment of the present disclosure, or the data storage method according to any embodiment of the present disclosure. Description of the Drawings
[0035] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0036] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU);
[0037] Figure 2A It is a schematic structure of a tensor;
[0038] Figure 2B It is a schematic diagram of the storage format of NDHWC;
[0039] Figure 2C It is a schematic diagram of the storage format of N(C / 32)DHW(32C);
[0040] Figure 3 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure;
[0041] Figure 4A It is a schematic diagram of the tensor to be processed provided by one embodiment of the present disclosure;
[0042] Figure 4B It is a schematic diagram of the tensor to be processed provided by another embodiment of the present disclosure;
[0043] Figure 4C It is a schematic diagram of the tensor to be processed provided by another embodiment of the present disclosure;
[0044] Figure 5 It is a schematic diagram of the state machine provided by one embodiment of the present disclosure;
[0045] Figure 6 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure;
[0046] Figure 7 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure;
[0047] Figure 8 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure;
[0048] Figure 9 It is a schematic block diagram of an electronic device provided by one embodiment of the present disclosure;
[0049] Figure 10 It is a schematic structural diagram of the processor provided by at least one embodiment of the present disclosure;
[0050] Figure 11 It is a schematic structural diagram of the processor provided by at least one embodiment of the present disclosure;
[0051] Figure 12 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0052] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0053] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "include" or "comprise" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, some details of known functions and known components are omitted in the present disclosure.
[0054] Figure 1 Schematic structural diagram of a general-purpose graphics processing unit (GPGPU).
[0055] As Figure 1 shown, the general-purpose graphics processing unit is actually an array of programmable multi-processors. For example, the programmable multi-processors can be Streaming Processor Clusters (SPCs), such as including Figure 1 the shown Streaming Processor Cluster 1,..., Streaming Processor Cluster M, where M is a positive integer greater than 1. In the general-purpose graphics processing unit, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing is performed between multiple streaming processor clusters through a global cache or global memory.
[0056] As Figure 1As shown, taking the streaming processor cluster 1 as an example, one streaming processor cluster includes multiple computing units, such as Figure 1 the computing unit 1, computing unit 2, ..., computing unit N in , where N is a positive integer. Each computing unit (Compute Unit, abbreviated as CU) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores or computing kernels), and each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The computing core is used to execute specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in ) and shared memory, which are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit.
[0057] As Figure 1 shown, each streaming processor cluster also provides a buffer for caching data of the N computing units in the streaming processor cluster.
[0058] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processing unit (or called a parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in ). All threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundle (or simply called a thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.
[0059] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in ) schedules and allocates the thread bundles so that multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the middle-level cache or global cache or global memory (such as Figure 1 the high bandwidth memory, High Bandwidth Memory, abbreviated as HBM in ) for read and write operations, etc.
[0060] As Figure 1As shown, general computing operations, such as those on matrices, typically require a large amount of data, which is usually stored in a memory, such as in a High Bandwidth Memory (HBM). When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the calculation result, data needs to be stored in the memory (Store operation). The way data is stored in the memory affects the memory access bandwidth, and thus affects the hardware utilization rate of the computing unit.
[0061] For example, general computing operations include General Matrix Multiplication (GEMM for short). The data required for general matrix multiplication includes a first tensor and a second tensor.
[0062] For example, for the first tensor or the second tensor, its shape dimensions can be represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor data to be processed in 5 dimensions, and a1, a2, a3, a4, a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, the number of data samples grabbed in one training. The D dimension represents the depth, the H dimension represents the height of the input data, the W dimension represents the width of the input data, and the C dimension represents the number of channels. For example, taking the first tensor as an example, a1 can be the N dimension size, a2 can be the D dimension size, a3 can be the H dimension size, a4 can be the W dimension size, and a5 can be the C dimension size. Of course, the present disclosure does not make specific limitations on this.
[0063] The elements of the tensor are arranged in various formats in the memory (such as Figure 1 the memory), which is called the data storage format (layout). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0064] For example, Figure 2A is a schematic structure of a tensor. In the Figure 2A shown tensor, a1 is the N dimension size and is equal to 1, a2 is the D dimension size and is equal to 1, a3 is the H dimension size and is equal to 5, a4 is the W dimension size and is equal to 4, and a5 is the C dimension size and is equal to 64.
[0065] For example, Figure 2A the pixel elements of the tensor in are represented as 0, 1, 2, 3,... and so on. The following uses the Figure 2A shown tensor to describe different data storage formats.
[0066] For example, the data storage format can include NDHWC, also known as the Linear mode. Figure 2BSchematic diagram of the storage format for NDHWC.
[0067] For example, for NDHWC, as Figure 2B shown, starting from the first element of the first channel (c0 in Figure 2B , where a5 = 0) (element 0 in Figure 2B ), then store the first element of the second channel (c1 in Figure 2B , where a5 = 1) (element 20 in Figure 2B ), and so on until the first elements of all channels are laid out. For example, after the first element of the 64th channel (c63 in Figure 2B , where a5 = 63) (element 1260 in Figure 2B ), select the second element of the first channel (c0 in Figure 2B , where a5 = 0) (element 1 in Figure 2B ), then store the second element of the second channel (c1 in Figure 2B , where a5 = 1) (element 21 in Figure 2B ), and so on until the second elements of all channels are laid out, and so on.
[0068] For example, the data storage format can also include N(C / x)DHW(xC), also known as the interleave mode, where x can take values such as 8, 16, 32, etc. according to needs.
[0069] N(C / x)DHW(xC) is similar to NDHWC, but there is a key difference. In the memory layout of N(C / x)DHW(xC), the a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.
[0070] Figure 2C Schematic diagram of the storage format for N(C / 32)DHW(32C).
[0071] As Figure 2C shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 (c0 in Figure 2C ) to a5 = 31 (c31 in Figure 2C ), and the second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format.
[0072] In memory, the original tensor is stored continuously in memory in the data storage format described above. The shape of the original tensor can be expressed as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in these 5 dimensions and are all positive integers. When extracting or storing a partial tensor in the original tensor, for example, the shape of the partial tensor can be expressed as a1×a2×a3×a4×a5. Since it is a partial tensor inside the original tensor, it may not be possible to continuously extract in each dimension during extraction, and it is impossible to accurately obtain the amount of data and the data location when loading or storing data to optimally implement data loading or storing, greatly reducing the data bandwidth when loading data and reducing the hardware computing efficiency.
[0073] At least one embodiment of the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium.
[0074] In at least one embodiment, the data loading method is used to load a tensor to be processed into a buffer area. The tensor to be processed is stored in the buffer area in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data loading method includes: obtaining the shape size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, where the data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1; combining the original tensor, the shape size of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for loading the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area; where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in each dimension lower than the first dimension, divide the requests in the second dimension, and the sub-data loaded by each request is all located in memory or none of it is located in memory, and in response to the sub-data loaded by the request being all located in memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.
[0075] In the data loading method provided by at least one embodiment of the present disclosure, the requests are divided according to whether the tensor can be continuously loaded in the first dimension. The sub-data used for loading by each request is all located in memory or none of it is located in memory. The division of data requests is more reasonable, more suitable for the loading and storing of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and improving the hardware performance.
[0076] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0077] Figure 3 It is a schematic flowchart of a data loading method provided for at least one embodiment of the present disclosure.
[0078] As Figure 3 shown, the data loading method provided for at least one embodiment of the present disclosure includes at least steps S10 - S30.
[0079] For example, the data loading method provided for at least one embodiment of the present disclosure is used to load a tensor to be processed into a buffer area, and the buffer area is, for example, a buffer in a streaming processor cluster.
[0080] For example, the tensor to be processed is stored in the buffer area in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension.
[0081] For example, the original tensor is stored in memory, and the tensor to be processed can be a partial tensor of the original tensor, or the tensor to be processed can be the original tensor itself. The shape and size of the tensor to be processed can be represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor data to be processed in 5 dimensions, and a1, a2, a3, a4, a5 are positive integers.
[0082] Figure 4A It is a schematic diagram of a tensor to be processed provided for an embodiment of the present disclosure.
[0083] Figure 4A In , each small solid cube represents an element, and the tensor composed of multiple small solid cubes is the original tensor. Figure 4A What is shown is one batch and one tensor at a certain depth in a batch, but the original tensor may have multiple batches, and each batch may have multiple tensors in the depth dimension, and its structure is the same as that shown in Figure 4A and will not be described again here.
[0084] Assume Figure 4A that the element pointed by the arrow is the origin of the coordinate system of the original tensor, C, W, and H respectively represent three coordinate axes, row is the coordinate value in the H - dimension direction, col is the coordinate value in the W - dimension direction, c is the coordinate value in the C - dimension direction, and the coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. The channel, width, and height values in the coordinates of other elements increase in the direction of the arrow. Of course, the original tensor may also include D - dimension and N - dimension coordinates, which are not shown here.
[0085] For example, the shape dimensions of the original tensor are represented by b1, b2, b3, b4, and b5. Taking the case where b1 is the N - dimension size, b2 is the D - dimension size, b3 is the H - dimension size, b4 is the W - dimension size, and b5 is the C - dimension size as an example for description, in Figure 4A the shown example, it is assumed that b1 = 1, b2 = 1, b3 = 4, b4 = 8, and b5 = 8.
[0086] For example, in Figure 4A the example, the dashed - box is the tensor to be processed, which includes some elements in the original tensor. The dimensions of the tensor to be processed in each dimension are represented by a1, a2, a3, a4, and a5. For example, it is assumed that both the N - dimension size a5 and the D - dimension size a4 are equal to 1. Figure 4A In the example of
[0087] the upper - left - hand element coordinates of the tensor to be processed are c = 0, row = 0, col = 3, and a1 = 5, a2 = 3, a3 = 3.
[0088] Figure 4B It is a schematic diagram of the tensor to be processed provided by another embodiment of the present disclosure.
[0089] Figure 4B In Figure 4B the tensor composed of multiple solid - line cubes is the original tensor. Figure 4A The element at the upper - left - hand corner of the original tensor is the origin of the coordinate system determined based on the original tensor. The coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. C, W, and H respectively represent the three coordinate axes, and their specific meanings are the same as those in
[0090] For example, in Figure 4B the example, the dashed - box is the tensor to be processed, which includes some elements (solid - line small cubes) in the original tensor. In addition, the tensor to be processed also includes dashed - line small cubes outside the edge of the original tensor. The dashed - line small cubes are obtained, for example, through a padding operation, and their values are, for example, 0. For example, it is assumed that both the N - dimension size a5 and the D - dimension size a4 are equal to 1. Figure 4B In the example of
[0091] Of course, in some other embodiments, the tensor to be processed may not even include any elements in the original tensor at all, and it can perform data loading operations using the coordinate system determined by the original tensor.
[0092] Figure 4C Schematic diagram of the tensor to be processed provided by another embodiment of the present disclosure.
[0093] Figure 4C In, the tensor composed of multiple solid-line cubes is the original tensor, Figure 4C In, the element located at the upper left corner of the original tensor is the origin of the coordinate system determined based on the original tensor. The coordinates of the element pointed by the arrow in the input data are: c = 0, row = 0, and col = 0. C, W, and H respectively represent the three coordinate axes, and their specific meanings are the same as those in Figure 4A which will not be elaborated here.
[0094] For example, in Figure 4C the example, the dashed box is the tensor to be processed, which does not include any elements (solid small cubes) in the original tensor. The tensor to be processed is composed of dashed small cubes outside the edge of the original tensor. The dashed small cubes are obtained, for example, through a filling operation, and their values are, for example, 0. For example, assume that both the N-dimensional size a5 and the D-dimensional size a4 are equal to 1. Figure 4C In the example, the coordinates of the upper left corner element of the tensor to be processed are c = 0, row = 0, col = -3, and a1 = 5, a2 = 3, a3 = 3.
[0095] As Figure 3 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S10 - S30.
[0096] First, in step S10, obtain the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor.
[0097] The relationship between the original tensor and the tensor to be processed is as described above. For example, refer to Figures 4A - 4C the relevant description, which will not be elaborated here.
[0098] For example, determine a coordinate system with a certain element in the original tensor as the origin of the coordinate system. For example, refer to Figures 4A - 4C the embodiment of, and use the upper left corner vertex of the original tensor as the origin of the coordinate system. Of course, the present disclosure is not limited thereto.
[0099] The starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor are, for example, Figures 4A - 4C the coordinates of the upper left corner vertex of the tensor to be processed in the embodiment of. The coordinate values of the starting coordinates of the tensor to be processed are the minimum coordinate values among the coordinate values of all elements in the tensor to be processed.
[0100] In step S20, multiple requests for loading the tensor to be processed are determined by combining the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed.
[0101] In step S30, multiple requests are sequentially sent, and the sub-data returned by each request is sequentially written into the buffer area to load the tensor to be processed into the buffer area.
[0102] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that the tensor to be processed cannot be continuously loaded in the first dimension but can be continuously loaded in each dimension lower than the first dimension, the requests are divided in the second dimension. The sub-data loaded by each request is either all in memory or all not in memory. And in response to the sub-data loaded by the request being all in memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in memory. The first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.
[0103] For example, the first dimension is the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. Specifically, if the first dimension is the W dimension, the second dimension is the H dimension; if the first dimension is the H dimension, the second dimension is the D dimension; if the first dimension is the D dimension, the second dimension is the C dimension.
[0104] In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension not being equal to x, the second dimension is the C dimension.
[0105] In response to the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is the N dimension.
[0106] In the present disclosure, the original tensor is stored in memory in the data storage format of N(C / x)DHW(xC). The data storage format of the original tensor in memory is different from the data storage format (NDHWC) of the tensor to be processed in the buffer area. In memory, the C dimension of the original tensor is the lowest dimension, followed by the W dimension, then the H dimension, then the D dimension, and the highest dimension is the N dimension. When the tensor to be processed is stored in the buffer area, the W dimension is the lowest dimension, followed by the H dimension, then the D dimension, then the C dimension, and the highest dimension is the N dimension. And for the W dimension, the continuous xC is substantially considered.
[0107] For example, in response to the size copy_t of the tensor to be processed in the first dimension not being equal to the size tensor_t of the original tensor in the first dimension, and / or the first coordinate value t_coord_b of the starting coordinates of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.
[0108] Taking the second coordinate value as 0 as an example, assuming the first dimension is the W dimension and the second dimension is the H dimension, it is determined that it is discontinuous in the W dimension when any of the following conditions is satisfied:
[0109] (1) The starting coordinate of the tensor to be processed in the W dimension is not equal to 0
[0110] (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension
[0111] (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension
[0112] At this time, it can be understood that the sub-data loaded by each request belongs to the same row, and the sub-data loaded by different requests is located in different rows, that is, the requests are divided in the H dimension, and the requests are split by row.
[0113] Assuming the first dimension is the H dimension and the second dimension is the D dimension, it is determined that it is discontinuous in the H dimension when any of the following conditions is satisfied:
[0114] (1) The starting coordinate of the tensor to be processed in the H dimension is not equal to 0
[0115] (2) The size copy_h of the tensor to be processed in the H dimension is greater than the size tensor_h of the original tensor in the H dimension
[0116] (3) The size copy_h of the tensor to be processed in the H dimension is less than the size tensor_h of the original tensor in the H dimension
[0117] At this time, it can be understood that the coordinates of the sub-data to be loaded by each request are different in the W dimension and the H dimension, but the coordinates in the D dimension and the N dimension are the same, and the requests are divided in the D dimension.
[0118] Assuming the first dimension is the D dimension and the second dimension is the C dimension, it is determined that it is discontinuous in the D dimension when any of the following conditions is satisfied:
[0119] (1) The starting coordinate of the tensor to be processed in the D dimension is not equal to 0
[0120] (2) The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension
[0121] (3) The size copy_d of the tensor to be processed in the D dimension is less than the size tensor_d of the original tensor in the D dimension
[0122] At this time, the requests are divided in the C dimension. Specifically, considering the particularity of the interleaving pattern, the requests are divided into groups of every x data elements in the C dimension, because every x data elements are stored continuously in the channel number dimension in the memory.
[0123] Assume that the first dimension is the C dimension and the second dimension is also the C dimension at this time. Determine that it is discontinuous in the C dimension when any of the following conditions is satisfied:
[0124] (1) The starting coordinate of the tensor to be processed in the C dimension is not equal to 0
[0125] (2) The size copy_c of the tensor to be processed in the C dimension is greater than the size tensor_c of the original tensor in the C dimension
[0126] (3) The size copy_c of the tensor to be processed in the C dimension is less than the size tensor_c of the original tensor in the C dimension
[0127] At this time, the requests are still divided in the C dimension. Specifically, considering the particularity of the interleaving pattern, the requests are divided into groups of every x data elements in the C dimension, because every x data elements are stored continuously in the channel number dimension in the memory.
[0128] Assume that the first dimension is the N dimension. Determine that it is discontinuous in the N dimension when any of the following conditions is satisfied:
[0129] (1) The starting coordinate of the tensor to be processed in the N dimension is not equal to 0
[0130] (2) The size copy_n of the tensor to be processed in the N dimension is greater than the size tensor_n of the original tensor in the N dimension
[0131] (3) The size copy_n of the tensor to be processed in the N dimension is less than the size tensor_n of the original tensor in the N dimension
[0132] If the size of the tensor to be processed in the C dimension is greater than x, and the second dimension is the C dimension at this time, the requests are still divided in the C dimension. Specifically, considering the particularity of the interleaving pattern, the requests are divided into groups of every x data elements in the C dimension, because every x data elements are stored continuously in the channel number dimension in the memory.
[0133] If the size of the tensor to be processed in the C dimension is equal to x, and the second dimension is the N dimension at this time, for the tensor data located in the memory, in fact, a request can be sent to the memory to load the data in the memory at this time, and other requests can be used to load the data not located in the memory.
[0134] Since the data in the N(C / x)DHW(xC) format needs to be written into the buffer in the NDHWC format, which is different from the storage format of the original data itself, and the dimension arrangement orders of the two data storage formats are also different, the update logic of the request initial coordinates during the request splitting is different from the dimension arrangement order of the data storage format itself.
[0135] Each request includes a data read address for indicating the starting position to read data from the memory, a data write address for indicating the starting position to write data into the buffer, and the length of the data to be requested for loading. The request initial coordinates corresponding to the next request are determined according to the request initial coordinates corresponding to the previous request, and the request initial coordinates corresponding to each request are used to determine the data read address, data write address, and the length of the data to be requested for loading of the request.
[0136] As described later, except for the first dimension, the coordinate values of the other dimensions higher than the first dimension are updated according to the coordinate values lower than that other dimension. Since the data of the original tensor needs to be written into the buffer in the NDHWC storage format, for the convenience of hardware implementation, the coordinate values of the request initial coordinates are incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and the coordinate value of the C dimension is incremented by x when updated. That is to say, the N dimension is updated prior to the C dimension during the update, but since in fact the N dimension of the tensor to be processed is originally the highest dimension, and the N dimension will be given priority when writing into the buffer, it is necessary to reserve in advance the data positions in the buffer that should be written into the buffer but have not been written yet, and write the data into the reserved data positions when writing the data subsequently, thereby ensuring that the tensors to be processed in the buffer can be arranged in the NDHWC dimension order.
[0137] For example, taking the first dimension as the W dimension as an example to illustrate the coordinate update dimension order. The coordinate value h_coord of the request initial coordinates corresponding to the next request in the H dimension is incremented by 1, and returns to the coordinate initial value h_coord_b when reaching the boundary of the tensor to be processed in the H dimension; the coordinate value d_coord of the request initial coordinates corresponding to the next request in the D dimension is incremented by 1 when h_coord returns to h_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the D dimension, and returns to the coordinate initial value d_coord_b when reaching the boundary of the tensor to be processed in the D dimension; the coordinate value n_coord of the request initial coordinates corresponding to the next request in the N dimension is incremented by 1 when d_coord returns to d_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the N dimension, and returns to the coordinate initial value n_coord_b when reaching the boundary of the tensor to be processed in the N dimension; the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension is incremented by x when n_coord returns to the coordinate initial value n_coord_b, and remains unchanged in other cases.
[0138] The calculation method of the data writing address also varies considering the need to reserve data positions in advance. For specific details, please refer to the relevant descriptions in the following text, which will not be elaborated here.
[0139] The object to be loaded in this disclosure, that is, the tensor to be processed, is split into multiple different requests according to a certain rule and sent sequentially. The returned data is written into the buffer area in order, thereby efficiently loading the tensor to be processed into memory. One principle during splitting is that the sub-data loaded by each request is either all in memory (the sub-data loaded by the request belongs to the data range of the original tensor) or all not in memory (the sub-data loaded by the request does not belong to the data range of the original tensor). This is because when the sub-data of the request is in memory, a request is sent to memory, and when the sub-data of the request is not in memory, the request can be converted into an operation such as writing a predetermined value to the buffer area by hardware. This splitting method can send corresponding loading requests to different hardware more reasonably.
[0140] In addition, during splitting, if the size relationship between the tensor to be processed and the original tensor in the first dimension causes that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, then the first dimension is used as a boundary for further division of requests, so that requests can be split more reasonably, the continuously stored data in memory can be retained as much as possible, the number of requests can be reduced, the tensor to be processed can be loaded efficiently, the bandwidth when loading or storing data from memory can be greatly increased, the performance can be improved, and the efficiency of the hardware computing unit can be improved.
[0141] The following specifically describes the method for determining multiple requests for loading the tensor to be processed.
[0142] For example, in some embodiments, step S20 may include: based on the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed, determining the first request sent among multiple requests and the initial state when the first request enters the state machine; based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the tensor to be processed, using the state machine to determine each request among multiple requests except the first request.
[0143] Each request includes three parameters: a data reading address for indicating the starting position to read data from memory, a data writing address for indicating the starting position to write data to the buffer area, and the length of the loaded data.
[0144] The data reading address of each request is obtained by summing up the sizes of all previous requests.
[0145] For example, taking the data storage format of NDHWC as an example, the calculation formula for the data reading address Addr_1 of the nth request is as follows:
[0146] Addr_1 = u_addr_base +
[0147] ((((n_coord × tensor_d + d_coord) × tensor_h + h_coord) × tensor_w + w_coord) × tensor_c + c_coord)
[0148] Among them, u_addr_base represents the storage address in memory of the element at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in five dimensions.
[0149] As mentioned above, since the N dimension takes precedence over the C dimension when loading sub - data, while in the buffer, the C dimension should be arranged prior to the N dimension. Taking Figure 4A the tensor to be processed shown as an example, assuming x = 2, all elements with coordinate values 0 and 1 on the C dimension will be written into the buffer first. But actually, for the elements with coordinate values 2 and 3, they should be stored next to the elements with coordinate values 0 and 1. Therefore, when writing the elements with coordinate values 0 and 1, positions in the buffer for the elements with coordinate values 2 and 3 need to be reserved, and then the elements with coordinate values 2 and 3 are written into the reserved positions.
[0150] For m requests with the same range of coordinate values of the requested sub - data on the C dimension, for example, the data with the same coordinate value range is split into m requests for sending. For example, the coordinate values of the m requests are all from 0 to x - 1, or from x to 2×x - 1, etc. Calculate the data write address Addr_2_i of the ith request sent among these m requests according to the following formula:
[0151] Addr_2_i = b_addr + req_size_0×copy_c + req_size_1×copy_c +…+ req_size_i - 1×copy_c
[0152] Among them, b_addr represents the data write address of the first request sent among the m requests, req_size_0, req_size_1,..., req_size_i - 1 represent the data lengths loaded by the previous i - 1 requests respectively, copy_c is the size of the tensor to be processed on the C dimension, and m and i are positive integers.
[0153] The sub-data corresponding to each of the m requests includes the first x data elements of the original tensor in the C dimension, for example, the coordinate values are from 0 to x - 1, and b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer.
[0154] The sub-data corresponding to each of the m requests includes the t-th data element to the (t + x)-th data element of the original tensor in the C dimension, where t is greater than x, for example, t is equal to 2×x or 3×x, etc., and b_addr is calculated according to the following formula:
[0155] b_addr = b_addr_base + (t / x - 1)×x
[0156] t, m, and i are positive integers.
[0157] For example, in some embodiments, based on the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed, determining the first request sent among multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor and the request initial coordinates corresponding to the first request; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and determining the data length loaded by the first request according to the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed.
[0158] For the first request sent, its corresponding request initial coordinates are the starting coordinates of the tensor to be processed, and thus the data read address of the first request can be calculated with reference to the above formula. The starting address for writing the tensor to be processed in the buffer is used as the data write address of the first request.
[0159] The data length loaded by the first request is determined according to the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is less than the second coordinate value (i.e., the coordinate value of the starting coordinates of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the shape size copy_t of the tensor to be processed in the first dimension. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value (i.e., the coordinate value of the starting coordinates of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the absolute value of the first coordinate value t_coord_b.
[0160] The state machine has a clear logical structure, the advantages of being easy to maintain and expand, and is especially suitable for processing logical scenarios with multiple conditions and multiple branches. In the present disclosure, the state machine automatically updates the request initial coordinates and the loaded data length corresponding to each request, avoiding complex conditional nesting, with clear logic, easy to maintain, and strong scalability.
[0161] Figure 5 It is a schematic diagram of the state machine provided by an embodiment of the present disclosure.
[0162] For example, the state machine includes a first state s0, a second state s1, and a third state s2. The initial state of entering the state machine is determined by the starting coordinate coord_b of the tensor to be processed.
[0163] For example, in response to the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, the initial state is determined to be the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, the initial state is determined to be the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, the initial state is determined to be the third state.
[0164] Reference Figure 5 , the tensor_t direction represents the data in the first dimension. The data with the coordinate of the t dimension less than the second coordinate value and greater than the third coordinate value is not in memory ( Figure 5 the black part in Figure 5 ), and the data with the coordinate of the t dimension between the second coordinate value and the third coordinate value (
[0165] the white part in Figure 5 ) is stored in memory, which is the actual size of the original tensor data in the first dimension.
[0166] Figure 5 The 6 cases (① to ⑥) in
[0167] For example, Figure 5 in case ① in
[0168] For example, Figure 5 in case ② in [reference], the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value and less than the third coordinate value. The state machine loops and jumps between the first state s0 and the second state s1.
[0169] For example, Figure 5 in case ③ in [reference], the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value. The state machine always loops and jumps in the second state s1.
[0170] For example, Figure 5 in case ④ in [reference], the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1, and the third state s2.
[0171] For example, Figure 5 in case ⑤ in [reference], the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps between the second state s1 and the third state s2.
[0172] For example, Figure 5 in case ⑥ in [reference], the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops and jumps in the third state s2.
[0173] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and determine the state entered by the second sent request in combination with the trigger condition. Then, based on the state entered by the second sent request, determine the request initial coordinates corresponding to the third sent request and the data length loaded by the second sent request, and determine the state entered by the third sent request in combination with the trigger condition, and so on.
[0174] For example, in some embodiments, based on the initial state, in combination with the shape and size of the original tensor and the data storage format of the tensor to be processed, using the state machine to determine each request except the first request among multiple requests may include: based on the initial state, using the state machine to determine the request initial coordinates corresponding to each request and the data length loaded by each request; according to the shape and size of the original tensor, the request initial coordinates corresponding to each request, and the data storage format of the tensor to be processed, determine the data read address of each request; according to the data length loaded by each request, determine the data write address of each request.
[0175] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request for each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state.
[0176] After determining the request initial coordinates corresponding to the current request, it is possible to determine whether the sub-data loaded by the current request is located in the memory according to the request initial coordinates.
[0177] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the sub-data to be loaded by the current request is located in the memory. At this time, the data read address and data write address of the current request can be determined with reference to the above calculation formula, and then the current request is sent to the memory to load the corresponding sub-data into the buffer.
[0178] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the sub-data to be loaded by the current request is not located in the memory. At this time, in response to all the sub-data to be loaded by the current request not being located in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values to the buffer, where the number of the multiple predetermined values is determined by the data length specified by the current request. For example, the predetermined value is 0.
[0179] The process of using the state machine to determine the request initial coordinates and the data length loaded by each request will be specifically described below.
[0180] When the first coordinate value of the tensor to be processed in the first dimension is less than the second coordinate value, it enters the first state s0.
[0181] For example, when the current request is the first request, the state of the current request is the initial state, and the state machine outputs the request initial coordinates and the data length corresponding to the current request, and prepares the request initial coordinates for the next state.
[0182] For example, in response to the state of the current request being the first state and satisfying the first condition, it is determined that the next request of the current request enters the first state. For example, the first condition includes that the sum of the size copy_t of the tensor to be processed in the first dimension and the first coordinate value is less than the second coordinate value. Taking the second coordinate value as 0 as an example, the first condition includes that the size copy_t of the tensor to be processed in the first dimension is less than the absolute value of the first coordinate value.
[0183] For example, if the current request is in the first state s0 and the size copy_t of the tensor to be processed in the first dimension is relatively small, for example, the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is less than the second coordinate value, the state of the next request is still the first state s0. For example Figure 5In case ① in
[0184] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value t_coord_b, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those in other dimensions. Specifically, for other dimensions, the update methods include incrementing (by 1 or x), returning to the initial coordinate, and remaining unchanged. For example, for the second dimension, the coordinates of each request in the second dimension are incremented, and the third dimension and higher dimensions remain unchanged until reaching the boundary of the tensor to be processed in the second dimension. After that, the coordinates of the second dimension are updated back to the starting coordinate and the coordinates of the third dimension are incremented. The third dimension is adjacent to and higher than the second dimension, and each dimension increments in this way. The coordinate values of other dimensions lower than the first dimension remain unchanged.
[0185] For example, in this case, the data length req_size loaded by the current request is the size copy_t of the tensor to be processed in the first dimension.
[0186] For example, in response to the current request being in the first state and satisfying the second condition, it is determined that the next request of the current request enters the second state s1. For example, the second condition includes that the size copy_t of the tensor to be processed in the first dimension is greater than or equal to the absolute value of the first coordinate value t_coord_b. For example, such as Figure 5 In case ② and case ④ in
[0187] For example, if the current request is in the first state s0 and the size copy_t of the tensor to be processed in the first dimension is relatively large, for example, the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the second coordinate value, the state of the next request jumps to the second state s1. For example, assuming the second coordinate value is 0, the second condition includes that the size copy_t of the tensor to be processed in the first dimension is greater than or equal to the absolute value of the first coordinate value t_coord_b.
[0188] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the second coordinate value, for example, 0, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged. Since it is a jump from the first state to the second state and the data in the first dimension has not been fully loaded, it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged.
[0189] For example, in this case, the data length req_size loaded by the current request is determined based on the first coordinate value and the second coordinate value, for example, it is the difference between the first coordinate value and the second coordinate value.
[0190] For example, in response to the status of the current request being the second state s1 and satisfying the third condition, it is determined that the next request of the current request enters the first state s0. The third condition includes that the first coordinate value t_coord_b is less than the second coordinate value and the sum of the first coordinate value t_coord_b and the size copy_t of the tensor to be processed in the first dimension is less than the third coordinate value (for example, the second coordinate value is 0 and the third coordinate value is the size tensor_t of the original tensor in the first dimension). For example, like the case of s1->s0 in case ② in Figure 5 Since the sum of the first coordinate value t_coord_b and the size copy_t of the tensor to be processed in the first dimension is less than the third coordinate value, it will not enter the third state s2.
[0191] In this case, it is determined that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value t_coord_b, and it is determined that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those in other dimensions. Specifically, for other dimensions, the update methods include incrementing (by 1 or x), returning to the initial coordinate, and remaining unchanged. For example, for the second dimension, the coordinates of each request in the second dimension are incremented, and the third dimension and higher dimensions remain unchanged until reaching the boundary of the tensor to be processed in the second dimension. After that, the coordinates in the second dimension are updated back to the starting coordinate and the coordinates in the third dimension are incremented. The third dimension is adjacent to and higher than the second dimension, and each dimension is incremented in this way. For other dimensions lower than the first dimension, the coordinates remain unchanged.
[0192] For example, in this case, the data length req_size loaded by the current request is determined based on the first coordinate value and the size copy_t of the tensor to be processed in the first dimension. For example, if the second coordinate value is 0, the data length req_size loaded by the current request is the sum of the first coordinate value and the size copy_t of the tensor to be processed in the first dimension.
[0193] For example, when the status of the current request is the second state s1 and satisfies the fourth condition, it is determined that the next request enters the second state s1. The fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension. For example, when the current request is in the second state, it means that the sub-data of the request are all in the memory, that is, they are all valid data. When the size copy_t of the tensor to be processed in the first dimension is relatively small, for example, the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is less than the third coordinate value, the status of the next request is still the second state s1 and will not jump to the third state s2. For example, like the case of Figure 5 case ③ in
[0194] In this case, it is determined that the coordinate value of the initial coordinate of the next request in the first dimension is the first coordinate value t_coord_b, and it is determined that the coordinate values of the initial coordinate of the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those in other dimensions. Specifically, for other dimensions, the update methods include incrementing, returning the initial coordinate, and remaining unchanged. For example, for the second dimension, the coordinates of each request in the second dimension are incremented, and the third dimension and higher dimensions remain unchanged until the boundary of the tensor to be processed in the second dimension is reached. After that, the coordinates of the second dimension are updated back to the starting coordinate and the coordinates of the third dimension are incremented. The third dimension is adjacent to and higher than the second dimension, and each dimension is incremented in this way. For other dimensions lower than the first dimension, the coordinates remain unchanged.
[0195] For example, in this case, the length of the data loaded by the current request req_size is the size copy_t of the tensor to be processed in the first dimension.
[0196] For example, in response to the state of the current request being the second state s1 and satisfying the fifth condition, it is determined that the next request enters the third state s2. The fifth condition includes that the sum of the size of the tensor to be processed in the first dimension and the first coordinate value is greater than or equal to the third coordinate value. At this time, the size copy_t of the tensor to be processed in the first dimension is relatively large, and the state of the next request jumps to the third state s2. For example, such as Figure 5 s1->s2 in case ④ and case ⑤ in
[0197] In this case, it is determined that the coordinate value of the initial coordinate of the next request in the first dimension is the third coordinate value, and it is determined that the coordinate values of the initial coordinate of the next request in other dimensions except the first dimension remain unchanged.
[0198] For example, in this case, the length of the data loaded by the current request req_size is determined based on the coordinate value t_coord and the third coordinate value of the initial coordinate of the current request in the first dimension. For example, the length of the data loaded by the current request req_size is the difference between the third coordinate value and the coordinate value t_coord.
[0199] For example, in response to the state of the current request being the third state, the state that the next request of the current request enters is determined according to the first coordinate value.
[0200] For example, when the first coordinate value is less than the second coordinate value, it is determined that the next request enters the first state s0. For example, Figure 5 s2->s0 in case ④ in Figure 5In case ⑤, s2->s1. When the first coordinate value is greater than or equal to the third coordinate value, it is determined that the next request enters the third state s2, for example Figure 5 In case ⑥, s2->s2.
[0201] In this case, the coordinate value of the initial coordinate of the request corresponding to the next request in the first dimension is determined to be the first coordinate value t_coord_b, and the coordinate value of the initial coordinate of the request corresponding to the next request in other dimensions except the first dimension is updated according to the coordinate value lower than the other dimensions. Specifically, for other dimensions, the update method includes self-increment (1 or x), return to the initial coordinate or remain unchanged. For example, for the second dimension, the coordinate of the second dimension of each request is self-incremented until it reaches the boundary of the second dimension of the tensor to be processed, and then it is updated back to the starting coordinate and the coordinate of the third dimension is self-incremented. The third dimension is adjacent to the second dimension and higher than the second dimension, and each dimension is incremented in this way. The coordinates of other dimensions lower than the first dimension remain unchanged.
[0202] For example, in this case, the data length req_size loaded by the current request is determined based on the coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the coordinate value of the third coordinate. For example, the data length req_size loaded by the current request is the difference between the coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the coordinate value of the third coordinate.
[0203] In some other embodiments, the state machine may also update the remaining size remaining_copy_t of the five dimensions, where the remaining size indicates the amount of data that has not been loaded in each dimension. When determining the state transition condition of the state machine, the determination is made based on the relationship between the fourth coordinate value of the initial coordinate of the request corresponding to the current request in the first dimension and the remaining size.
[0204] For example, taking the jump from the first state s0 to the first state s0 as an example, the first condition may include that the sum of the fourth coordinate value t_coord of the requested initial coordinate in the first dimension corresponding to the current request and the remaining size remain_copy_t of the tensor to be processed in the first dimension is less than the second coordinate value, and the fourth coordinate value t_coord is less than the second coordinate value. The remaining size remain_copy_t is the number of remaining tensor data in the first dimension of the tensor to be processed that has not been requested to be loaded.
[0205] For example, taking the jump from the first state s0 to the second state s1 as an example, the second condition may include that the sum of the fourth coordinate value t_coord of the request initial coordinate corresponding to the current request in the first dimension and the remaining size remaining_copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value.
[0206] For example, taking the jump from the second state s1 to the first state s0 as an example, the third condition may include that the sum of the fourth coordinate value t_coord of the request initial coordinate corresponding to the current request in the first dimension and the remaining size remain_copy_t of the tensor to be processed in the first dimension is less than or equal to the third coordinate value, and the first coordinate value is less than the second coordinate value.
[0207] For example, taking the jump from the second state s1 to the second state s1 as an example, the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value.
[0208] For example, taking the jump from the second state s1 to the third state s2 as an example, the fifth condition may include that the sum of the request initial coordinate corresponding to the current request and the remaining size remain_copy_t of the tensor to be processed in the first dimension is greater than the third coordinate value.
[0209] For example, for the third state s2, the state that the next request of the current request enters is determined according to the first coordinate value, and specific reference can be made to the foregoing content.
[0210] Of course, as the coordinates continuously increase, when the remaining sizes of all dimensions are equal to 1, it is determined that the state machine can end the coordinate determination of each request, and the state machine can enter the Idle state.
[0211] The state transition conditions of the specific state machine and the updated coordinate content can be changed and set according to actual needs. The logic is similar to the state machine logic described above, and no further examples are given here.
[0212] The data loading process is described in detail below when the first dimension is the W dimension. In this example, the first coordinate value is the coordinate value w_coord_b of the starting coordinate of the tensor to be processed in the W dimension, the second coordinate value is 0, the third coordinate value is the size tensor_w of the original tensor in the W dimension, and the predetermined value is 0.
[0213] For example, when the size copy_w of the tensor to be processed in the W dimension is less than or greater than the size tensor_w of the original tensor in the W dimension, and / or the first coordinate value w_coord_b is not equal to 0, it is determined at this time that the tensor to be processed cannot be continuously loaded in the W dimension.
[0214] For example, first in step S10, the shape and size of the tensor to be processed are obtained, that is, copy_c / h / w / d / n, etc., and the starting coordinate of the tensor to be processed in the coordinate system determined by the original tensor, that is, c / h / w / d / n_coord_b.
[0215] After that, in step S20, multiple requests for loading the tensor to be processed are determined in combination with the above information.
[0216] Referring to the process described above, a state machine is used to determine the request initial coordinates and the data length loaded for each request, and based on this, the data read address and data write address for each request are determined. For example, in this example, the coordinates of the sub-data loaded by each request are different in the W dimension, but the same in the H dimension, D dimension, and N dimension. Also, the coordinates of the sub-data loaded by each request are different in the C dimension, which is caused by xC in the interleaving pattern.
[0217] Table 1 shows the update process of the request initial coordinates and the data length loaded for each request provided in an embodiment of the present disclosure.
[0218] Table 1
[0219]
[0220] As shown in Table 1, if the first coordinate value w_coord_b is less than 0, it is determined that the first request enters the first state s0. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0221] If the first condition is satisfied, it is determined that the next request enters the first state s0. For the first condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value w_coord of the request initial coordinates corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the request initial coordinates corresponding to the next request in the H dimension is incremented by 1 and returns h_coord_b when reaching the boundary of the tensor to be processed in the H dimension; the coordinate value d_coord of the request initial coordinates corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the D dimension, and returns d_coord_b when reaching the boundary of the tensor to be processed in the D dimension; the coordinate value n_coord of the request initial coordinates corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the N dimension, and returns n_coord_b when reaching the boundary of the tensor to be processed in the N dimension; the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension is incremented by x when n_coord returns n_coord_b, and remains unchanged in other cases.
[0222] Also, at this time, the data length loaded by the first request is the size copy_w of the tensor to be processed in the W dimension.
[0223] If the second condition is satisfied, it is determined that the next request enters the second state s1. For the second condition, refer to the foregoing description. At this time, the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is 0; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension, the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension, the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension, and the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension all remain unchanged.
[0224] Moreover, at this time, the data length loaded by the first request is the absolute value of the first coordinate value.
[0225] Since the coordinate value w_coord of the request initial coordinate corresponding to the first request in the W dimension is less than 0, all the sub-data loaded by the first request are not in the memory. In step S30, the first request is converted into an operation of writing copy_w or |w_coord_b| zeros to the buffer area, where |w_coord_b| represents the absolute value of w_coord_b.
[0226] As shown in Table 1, when the first coordinate value w_coord_b is greater than or equal to 0 and less than tensor_w, it is determined that the first request enters the second state s1. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0227] If the third condition is satisfied, it is determined that the next request enters the first state s0. For the third condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value w_coord of the request initial coordinate corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1 and returns h_coord_b when reaching the boundary of the tensor to be processed in the W dimension; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If it does not reach the boundary of the tensor to be processed in the D dimension, it remains unchanged and returns d_coord_b when reaching the boundary of the tensor to be processed in the D dimension; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b. If it does not reach the boundary of the tensor to be processed in the N dimension, it remains unchanged and returns n_coord_b when reaching the boundary of the tensor to be processed in the N dimension; the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is incremented by x when n_coord returns n_coord_b, and remains unchanged in other cases.
[0228] And at this time, the length of the data loaded by the first request is the sum of the size copy_w of the tensor to be processed in the W dimension and the first coordinate value w_coord_b.
[0229] If the fourth condition is satisfied, it is determined that the next request enters the second state s1. For the fourth condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value w_coord of the initial coordinate of the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the initial coordinate of the next request in the H dimension is incremented by 1, and when reaching the boundary of the tensor to be processed in the W dimension, it returns h_coord_b; the coordinate value d_coord of the initial coordinate of the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If it does not reach the boundary of the tensor to be processed in the D dimension, it remains unchanged, and when reaching the boundary of the tensor to be processed in the D dimension, it returns d_coord_b; the coordinate value n_coord of the initial coordinate of the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b. If it does not reach the boundary of the tensor to be processed in the N dimension, it remains unchanged, and when reaching the boundary of the tensor to be processed in the N dimension, it returns n_coord_b; the coordinate value c_coord of the initial coordinate of the next request in the C dimension is incremented by x when n_coord returns n_coord_b, and remains unchanged in other cases.
[0230] And at this time, the length of the data loaded by the first request is the size copy_w of the tensor to be processed in the W dimension.
[0231] If the fifth condition is satisfied, it is determined that the next request enters the third state s2. For the fifth condition, refer to the foregoing description. As shown in Table 1, at this time, the coordinate value w_coord of the initial coordinate of the next request in the W dimension is the third coordinate value tensor_w; the coordinate values h_coord of the initial coordinate of the next request in the H dimension, d_coord of the initial coordinate of the next request in the D dimension, c_coord of the initial coordinate of the next request in the C dimension, and n_coord of the initial coordinate of the next request in the N dimension all remain unchanged.
[0232] And at this time, the length of the data loaded by the first request is the difference between the size tensor_w of the original tensor in the W dimension and the coordinate value w_coord of the initial coordinate of the next request in the W dimension.
[0233] Since the state of the first request is the second state s1, all the sub-data it loads are located in the memory. In step S30, the first request is sent to the memory to load the sub-data in the original tensor stored in the memory into the memory.
[0234] As shown in Table 1, when the first coordinate value w_coord_b is greater than or equal to tensor_w, it is determined that the first request enters the third state s2. The data read address and data write address of the first request refer to the foregoing description and will not be elaborated here.
[0235] Determine the state that the next request enters according to the first coordinate value w_coord_b. For example, when the first coordinate value w_coord_b is less than 0, it is determined that the next request enters the first state s0; when the first coordinate value w_coord_b is greater than or equal to 0 but less than tensor_w, it is determined that the next request enters the second state s1; when the first coordinate value w_coord_b is greater than or equal to tensor_w, it is determined that the next request enters the third state s2.
[0236] As shown in Table 1, regardless of which state the next request enters, the update logic of the request initial coordinates and the data length loaded by the first request are exactly the same.
[0237] For example, taking the next request entering the first state s0 as an example, at this time, the coordinate value w_coord of the request initial coordinates corresponding to the next request in the W dimension is the first coordinate value w_coord_b; the coordinate value h_coord of the request initial coordinates corresponding to the next request in the H dimension is incremented by 1 and returns h_coord_b when reaching the boundary of the tensor to be processed in the W dimension; the coordinate value d_coord of the request initial coordinates corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the D dimension, and returns d_coord_b when reaching the boundary of the tensor to be processed in the D dimension; the coordinate value n_coord of the request initial coordinates corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b, remains unchanged if it does not reach the boundary of the tensor to be processed in the N dimension, and returns n_coord_b when reaching the boundary of the tensor to be processed in the N dimension; the coordinate value c_coord of the request initial coordinates corresponding to the next request in the C dimension is incremented by x when n_coord returns n_coord_b, and remains unchanged in other cases. The data length loaded by the first request is the difference between the size tensor_w of the original tensor in the W dimension and the coordinate value w_coord of the request initial coordinates corresponding to the next request in the W dimension.
[0238] Since the coordinate value w_coord of the request initial coordinates corresponding to the first request in the W dimension is greater than tensor_w, all the sub-data loaded by the first request are not in the memory. In step S30, the first request is converted into an operation of writing w_coord - tensor_w zeros to the buffer area.
[0239] After that, continue the above process to determine the initial coordinates of each subsequent request and the length of the data to be loaded, and execute step S30 to send multiple requests. The specific process will not be elaborated here.
[0240] In the above embodiment, requests are divided according to whether the tensors can be continuously loaded in the first dimension. Each request is used to load sub-data that is either all located in memory or all not located in memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, and thus enhancing the hardware utilization rate of the computing unit and improving the hardware performance.
[0241] At least one embodiment of the present disclosure further provides a data storage method for writing a tensor to be processed in a buffer into memory.
[0242] For example, the shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, depth dimension, height dimension, width dimension, and number of channels dimension. For the relevant content of the tensor to be processed, reference can be made to the description of the foregoing data loading method, which will not be elaborated here.
[0243] Figure 6 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure.
[0244] As Figure 6 shown, the data storage method provided by at least one embodiment of the present disclosure includes at least steps S40 - S60.
[0245] This data storage method is used to write the tensor to be processed in the buffer into memory in the NDHWC data storage format, where the data storage format of the tensor to be processed in the buffer is N(C / x)DHW(xC), and x is a positive integer greater than 1.
[0246] As Figure 6 shown, first, in step S40, obtain the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor.
[0247] The data storage format of the original tensor is NDHWC.
[0248] The shape and size of the original tensor are represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component.
[0249] For the related descriptions about the original tensor and the relationship between the original tensor and the tensor to be processed, reference can be made to the relevant content of the foregoing data loading method, which will not be elaborated here.
[0250] In step S50, based on the shape and size of the original tensor and the tensor to be processed, and the starting coordinates of the tensor to be processed, multiple requests for storing the tensor to be processed are determined.
[0251] In step S60, at least some of the multiple requests are sequentially sent, and the data to be written indicated by each request is sequentially written into the memory to store the tensor to be processed in the memory.
[0252] For example, in some embodiments, step S50 may include: determining the first request to be sent among the multiple requests and the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed and the shape and size of the tensor to be processed; based on the initial state, combining the shape and size of the original tensor and the data storage format of the tensor to be processed, and using the state machine to determine each request among the multiple requests except the first request.
[0253] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when storing the tensor to be processed, it cannot be continuously stored in the first dimension but can be continuously stored in each dimension lower than the first dimension, requests are divided in the second dimension, and the sub-data to be written by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory.
[0254] For example, referring to Figure 5 , the tensor_t direction represents the data in the first dimension, and the data with the coordinate of the t dimension less than the second coordinate value and greater than the third coordinate value is not in the memory or the data range of the t dimension is not within the data range of the original tensor ( Figure 5 the part shown in the black background in Figure 5 this part of the data can be called invalid data), and the data with the coordinate of the t dimension between the second coordinate value and the third coordinate value (
[0255] the white part in Figure 5 is stored in the memory, and the data range of the t dimension is within the data range of the original tensor.
[0256] For example, if the sub-data for writing does not fall within the data range of the original tensor, it represents the coordinate range of the sub-data in any dimension. For example, referring to Figure 5 the t-dimension, which is less than the second coordinate value or greater than the third coordinate value, that is, it does not belong to the original tensor. In this case, the sub-data for writing does not belong to the original tensor and may not need to be stored in memory subsequently.
[0257] The determination method for each request can refer to the relevant content of step S20 in the foregoing data loading method, and the repeated parts will not be elaborated here.
[0258] As Figure 4A shown in the example, the tensor to be processed consists of partial tensors of the original tensor. At this time, all tensors to be processed in the buffer need to be stored at the corresponding positions in the original tensor to update the original tensor. Therefore, at this time, multiple requests need to be sent to the memory to store the sub-data for each request at the corresponding positions in the memory.
[0259] As Figure 4B shown in the example, the tensor to be processed not only includes partial tensors of the original tensor, but also includes data outside the edge of the original tensor caused by filling operations, etc., which are not stored in the memory. Therefore, requests need to be distinguished before being sent. If the sub-data for which the request is used to store does not belong to the content in the original tensor, for example Figure 4B the small dashed cube in Figure 4B then there is no need to send a request, and these sub-data do not need to be stored in the memory; if the sub-data for which the request is used to store belongs to the original tensor, for example
[0260] the solid small cube in the tensor to be processed that belongs to the original tensor in Figure 4C then send a request to the memory to store the sub-data at the corresponding position in the original tensor.
[0261] For example, in some embodiments, step S60 may include: for any request, in response to the first request coordinate value of the request initial coordinate corresponding to any request in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determine not to send any request; in response to the first request coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determine to send any request; where the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
[0262] For example, if the first request coordinate value of the request corresponding to the initial coordinate in the first dimension is less than the second coordinate value, or the first request coordinate value is greater than or equal to the third coordinate value, it means at this time that the sub-data requested for storage does not belong to the original tensor and does not need to be stored in the memory, and the request is not sent.
[0263] For example, if the first request coordinate value is greater than or equal to the second coordinate value and less than the third coordinate value, it means at this time that the sub-data requested for storage belongs to the original tensor and needs to be stored in the memory, and the request is sent to store the sub-data at the corresponding position in the original tensor.
[0264] The data storage method provided by at least one embodiment of the present disclosure divides requests according to whether the tensor can be continuously stored in the first dimension. Each sub-data requested for storage is either entirely located in the memory or entirely not located in the memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improves the bandwidth and efficiency during data access, and further improves the hardware utilization rate of the computing unit and the hardware performance.
[0265] At least one embodiment of the present disclosure also provides a data loading method. Figure 7 It is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure.
[0266] As Figure 7 shown, the data loading method provided by at least one embodiment of the present disclosure includes steps S201 - S202.
[0267] For example, in step S201, a data loading instruction for instructing to load the tensor to be processed into the buffer is received.
[0268] For example, the data loading instruction includes the shape size of the tensor to be processed and the starting coordinate of the tensor to be processed in the coordinate system determined by the original tensor as input parameters.
[0269] The tensor to be processed is stored in the buffer in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1.
[0270] For the related descriptions of the tensor to be processed and the original tensor, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0271] For example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction can be a Load instruction.
[0272] In step S202, after parsing the data loading instruction, the execution unit executes the data loading instruction.
[0273] For example, after the processor receives the data loading instruction, it parses the data loading instruction, such as decoding the data loading instruction, generating microinstructions, and sending the microinstructions to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the microinstruction category; in response to the microinstruction, when the input parameters are ready, the execution unit executes the relevant operations of the data loading instruction.
[0274] For example, step S202 may include: determining a plurality of requests for loading the tensor to be processed in combination with the shape and size of the original tensor and the starting coordinates of the tensor to be processed; sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area.
[0275] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in each dimension lower than the first dimension, the requests are divided in the second dimension, and the sub-data loaded by each request are all located in the memory or none of them are located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory, and the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.
[0276] Regarding "determining a plurality of requests for loading the tensor to be processed in combination with the shape and size of the original tensor and the starting coordinates of the tensor to be processed", reference may be made to the relevant content of the foregoing step S20, which will not be elaborated here.
[0277] Regarding "sequentially sending the plurality of requests, and sequentially writing the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area", reference may be made to the relevant content of the foregoing step S30, which will not be elaborated here.
[0278] In the present disclosure, the data loading instruction is divided into a plurality of requests. According to whether the tensor can be continuously loaded in the first dimension, the data loading instruction is converted into a plurality of requests, and the sub-data for loading by each request are all located in the memory or none of them are located in the memory. The division of the data requests is more reasonable, more suitable for the storage of tensor data, greatly improves the bandwidth and efficiency during data access, and further improves the hardware utilization rate of the computing unit and improves the hardware performance.
[0279] At least one embodiment of the present disclosure further provides a data storage method. Figure 8 It is a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure.
[0280] AsFigure 8 As shown, the data storage method provided by at least one embodiment of the present disclosure includes steps S203 - S204.
[0281] For example, in step S203, a data storage instruction indicating to execute writing a tensor to be processed in the buffer into the memory is received.
[0282] For example, the data storage instruction includes the shape dimensions of the tensor to be processed as an input parameter, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. The data storage instruction is used to write the tensor to be processed in the N(C / x)DHW(xC) format in the buffer into the memory in the NDHWC data storage format. The data storage format of the original tensor is NDHWC.
[0283] For the relevant descriptions of the tensor to be processed and the original tensor, reference may be made to the relevant descriptions of the foregoing data storage method, and repeated parts will not be elaborated.
[0284] For example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction can be a Store instruction.
[0285] In step S204, after parsing the data storage instruction, the execution unit is used to execute the data storage instruction.
[0286] For example, after the processor receives the data storage instruction, it parses the data storage instruction, for example, decodes the data storage instruction, generates a micro-instruction and sends the micro-instruction to the instruction distribution unit; the instruction distribution unit sends it to the corresponding scheduling queue according to the micro-instruction category; in response to the micro-instruction, when the input parameters are ready, the execution unit executes the relevant operations of the data storage instruction.
[0287] For example, step S204 may include: combining the original tensor, the shape dimensions of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine a plurality of requests for storing the tensor to be processed; sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in order to store the tensor to be processed into the memory.
[0288] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when storing the tensor to be processed, it cannot be continuously stored in the first dimension but can be continuously stored in each dimension lower than the first dimension, the requests are divided in the second dimension, and the sub-data to be written by each request all belong to the data range of the original tensor or all do not belong to the data range of the original tensor. And in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory.
[0289] Regarding "determining multiple requests for storing the tensor to be processed by combining the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed", reference may be made to the relevant content of the foregoing step S50, which will not be elaborated here.
[0290] Regarding "sequentially sending at least some of the multiple requests and writing the data to be written indicated by each request into the memory in sequence to store the tensor to be processed", reference may be made to the relevant content of the foregoing step S60, which will not be elaborated here.
[0291] In the present disclosure, the data storage instruction is divided into multiple requests. According to whether the tensor can be continuously stored in the first dimension, the data storage instruction is converted into multiple requests, and the sub-data for storage by each request is all located in the memory or all not located in the memory. The division of the data requests is more reasonable and more suitable for storing tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.
[0292] Figure 9 This is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 9 shown, the electronic device 300 is, for example, suitable for implementing the data loading method or the data storage method provided by the embodiments of the present disclosure. It should be noted that Figure 9 the components of the electronic device 300 shown are merely exemplary and not restrictive. According to actual application requirements, the electronic device 300 may further include other components.
[0293] As Figure 9 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.
[0294] For example, when the computer-readable instructions are run by the processing device 301, one or more steps of the data loading method described in any of the foregoing embodiments, or one or more steps of the data storage method described in any of the foregoing embodiments may be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference may be made to the relevant descriptions in the foregoing embodiments of the data loading method, and for a detailed description of the processing process of the data storage method, reference may be made to the relevant descriptions in the foregoing embodiments of the data storage method.
[0295] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 303 and / or cache memory, etc. For example, computer-readable instructions may be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage media, such as style images, and various data used and / or generated by the application programs, etc.
[0296] For example, the processing device 301, read-only memory (ROM) 302, and random access memory (RAM) 303 are connected to each other via the bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0297] Generally, the following devices may be connected to the input / output (I / O) interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 may alternatively implement or have more or fewer devices. For example, the processing device 301 may control other components in the electronic device 300 to perform desired functions. The processing device 301 may be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit GPU, etc., which has data processing capabilities and / or program execution capabilities. The central processing unit (CPU) may be of the X86, ARM, RISC-V architecture, etc. The GPU may be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.
[0298] Figure 10 Schematic structural diagram of a processor provided for at least one embodiment of the present disclosure. As Figure 10 shown, the processor 400 includes an instruction parsing unit 401 and a first execution unit 402.
[0299] For example, the instruction parsing unit 401 is configured to receive and parse data loading instructions.
[0300] For example, the data loading instruction includes the shape and size of the to-be-processed tensor as an input parameter, the starting coordinates of the to-be-processed tensor in the coordinate system determined by the original tensor. The to-be-processed tensor is stored in the buffer in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1.
[0301] For the related descriptions of the to-be-processed tensor and the original tensor, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated here.
[0302] For example, after the instruction parsing unit parses the data loading instruction, the first execution unit 402 executes the data loading instruction.
[0303] For example, when the first execution unit 402 executes the data loading instruction, it includes performing the following operations: determining a plurality of requests for loading the to-be-processed tensor in combination with the original tensor, the shape and size of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor; sequentially sending the plurality of requests, and writing the sub-data returned by each request into the buffer in order to load the to-be-processed tensor into the buffer.
[0304] For example, in response to the size relationship between the to-be-processed tensor and the original tensor in the first dimension such that when loading the to-be-processed tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, requests are divided in the second dimension. The sub-data loaded by each request is either all located in the memory or none of it is located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.
[0305] Specifically, when upper-layer software based on a processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send data loading instructions for computing and processing to the processor (such as a CPU or GPU) through a uniformly encapsulated function library, the data loading instructions can carry the shape and size of the tensor to be processed as input parameters, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; when the processor receives the data loading instruction, the instruction parsing unit 401 parses the data loading instruction to obtain the shape and size of the tensor to be processed as input parameters, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, and the processor schedules the operation unit to execute the data loading task for the input parameters. For example, after parsing the data loading instruction, the processor can store the input parameters in the data loading instruction in a register or memory, and the first execution unit 402 can obtain the input parameters from the register or memory when performing computing and processing.
[0306] Regarding the specific process of using the first execution unit 402 to execute the data loading instruction, reference can be made to steps S20 - S30 in the data loading method described above, and the repeated parts will not be elaborated.
[0307] The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method, and the repeated parts will not be elaborated.
[0308] Figure 11 It is a schematic structural diagram of the processor provided by at least one embodiment of the present disclosure. As Figure 11 shown, the processor 500 includes an instruction parsing unit 501 and a second execution unit 502.
[0309] For example, the instruction parsing unit 501 is used to receive and parse data storage instructions.
[0310] For example, the data storage instruction includes the shape and size of the tensor to be processed as input parameters, the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, and is used to store the tensor to be processed in the N(C / x)DHW(xC) format in the memory in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, and the data storage format of the original tensor is NDHWC.
[0311] Regarding the related descriptions of the tensor to be processed and the original tensor, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.
[0312] For example, after the instruction parsing unit parses the data storage instruction, the second execution unit 502 executes the data storage instruction.
[0313] For example, when the second execution unit 502 executes a data storage instruction, it includes performing the following operations: determining a plurality of requests for storing the tensor to be processed in combination with the shape and size of the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed; sequentially sending at least some of the plurality of requests, and writing the data to be written indicated by each request into the memory in sequence to store the tensor to be processed.
[0314] For example, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in each dimension lower than the first dimension, the requests are divided in the second dimension, and the sub-data loaded by each request is either all located in the memory or none of it is located in the memory, and in response to the sub-data loaded by the request being all located in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.
[0315] Specifically, when the upper-layer software based on the processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send a data storage instruction for computing and processing to the processor (such as a CPU or a GPU) through a unified encapsulated function library, the data storage instruction can carry the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters; when the processor receives the data storage instruction, the instruction parsing unit 401 parses the data storage instruction to obtain the shape and size of the tensor to be processed, the data storage format of the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor as input parameters, and the processor schedules the operation unit to execute the data loading task for the input parameters. For example, after parsing the data storage instruction, the processor can store the input parameters in the data storage instruction in a register or memory, and when the second execution unit 502 performs computing and processing, it can obtain the input parameters from the register or memory.
[0316] Regarding the specific process of using the second execution unit 502 to execute the data storage instruction, reference can be made to steps S50 - S60 in the data loading method described above, and the repeated parts will not be elaborated.
[0317] The processor provided in at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data storage method, and the repeated parts will not be elaborated.
[0318] Figure 12 It is a schematic diagram of a non-transitory computer-readable storage medium provided in at least one embodiment of the present disclosure. For example, as Figure 12As shown, the storage medium 600 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 601 can be non-temporarily stored on the storage medium 600. For example, when the computer-readable instructions 601 are executed by a processor, one or more steps of the data loading method described above can be performed. For example, when the computer-readable instructions 601 are executed by a processor, one or more steps of the data storage method described above can be performed.
[0319] For example, the storage medium 600 can be applied to the electronic device 300. For example, the storage medium 600 can include the storage device 308 in the electronic device 300.
[0320] For example, the storage device can include any combination of one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions can be stored on the computer-readable storage medium, and the processor can run the computer-readable instructions to implement various functions of the processor. Various application programs and various data can also be stored in the storage medium.
[0321] For example, the storage medium can include the memory card of a smart phone, the cache component of a tablet computer, the hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and can also be other applicable storage media.
[0322] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0323] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.
[0324] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0325] The above description is only the preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.
[0326] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0327] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0328] The following points also need to be noted regarding the present disclosure:
[0329] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.
[0330] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.
[0331] The above description is only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A data loading method for loading a tensor to be processed into a buffer, where, The to-be-processed tensor is stored in the buffer in the NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data loading method includes: Obtaining the shape size of the to-be-processed tensor and the starting coordinates of the to-be-processed tensor in the coordinate system determined by the original tensor, where the data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1; Combining the original tensor, the shape size of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor to determine multiple requests for loading the to-be-processed tensor; Sequentially sending the multiple requests and sequentially writing the sub-data returned by each request into the buffer to load the to-be-processed tensor into the buffer; Wherein, in response to the size relationship between the to-be-processed tensor and the original tensor in the first dimension such that when loading the to-be-processed tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, requests are divided in the second dimension, the sub-data loaded by each request is either all in the memory or none of it is in the memory, and in response to the sub-data loaded by the request being all in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory. Wherein, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.
2. The data loading method according to claim 1, wherein, In response to the first dimension being the W dimension, H dimension, or D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension, or the first dimension being the N dimension and the size of the to-be-processed tensor in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the to-be-processed tensor in the C dimension being equal to x, the second dimension is the N dimension.
3. The data loading method according to claim 1, wherein, Each request includes a data read address for indicating the starting position for reading data from the memory, a data write address for indicating the starting position for writing data into the buffer, and the length of the data loaded by the request. Determining the request initial coordinates corresponding to the next request adjacent to the previous request in the sending order according to the request initial coordinates corresponding to the previous request, and the request initial coordinates corresponding to each request are used to determine the data read address, data write address, and the length of the data loaded by the request. Wherein, when sequentially determining the request initial coordinates corresponding to each request, the coordinate values of the request initial coordinates are sequentially incremented and updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.
4. The data loading method according to claim 1, wherein, In response to the size of the tensor to be processed in the first dimension being not equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension being not equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.
5. The data loading method according to claim 1, wherein, Combining the original tensor, the shape and size of the tensor to be processed, and the starting coordinate of the tensor to be processed, multiple requests for loading the tensor to be processed are determined, including: Based on the starting coordinate of the tensor to be processed and the shape and size of the tensor to be processed, the first request sent in the multiple requests and the initial state when the first request enters the state machine are determined; Combining the initial state, based on the shape and size of the original tensor, each request other than the first request in the multiple requests is determined using the state machine.
6. The data loading method according to claim 5, wherein, Each request includes a data read address for indicating the starting position for reading data from the memory, a data write address for indicating the starting position for writing data to the buffer, and the data length loaded by the request. Based on the starting coordinate of the tensor to be processed, the shape and size of the tensor to be processed, and the data storage format of the tensor to be processed, the first request sent in the multiple requests and the initial state when the first request enters the state machine are determined, including: Based on the starting coordinate of the tensor to be processed, it is determined that the first request enters the initial state in the state machine; The starting coordinate of the tensor to be processed is used as the request initial coordinate corresponding to the first request; According to the shape and size of the original tensor and the request initial coordinate corresponding to the first request, the data read address of the first request is determined; The starting address for writing the tensor to be processed in the buffer is determined as the data write address of the first request; According to the starting coordinate of the tensor to be processed and the shape and size of the tensor to be processed, the data length loaded by the first request is determined.
7. The data loading method according to claim 6, wherein, Based on the starting coordinate of the tensor to be processed, it is determined that the first request enters the initial state in the state machine, including: In response to the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the initial state is the first state; In response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that the initial state is the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; In response to the first coordinate value being greater than or equal to the third coordinate value, it is determined that the initial state is the third state.
8. The data loading method according to claim 5, wherein, Combining the initial state, based on the shape and size of the original tensor, each request other than the first request in the multiple requests is determined using the state machine, including: Based on the initial state, the request initial coordinates corresponding to each request and the data length loaded by each request are determined using the state machine; Determine the data reading addresses of the respective requests according to the shape and size of the original tensor and the request initial coordinates corresponding to the respective requests; Determine the data writing addresses of the respective requests according to the data lengths loaded by the respective requests.
9. The data loading method according to claim 8, wherein, The state machine includes a first state, Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: In response to the state of the current request being the first state and satisfying a first condition, determine that the next request of the current request enters the first state, wherein, in response to the current request being the first request, the state of the current request is the initial state; In response to the next request entering the first state, determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is a first coordinate value, and determine that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; Determine that the data length loaded by the current request is the size of the to-be-processed tensor in the first dimension; Wherein, the first condition includes that the sum of the size of the to-be-processed tensor in the first dimension and the first coordinate value is less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, or, the first condition includes that the sum of the fourth coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the remaining size is less than the second coordinate value and the fourth coordinate value is less than the second coordinate value, and the remaining size is the number of remaining tensor data in the first dimension of the to-be-processed tensor that has not been requested to be loaded, The first coordinate value is the coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension.
10. The data loading method according to claim 9, wherein, The state machine further includes a second state, Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests further includes: In response to the state of the current request being the first state and satisfying a second condition, determine that the next request of the current request enters the second state, In response to the next request entering the second state, determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the second coordinate value, and determine that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged; Determine the data length loaded by the current request based on the first coordinate value and the second coordinate value; Wherein, the second condition includes that the sum of the size of the to-be-processed tensor in the first dimension and the first coordinate value is greater than or equal to the second coordinate value, or, the second condition includes that the sum of the fourth coordinate value and the remaining size is greater than the second coordinate value.
11. The data loading method according to claim 8, wherein, The state machine includes a first state and a second state, Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests includes: In response to the status of the current request being the second status and meeting the third condition, determine that the next request of the current request enters the first status, where, in response to the current request being the first request, the status of the current request is the initial status; In response to the next request entering the first status, determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determine that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; Based on the first coordinate value and the size of the to-be-processed tensor in the first dimension, determine the data length loaded by the current request; Wherein, the third condition includes that the first coordinate value is less than the second coordinate value and the sum of the first coordinate value and the size of the to-be-processed tensor in the first dimension is less than the third coordinate value, or, the third condition includes that the sum of the fourth coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the remaining size is less than the third coordinate value and the fourth coordinate value is less than the second coordinate value, and the remaining size is the number of remaining tensor data in the first dimension of the to-be-processed tensor that has not been requested to be loaded, The first coordinate value is the coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension, the second coordinate value is the coordinate value of the starting coordinate of the original tensor in the first dimension, and the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension.
12. The data loading method according to claim 11, wherein, Based on the initial status, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests, further includes: In response to the status of the current request being the second status and meeting the fourth condition, determine that the next request enters the second status, In response to the next request entering the second status, determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determine that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than those of the other dimensions; Determine that the data length loaded by the current request is the size of the to-be-processed tensor in the first dimension; Wherein, the fourth condition includes that the first coordinate value is greater than or equal to the second coordinate value and the sum of the first coordinate value and the size of the to-be-processed tensor in the first dimension is less than the third coordinate value, or, the fourth condition includes that the sum of the fourth coordinate value and the remaining size is less than the third coordinate value and the first coordinate value is greater than or equal to the second coordinate value.
13. The data loading method according to claim 12, wherein, The state machine further includes a third status, Based on the initial status, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests, further includes: In response to the status of the current request being the second status and meeting the fifth condition, determine that the next request enters the third status, In response to the next request entering the third state, determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the third coordinate value, and determine that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension remain unchanged; Determine the data length loaded by the current request based on the coordinate value of the request initial coordinate corresponding to the current request in the first dimension and the third coordinate value; Wherein, the fifth condition includes that the sum of the size of the to-be-processed tensor in the first dimension and the first coordinate value is greater than or equal to the third coordinate value, or the fifth condition includes that the sum of the fourth coordinate value and the remaining size is greater than the third coordinate value.
14. The data loading method according to claim 8, wherein, The state machine includes a third state, Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the respective requests and the data lengths loaded by the respective requests, including: In response to the state of the current request being the third state, determine the state entered by the next request of the current request according to the first coordinate value, wherein, in response to the current request being the first request, the state of the current request is the initial state; Determine that the coordinate value of the request initial coordinate corresponding to the next request in the first dimension is the first coordinate value, and determine that the coordinate values of the request initial coordinate corresponding to the next request in other dimensions except the first dimension are updated according to the coordinate values lower than the other dimensions; Determine that the data length loaded by the current request is determined based on the fourth coordinate value and the third coordinate value of the request initial coordinate corresponding to the current request in the first dimension, Wherein, the first coordinate value is the coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension, and the third coordinate value is the sum of the second coordinate value of the starting coordinate of the original tensor in the first dimension and the size of the original tensor in the first dimension.
15. The method according to any one of claims 8 - 14, wherein, According to the shape size of the original tensor and the request initial coordinates corresponding to the respective requests, determine the data reading addresses of the respective requests, including: Calculate the data reading address Addr_1 of the nth request according to the following formula: Addr_1 = u_addr_base + ((((n_coord × tensor_d + d_coord) × tensor_h + h_coord) × tensor_w + w_coord) × tensor_c + c_coord) Wherein, u_addr_base represents the storage address in the memory of the element located at the starting coordinate position of the original tensor in the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in 5 dimensions of NDHWC, and n is a positive integer.
16. The method according to any one of claims 8-14, wherein, Determine the data write addresses of the respective requests according to the data lengths loaded by the respective requests, including: For m requests with the same range of coordinate values of the sub-data of the requests in the C dimension, calculate the data write address Addr_2_i of the i-th request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_0×copy_c + req_size_1×copy_c +…+ req_size_i-1×copy_c where b_addr represents the data write address of the first request sent among the m requests, req_size_0, req_size_1,..., req_size_i-1 represent the data lengths loaded by the respective first i-1 requests sent, and copy_c is the size of the to-be-processed tensor in the C dimension; where, in response to the sub-data of each of the m requests including the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base for writing the to-be-processed tensor into the buffer; in response to the sub-data of each of the m requests including the t-th data element to the (t + x)-th data element of the original tensor in the C dimension, where t > x, the b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1)×x t, m, and i are positive integers.
17. The data loading method according to claim 1, wherein, Send the multiple requests in sequence, and write the sub-data returned by each request into the buffer in sequence to load the to-be-processed tensor into the buffer, including: For any one request, in response to the sub-data loaded by the any one request being all located in the memory, send the any one request to the memory; in response to the sub-data loaded by the any one request not being all located in the memory, convert the any one request into writing a plurality of predetermined values into the buffer, where the number of the plurality of predetermined values is determined by the data length specified by the any one request.
18. A data loading method, including: Receive a data loading instruction indicating to execute loading a to-be-processed tensor into a buffer, where the data loading instruction includes the shape size of the to-be-processed tensor, the starting coordinates of the to-be-processed tensor in a coordinate system determined by an original tensor, as input parameters, and the to-be-processed tensor is stored in the buffer in an NDHWC data storage format, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, and the data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1; After parsing the data loading instruction, use an execution unit to execute the data loading instruction; where using the execution unit to execute the data loading instruction includes: Combine the original tensor, the shape size of the to-be-processed tensor, and the starting coordinates of the to-be-processed tensor to determine a plurality of requests for loading the to-be-processed tensor; Send the multiple requests in sequence, and write the sub-data returned by each request into the buffer in order to load the tensor to be processed into the buffer; Among them, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in each dimension lower than the first dimension. Divide the requests in the second dimension. The sub-data loaded by each request is either all in the memory or none of them are in the memory. And in response to the sub-data loaded by the request being all in the memory, the sub-data loaded by the request comes from the original tensor and is continuously stored in the memory. Among them, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.
19. A data storage method for writing a to-be-processed tensor in a buffer to memory in an NDHWC data storage format, where, The data storage format of the tensor to be processed in the buffer is N(C / x)DHW(xC), where x is a positive integer greater than 1, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension. The data storage method includes: Obtain the shape size of the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Among them, the data storage format of the original tensor is NDHWC; Combine the original tensor, the shape size of the tensor to be processed, and the starting coordinates of the tensor to be processed to determine multiple requests for storing the tensor to be processed; Send at least some of the multiple requests in sequence, and write the data to be written indicated by each request into the memory in order to store the tensor to be processed into the memory; Among them, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when storing the tensor to be processed, it cannot be continuously stored in the first dimension but can be continuously stored in each dimension lower than the first dimension. Divide the requests in the second dimension. The sub-data to be written by each request either all belong to the data range of the original tensor or none of them belong to the data range of the original tensor. And in response to the sub-data to be written by the request all belonging to the data range of the original tensor, the sub-data to be written by the request belongs to the original tensor and is continuously stored in the memory. Among them, the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.
20. The data storage method according to claim 19, wherein, The request initial coordinates corresponding to each request are used to determine the data storage address of the sub-data to be stored by the request in the memory. Send at least some of the multiple requests in sequence, and write the data to be written indicated by each request into the memory to store the tensor to be processed, including: For any request, in response to the first request coordinate value of the request initial coordinates corresponding to the any request in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, or the first request coordinate value being greater than or equal to the third coordinate value, determine not to send the any request; Determine to send any of the requests in response to the first requested coordinate value being greater than or equal to the second coordinate value and the first requested coordinate value being less than the third coordinate value; wherein the difference between the third coordinate value and the second coordinate value is equal to the size of the original tensor in the first dimension.
21. A processor, comprising an instruction parsing unit and an execution unit, wherein, the instruction parsing unit is configured to receive and parse a data loading instruction, wherein the data loading instruction includes the shape and size of a tensor to be processed as an input parameter, the starting coordinates of the tensor to be processed in a coordinate system determined by an original tensor, and the tensor to be processed is stored in a buffer area in an NDHWC data storage format, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, and C represents the channel number dimension, and the data storage format of the original tensor is N(C / x)DHW(xC), and x is a positive integer greater than 1; the execution unit executes the data loading instruction after the instruction parsing unit parses the data loading instruction, wherein when the execution unit executes the data loading instruction, it includes performing the following operations: determine a plurality of requests for loading the tensor to be processed in combination with the original tensor, the shape and size of the tensor to be processed, and the starting coordinates of the tensor to be processed; send the plurality of requests in sequence, and sequentially write the sub-data returned by each request into the buffer area to load the tensor to be processed into the buffer area; wherein in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in each dimension lower than the first dimension, divide the requests in the second dimension, the sub-data loaded by each request is either all in the memory or none of it is in the memory, and in response to the sub-data loaded by the request being all in the memory, the sub-data loaded by the request is from the original tensor and is continuously stored in the memory, wherein the first dimension is the same as or adjacent to the second dimension.
22. An electronic device, comprising: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, wherein when the computer-executable instructions are run by the processor, they implement the data loading method according to any one of claims 1-18, or the data storage method according to any one of claims 19-20.
23. A non-transitory computer-readable storage medium, wherein, The non-transient computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they implement the data loading method according to any one of claims 1-18, or the data storage method according to any one of claims 19-20.
Citation Information
Patent Citations
Data processing method and device, processor, electronic equipment and storage medium
CN119089948A
Compilation optimization method and apparatus, computer device and storage medium
WO2023030507A1