Data Loading Method, Data Storage Method, Processor, Electronic Device, and Medium

Through the sequential data acquisition method based on pixel points, the problem of inefficient data loading and storage in the prior art is solved, the data handling efficiency and processor performance are improved, and it is suitable for data loading and storage in parallel processors.

CN120216402BActive Publication Date: 2025-07-25SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607276.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-25
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, data loading and storage methods are inefficient in parallel processors, especially in pixel-by-pixel point computing operations such as convolution operations, which leads to an increase in memory access time and reduces the computing performance of the processor.

Method used

Using the sequential data acquisition method of pixel points, the tensors to be processed are loaded from the original tensor in memory to the cache area. By determining multiple requests to carry data in a continuous or discontinuous manner, ensuring that data is accessed in a specific order in the dimension, improving data handling efficiency.

Benefits of technology

It improves the bandwidth and efficiency of data loading and storage, improves the hardware utilization of computing units, and improves the overall performance of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216402B_ABST
    Figure CN120216402B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a medium. The data loading method is used to load a tensor to be processed from an original tensor in a memory into a buffer area, and includes: determining a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed, and where the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a medium. Background Art

[0002] A tensor is a data structure of a multi-dimensional array. Tensor operations are widely used in processors such as parallel processors. For example, in the field of deep learning, the dimensions of the input data, the intermediate data processed during the deep learning process, and the output data are elastic and not exact. Therefore, an elastic data form is needed to describe various types of data, and thus the concept of a tensor is generated. In the field of deep learning, all data to be operated on can be stored and exist in the form of a tensor. If the data is not in the tensor form, it needs to be first converted into the data structure form of a tensor. As an example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor itself can have any number of dimensions. For example, it can be represented as a 5-dimensional array.

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processing devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of a parallel processor need to process a large amount of data, and this data is generally stored in the storage component of the parallel processor. For example, the storage component can be a high-speed memory. Through a data loading instruction, this data can be loaded from the storage component to the buffer for calculation, and through a data storage instruction, the data in the buffer can be stored in the memory.

[0004] How to provide a fast and efficient data loading / storing method is crucial for the computing performance of the device. Summary of the Invention

[0005] Embodiments of the present disclosure provide a data loading method, a data storage method, a processor, an electronic device, and a medium, which are used to provide a fast and efficient data loading / storing method, improve the memory access bandwidth, and increase the hardware utilization rate of the computing device.

[0006] According to a first aspect of the present disclosure, there is provided a data loading method for loading a tensor to be processed from an original tensor in memory into a buffer. The data storage format of the original tensor in memory is the same as that of the tensor to be processed in the buffer. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape of the tensor to be processed is represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape of the original tensor is represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in 5 dimensions and are all positive integers. The data loading method includes: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests, and writing the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer. When loading the tensor to be processed, due to the dimensional size relationship between the tensor to be processed and the original tensor, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And when the data loaded by the request all belongs to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the data storage format of the original tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0007] According to a second aspect of the present disclosure, there is provided a data loading method, including: receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in a memory to a buffer, where the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component; the shape of the tensor to be processed is represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers, and the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension; the shape of the original tensor is represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers; and after parsing the data loading instruction, using an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; determining a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests, and writing the data corresponding to each request into the buffer in sequence to load the tensor to be processed to the buffer, where in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, partitioning requests are made in the second dimension, the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension in storage or loading, and the first dimension and the second dimension are adjacent.

[0008] According to a third aspect of the present disclosure, there is provided a data storage method for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, wherein the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in memory, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the second tensor are represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape and size of the first tensor are represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers. The data storage method includes: obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for storing the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests and writing the data corresponding to each request into memory in sequence to store the second tensor into memory, wherein in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but continuous acquisition can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data obtained by each request either all belong to the data range of the first tensor or all do not belong to the data range of the first tensor, and in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, wherein the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0009] According to a fourth aspect of the present disclosure, a data storage method is provided, including: receiving a data storage instruction indicating to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, wherein a data storage format of the first tensor in the buffer is the same as a data storage format of the second tensor in memory, the data storage format is used to indicate a storage order and a dimension arrangement of the tensor in a storage component, a shape size of the second tensor is represented by a1, a2, a3, a4, a5, a1, a2, a3, a4, a5 respectively indicate sizes of the second tensor in 5 dimensions and are all positive integers, the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, a shape size of the first tensor is represented by b1, b2, b3, b4, b5, b1, b2, b3, b4, b5 respectively indicate sizes of the first tensor in 5 dimensions and are all positive integers; and after parsing the data storage instruction, executing the data storage instruction using an execution unit, wherein executing the data storage instruction using the execution unit includes: obtaining a number of pixels included in the second tensor, a data storage format of the first tensor, and a starting coordinate of the second tensor in a coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinate of the second tensor in the coordinate system determined by the first tensor, determining a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinate until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests, and sequentially writing data corresponding to each request into memory to store the second tensor into memory, wherein in response to a size relationship between the second tensor and the first tensor in a first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but continuous acquisition can be performed in dimensions lower than the first dimension, partitioning requests in a second dimension, data obtained by each request all belongs to a data range of the first tensor or all does not belong to the data range of the first tensor, and in response to data obtained by a request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, wherein the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0010] According to a fifth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, where the data loading instruction instructs to execute loading a tensor to be processed from an original tensor in a memory to a buffer, where the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component. The shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in five dimensions and are all positive integers; and the execution unit is configured to: execute the data loading instruction. Wherein, the execution unit executes the data loading instruction, including: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor, determining a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor with the starting coordinates as the starting point until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests, and writing the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer, where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0011] According to a sixth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, where the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory, the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component, the shape size of the second tensor is represented by a1, a2, a3, a4, a5, a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in five dimensions and are all positive integers, the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, the shape size of the first tensor is represented by b1, b2, b3, b4, b5, b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in five dimensions and are all positive integers; and the execution unit is configured to: execute the data storage instruction. Wherein, the execution unit executes the data storage instruction, including: obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in a coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in a coordinate system determined by the first tensor, determining a plurality of requests for storing the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into the memory to store the second tensor into the memory, where in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but continuous acquisition can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, the data obtained by each request all belong to the data range of the first tensor or all do not belong to the data range of the first tensor, and in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, where the data storage format of the first tensor indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0012] According to a seventh aspect of the present disclosure, an electronic device is provided, including a processor and a memory connected to the processor, where the processor includes a buffer, and the processor is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor, they implement the data loading method according to the embodiments of the present disclosure to load a tensor to be processed from an original tensor in the memory into the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain a second tensor based on a first tensor in the buffer and write the second tensor into the memory.

[0013] According to an eighth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, a data loading method according to an embodiment of the present disclosure is implemented, or a data storage method according to an embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 FIG. shows a schematic structural diagram of a General-Purpose computing on Graphics Processing Unit (GPGPU);

[0016] Figure 2 FIG. shows a schematic structure of a tensor;

[0017] Figure 3A FIG. shows a schematic diagram of the data storage format of NDHWC;

[0018] Figure 3B FIG. shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32;

[0019] Figure 4 FIG. shows a schematic diagram of obtaining a tensor in the related art;

[0020] Figure 5 FIG. shows a schematic flowchart of a data loading method provided by at least one embodiment of the present disclosure;

[0021] Figure 6 FIG. shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to an embodiment of the present disclosure;

[0022] Figure 7A FIG. shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to the data storage format of NDHWC according to an embodiment of the present disclosure;

[0023] Figure 7B FIG. shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure;

[0024] Figure 8A Shows a schematic diagram of setting boundary values for an original tensor according to an embodiment of the present disclosure;

[0025] Figure 8B Shows a schematic diagram of continuous data acquisition in the case of setting boundaries for an original tensor according to an embodiment of the present disclosure;

[0026] Figure 8C Shows a schematic diagram of continuously acquiring a tensor to be processed in a continuous manner according to the set step size according to an embodiment of the present disclosure;

[0027] Figure 9A Shows a state schematic diagram of a state machine provided according to some embodiments of the present disclosure;

[0028] Figure 9B Shows the state transition of the NDHWC data storage format according to the PerW division method;

[0029] Figure 10 Shows a schematic flowchart of a data loading method provided according to at least one embodiment of the present disclosure;

[0030] Figure 11 Shows a schematic flowchart of a data storage method provided according to at least one embodiment of the present disclosure;

[0031] Figure 12 Shows a schematic flowchart of a data storage method provided according to at least one embodiment of the present disclosure;

[0032] Figure 13 Shows a schematic block diagram of a processor according to some embodiments of the present disclosure;

[0033] Figure 14 Shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure;

[0034] Figure 15 Shows a block diagram of an example computing device implementing some embodiments of the present disclosure; and

[0035] Figure 16 Shows a schematic block diagram of a computer-readable storage medium according to some embodiments of the present disclosure. Detailed implementation manners

[0036] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0037] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Words such as "upper", "lower", "left", "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of this disclosure clear and concise, some detailed descriptions of known functions and known components are omitted in this disclosure.

[0038] Figure 1 A structural schematic diagram of a GPGPU is shown. As Figure 1 shown, a GPGPU is actually an array of programmable multi-processors. For example, the programmable multi-processor can be a Streaming Processor Cluster (SPC for short), for example, including Figure 1 the streaming processor cluster 1 shown, ..., the streaming processor cluster M, where M is a positive integer. In a general-purpose graphics processor, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. As an example, data sharing between multiple streaming processor clusters is performed through a global cache or High Bandwidth Memory (HBM).

[0039] As Figure 1 shown, taking the streaming processor cluster 1 as an example, 1 streaming processor cluster can include multiple Compute Units (CUs for short), for example Figure 1 the compute unit 1, compute unit 2, ..., compute unit K shown, where K is a positive integer. Each compute unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A compute unit can include multiple cores (also called computing cores or compute cores), and each compute core includes an Arithmetic Logic Unit (ALU), a floating-point computing unit, etc., and the compute core is used to perform specific computing tasks. In addition, the compute unit also includes registers (such as Figure 1a register bank) and shared memory for hierarchically storing source data and destination data related to computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit. In addition, the buffer can be understood as being used to share data among the computing units within the streaming processor cluster.

[0040] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processing unit (or parallel computing processor), and then multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block is split into the smallest execution thread bundle (or simply called thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0041] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates thread bundles so that multiple computing cores in the computing unit run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instruction. Memory execution instructions are issued to the shared memory in the computing unit or further issued to the middle-level cache or global cache or high-bandwidth memory for read / write operations, etc.

[0042] As Figure 1 shown, general computing operations, such as matrix computing operations in the field of artificial intelligence, usually require a large amount of data. These data are usually stored in a memory, such as in a high-bandwidth memory HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing result, data needs to be stored in the memory (Store operation). The storage method of data in the memory affects the memory access bandwidth, which in turn affects the hardware utilization rate of the computing unit.

[0043] For example, general computing operations include general matrix multiplication (abbreviated as GEMM). As an example, the data required for general matrix multiplication is represented as two 5D arrays, such as the matrix multiplication calculation of two tensors A and tensor B. In addition, general computing operations also include convolution operations, which are manifested as data dot products. It can be understood that in the field of artificial intelligence, other computing operations are also involved, which will not be listed one by one here. The data involved in these calculations is usually embodied in the form of tensors.

[0044] For example, for a certain tensor A in the buffer, its shape and size can be represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor data in 5 dimensions, and a1, a2, a3, a4, a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, N represents the batch dimension, which is the number of data samples grabbed in one training. The D dimension represents the depth dimension, the H dimension represents the height dimension of the input data, the W dimension represents the width dimension of the input data, and the C dimension represents the number of channels dimension. For example, taking the tensor A as an example, a1 can be the size of the N dimension, a2 can be the size of the D dimension, a3 can be the size of the H dimension, a4 can be the size of the W dimension, and a5 can be the size of the C dimension. Of course, the present disclosure does not make specific limitations on this.

[0045] As an example, Figure 2 shows a schematic structure of a tensor. In Figure 2 the shown tensor, a1 is the size of the N dimension and is equal to 1, a2 is the size of the D dimension and is equal to 1, a3 is the size of the H dimension and is equal to 5, a4 is the size of the W dimension and is equal to 4, and a5 is the size of the C dimension and is equal to 64. For example, Figure 2 the pixel elements of the tensor in

[0046] are represented as 0, 1, 2, 3,... and so on. Figure 2 The placement of the tensor in the memory (such as the memory or the buffer) can have various formats, which are called data storage formats (layout). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The following uses the

[0047] shown tensor to describe different data storage formats. Figure 3A shows a schematic diagram of the NDHWC data storage format.

[0048] For example, for the NDHWC linear mode, as Figure 3A shown, starting from the first element (element 0 in Figure 3A ) of the first channel (a5 = 0, c0 in Figure 3A ), then storing the first element (element 20 in Figure 3A ) of the second channel (a5 = 1, c1 in Figure 3A ), and so on, until all the first elements of all channels are laid out, for example, until the first element of the 64th channel (a5 = 63, c63 in Figure 3A ) (element in Figure 3AAfter the element 1260) in, select the first channel (a5 = 0, Figure 3A the second element of c0) in Figure 3A element 1) in, and then store the second element of the second channel (a5 = 1, Figure 3A c1) in Figure 3A element 21) in, and so on until the second elements of all channels are laid out, and so on.

[0049] In the related art, the data storage format can also include N(C / x)DHW(xC), also known as the Interleave mode, where x can be set to 8, 16, 32, etc. as needed.

[0050] The N(C / x)DHW(xC) data storage format is similar to the NDHWC data storage format, but a key difference is that in the layout of N(C / x)DHW(xC), the a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.

[0051] Figure 3B Shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32.

[0052] As Figure 3B shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 3B c0) in to a5 = 31 ( Figure 3B c31) in, and the second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format.

[0053] In memory, for example, a certain tensor B, similar to a certain tensor A in the buffer, this tensor B can be stored in memory in one of the two data storage formats described above, and the shape dimensions of the tensor can be similarly represented as b1×b2×b3×b4×b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the tensor B in these 5 dimensions and are all positive integers.

[0054] It can be understood that in the related art and possible future developments, the data storage format of tensors is not limited to the above-described N(C / x)DHW(xC) data storage format and NDHWC data storage format. The method described in this disclosure does not limit this. Further, for multiple tensors that can be stored in memory and the buffer during the calculation process, usually the storage space of memory is much larger than that of the buffer, but it is farther from the calculation unit, and the data transfer efficiency is lower compared to the buffer.

[0055] In the related art, during the computing process of a processing device, a large amount of computing data will be generated, for example, in the form of tensors, which can be temporarily stored in a buffer. For example, the buffer here can refer to Figure 1 the buffer in the streaming processor cluster shown in Figure 1 Further, this data can also be transferred from the buffer or directly stored in the memory. For example, the memory can be

[0056] It can be understood that in this article, the data storage process of storing the tensors in the buffer into the memory and the data loading process of loading the tensors in the memory into the buffer can be implemented in a similar manner. Therefore, for the sake of convenience of description, in some embodiments or examples, only the data loading process will be described as an example, and those skilled in the art can apply it similarly to the data storage process. For the differences between the two, additional descriptions will be made.

[0057] As an example, Figure 4 shows a schematic diagram of obtaining a tensor in the related art. As Figure 4 shown, there is a tensor A stored in the memory, which can be placed in the memory in the above N(C / x)DHW(xC) or NDHWC data storage format. That is, tensor A is a 5D array. In Figure 4 only three dimensions W, H, and C are schematically shown. Schematically, on the left side of Figure 4 a dimension coordinate system of tensor A is shown, where they are the W dimension, H dimension, and C dimension respectively. During the data loading process, all or part of the data in tensor A can be loaded into the buffer through a loading instruction. For example, the tensor B in Figure 4 is loaded into the buffer as a whole. Specifically, the loading instruction can indicate the first starting point (C = 0, W = 0, H = 0) of tensor A to be loaded, the second starting point (C = 0, W = 3, H = 0) of tensor B in the coordinate system of tensor A, and the sizes of tensor B in each dimension. Schematically, in Figure 4In the 3D schematic diagram shown, the tensor B is a cuboid determined by the above-mentioned second starting point and the respective dimensional sizes. Based on the information about the second starting point and the respective dimensional sizes of the tensor to be loaded, the memory can load the tensor B into the buffer area.

[0058] In the above-related technologies, the block-based data loading method can be applied to general matrix multiplication GEMM calculations in general computing operations. However, the above block-based data loading method is not applicable to convolution operations, whose computational characteristics are data dot products. The above block-based data loading method will limit the computational efficiency of such pixel-by-pixel type convolution operations, increase data access time, and reduce the overall performance of the processor, which limits the further development space of efficient and general-purpose processors.

[0059] Furthermore, in the memory, the original tensor is continuously stored in the memory in the data storage format as described above. The shape of the original tensor is represented as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in these 5 dimensions and are all positive integers. When extracting or storing a partial tensor in the original tensor, since it is a partial tensor inside the original tensor, it may not be possible to continuously extract in each dimension during extraction, and it is impossible to accurately obtain the amount of data and the data position requested when loading or storing data to optimally achieve data loading or storing, greatly reducing the data bandwidth during data loading and reducing the hardware computational efficiency.

[0060] In view of the above technical problems in the related technologies, the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. In the present disclosure, first, the tensor is no longer obtained and transported in a block form, but is sequentially obtained and transported in units of pixels to be applicable to computational operations such as convolution operations. Further, for the operation of loading the tensor to be processed from the original tensor in the memory into the buffer area, multiple requests for loading the tensor to be processed are determined according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, so as to sequentially obtain data from the original tensor starting from the starting coordinates using the determined multiple requests until the number of obtained data is equal to the number of pixels included in the tensor to be processed, thereby improving the data transportation efficiency for the data loading operation, reducing the memory access time, and improving the overall performance of the processor.

[0061] In the data loading method provided by at least one embodiment of the present disclosure, requests are divided according to whether tensors can be continuously loaded in the first dimension. The data to be loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.

[0062] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0063] The data loading method provided by at least one embodiment of the present disclosure is used to load a tensor to be processed from an original tensor in memory into a buffer area. The memory according to the embodiments of the present disclosure may be, for example, a high-bandwidth memory (HBM), and the buffer area may be, for example, a buffer (buffer) in a streaming processor cluster, which is not limited herein.

[0064] In the method according to the embodiments of the present disclosure, the data storage format of the original tensor in memory is the same as the data storage format of the tensor to be processed in the buffer area. The data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component. For example, the storage component refers to the above-mentioned HBM or buffer area. That is to say, in the data loading method according to the embodiments of the present disclosure, the placement manner of the original tensor in memory is the same as the placement manner of the tensor to be processed to be loaded next in the buffer area. For example, both are arranged in the NDHWC linear manner or both are arranged in the N(C / x)DHW(xC) interleaved manner, which is not limited herein. The method proposed by the present disclosure is applicable to the above two data storage formats.

[0065] The original tensor in memory is a 5D tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, depth dimension, height dimension, width dimension, and number of channels dimension. Similarly, the shape size of the tensor to be processed is represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers.

[0066] For the NDHWC data storage format, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the number of channels dimension.

[0067] For the data storage format of N(C / x)DHW(xC), N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension. Generally, x can be set to an integer multiple of 4.

[0068] Regarding the characteristics of the above two data storage formats, reference can be made to the above in combination with Figure 3A - Figure 3B the description, which will not be repeated here.

[0069] Furthermore, for the 5D data of the original tensor, the data determined by the height dimension and the width dimension represents a pixel (or, it can also be called an element), and it accumulates step by step towards the higher dimension. Among them, the number of pixels is not calculated for the number of channels dimension. As an example, as Figure 4 shown in, the values of each W dimension and H dimension can determine a pixel. For example, W = 0 and H = 0 correspond to the first pixel (or element) in tensor A, and W = 1 and H = 0 correspond to the second pixel in tensor A. In the tensor, the C dimension does not affect the number of pixels.

[0070] Figure 5 is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 5 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S101 - S103.

[0071] In step S101, obtain the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Then, in step S102, combine the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine multiple requests for loading the tensor to be processed, where the multiple requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed. In step S103, sequentially send the multiple requests and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area. According to the embodiment of the present disclosure, the tensor to be processed can be used for convolution operations in the computing units within the processor. In this article, the number of data can be expressed as the number of pixels.

[0072] In the data loading method according to an embodiment of the present disclosure, in order to obtain a tensor to be processed from the original tensor in the memory, it is necessary to indicate the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. For example, the starting coordinates may be the second starting point (C = 0, W = 3, H = 0) described in conjunction with Figure 4 Furthermore, in the data loading method according to an embodiment of the present disclosure, it is also necessary to indicate the number of pixels included in the tensor to be processed (for example, denoted as copy_pixel_num), that is, the total number of pixels in the tensor to be loaded into the buffer. This continuous data copying method is different from the block-based data loading method adopted in the related art above (wherein it is necessary to indicate the sizes of the tensor to be processed in each dimension). According to the data loading method of the embodiment of the present disclosure, the range of the tensor to be processed is determined by the number of pixels to be obtained. Further, instead of obtaining the tensor in a block from the original tensor, data is sequentially obtained from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed.

[0073] Next, first, the implementation process of implementing the above continuous data loading based on copy_pixel_num and starting coordinates will be described, and then, the implementation process of determining multiple requests for loading the tensor to be processed will be described later.

[0074] As an example, Figure 6 shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to an embodiment of the present disclosure. In Figure 6 , tensor A represents the original tensor in the memory. In Figure 6 , each pixel point is shown as a square. Tensor A covers all the squares, that is, it includes both white squares and gray-shaded squares. The starting point of the original tensor is denoted as (C = 0, W = 0, H = 0), for example, to indicate the position of the original tensor in the memory. Then, in Figure 6 , tensor B represents the tensor to be obtained, and its starting point is denoted as (C = 0, W = 3, H = 0), that is, data is obtained starting from the 3rd pixel point in the first row of tensor A. Then, in Figure 6 , the process of obtaining the tensor to be processed in a sequential (or also called continuous) manner according to the embodiment of the present disclosure is schematically shown. That is, starting from the starting point (C = 0, W = 3, H = 0) as shown by the dashed arrow, data is obtained from tensor A pixel by pixel until the number of obtained pixels reaches copy_pixel_num. In the example of Figure 6 , only the three dimensions of W, H, and C are shown. It can be understood that if the data in these 3 dimensions still does not reach copy_pixel_num, data in higher dimensions can be further obtained, which is not limited herein. In addition, asFigure 6 As shown, the C dimension itself does not affect the number of pixels. Suppose, Figure 6 the data storage format of tensor A in [reference] in memory is NDHWC. In Figure 6 the example shown, the total number of pixels of the tensor to be processed is copy_pixel_num = 27.

[0075] In an embodiment according to the present disclosure, the step of sequentially obtaining data from the original tensor starting from the starting coordinates includes: removing the channel number dimension from the 5 dimensions of the original tensor, and in the order of dimensions from low to high, using the starting coordinates as the starting point for data acquisition, and sequentially obtaining data until the number of pixels obtained reaches copy_pixel_num. The reason for removing the channel number dimension from the 5 dimensions is that for a tensor, the channel number dimension does not count the number of pixels. That is to say, S102 may include obtaining pixel data from the original tensor starting from the starting coordinates in the order of the width dimension, height dimension, depth dimension, and batch dimension until the number of pixels obtained reaches copy_pixel_num.

[0076] Compare Figure 4 and Figure 6 the two ways of obtaining the tensors to be processed shown, the data loading method provided by the embodiments of the present disclosure can achieve per-pixel data loading, rather than Figure 4 the block-by-block acquisition in [reference]. This data loading method is more conducive to the calculation process such as convolution operations and is beneficial to improving the operation efficiency. Thus, based on the method provided by the embodiments of the present disclosure, for the data to be subjected to convolution operations next, the processor can, for example, indicate in the form of an instruction to fetch this part of the data from the memory to the buffer area according to the above continuous data loading method for use in convolution operations. It can be understood that the above processor can reasonably use whether to use the continuous data loading method or the block-by-block data loading method according to the type of operation to be performed or the data processing characteristics, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer area may also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0077] According to some embodiments of the present disclosure, there are multiple tensors stored in the memory, and the data loading method may further include: obtaining indication information of the storage location of the original tensor in the memory. As an example, the indication information may include the starting coordinates of the original tensor in the memory and the size values in five dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. It can be understood that the storage space of the memory is generally much larger than the buffer, and various tensor data generated during the calculation process can be stored therein. These tensor data can be placed in the memory according to any one of the above-mentioned N(C / x)DHW(xC) or NDHWC data storage formats. In the embodiments according to the present disclosure, in order to enable the memory to know the specific location of the tensor to be obtained in the memory, indication information of the storage location of the original tensor in the memory may also be obtained. The indication information includes the starting coordinates of the original tensor in the memory and the size values in five dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. As an example, in the case where the original tensor is in the NDHWC data storage format, tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5 respectively represent the values of the original tensor in the five dimensions of N, D, H, W, and C.

[0078] Regarding the influence of the data storage format on the loading process of the tensor to be processed ("tensor B"), it will be described below in conjunction with Figure 7A - Figure 7B for description. In conjunction with Figure 7A and Figure 7B In the example shown, the data storage method of the tensor to be processed in the buffer is the same as the data placement method of the original tensor in the memory.

[0079] Figure 7A FIG. shows a schematic diagram of continuously obtaining the tensor to be processed according to the NDHWC data storage format in an embodiment of the present disclosure. For Figure 7A the tensor shown, its data storage format in the memory is in the form of NDHWC, or is arranged according to the NDHWC placement method.

[0080] In Figure 7A In the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 8 (the C dimension is not shown in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Taking the N dimension as an example, N_dim = 2 corresponds to Figure 7AFor N0 and N1 among them, based on the above parameters, the total number of pixel points included in the original tensor can be obtained as 64.

[0081] Next, in Figure 7A the example of, the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixel points included in the tensor to be processed copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7A the pixel points where W = 2 and H = 2 under the dimensions of N = 0 and D = 0 shown in, starting from this pixel point, according to the data arrangement method, in the order from low dimension to high dimension (that is, in the order of W, H, D, N, as shown by the data acquisition schematic arrow in Figure 7A ), sequentially acquire data until the number of acquired data is equal to the number of pixel points included in the tensor to be processed. In Figure 7A the example shown in, starting from the determined starting point, acquire data pixel by pixel according to the arrow order, so as to load this part of the data in the original tensor in memory into the buffer area as the tensor to be processed for subsequent operations such as convolution calculations.

[0082] Figure 7B shows a schematic diagram of continuously acquiring the tensor to be processed for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure. For Figure 7B the tensor shown, its data storage format in memory is in the form of N(C / x)DHW(xC), or is referred to as the arrangement method of N(C / x)DHW(xC). In Figure 7B the example of, x = 8, that is, every 8 C channels are grouped and bound to the W dimension. Compared with Figure 7A the NDHWC data storage format shown in, for the data storage format of N(C / x)DHW(xC), the level of the C dimension in the tensor is higher and is located between the N dimension and the D dimension.

[0083] In Figure 7B the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 16 (shown as two 8*C dimensions in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Thus, the total number of pixel points included in the original tensor can be obtained as 64.

[0084] Next, in Figure 7B the example of Figure 7B , the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixels included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7B the pixel point where W = 2 and H = 2 under N = 0 and D = 0 and the first 8*C dimension shown in Figure 7A . Starting from this pixel point, according to the data arrangement method, in the order from low dimension to high dimension (i.e., in the order of W, H, D, C, N, as shown by the data acquisition schematic arrow in Figure 7B ), data is sequentially acquired until the number of acquired data is equal to the number of pixels included in the tensor to be processed. Compared with the acquisition order shown in Figure 7A and Figure 7B , since the level of the C dimension is higher, in the example of Figure 7B , first, data of W, H, D, and the first 8*C dimension are acquired according to the starting point, and then, data of the second 8*C dimension are acquired in the order of W, H, D, and finally, data of the N dimension are acquired. It can be understood that in the examples of

[0085] and Figure 7B , the number of pixels of the tensor to be processed acquired is 40, and the difference is only caused by the different arrangement methods of the C dimension. In addition, it should be noted that for the second 8*C dimension shown in

[0086] Figure 7B , the starting points at the W and H dimensions should be aligned with the starting point of the first 8*C dimension.

[0085] In the example shown in Figure 7B , starting from the determined starting coordinates, data is acquired pixel by pixel in the order of the arrow for loading this part of the data in the original tensor in the memory into the buffer as the tensor to be processed for subsequent operations such as convolution calculations.

[0086] In some embodiments according to the present disclosure, boundary values may also be defined for at least a part of the five dimensions of the original tensor respectively. As an example, left and right boundary values may be set for, for example, the C dimension, the W dimension, the H dimension, and the D dimension, which may be expressed as (L-tensor_C, R-tensor_C), (L-tensor_W, R-tensor_W), (L-tensor_H, R-tensor_H), (L-tensor_D, R-tensor_D). The above boundary values are used to define the boundary range for obtaining data from the original tensor, where the boundary values are arbitrary values compared to the size values of the original tensor in this dimension.

[0087] In the method described in combination with Figure 7A and Figure 7B the range (i.e., the boundary value) of the original tensor is not defined, that is, the tensor to be processed is obtained from the complete data of the original tensor. In the method according to the embodiments of the present disclosure, it is also proposed that the range for obtaining the tensor to be processed from the original tensor can be delimited by setting boundary values. Further, in the implementation process, the boundary values can be set to arbitrary values compared to the size values of the original tensor in this dimension. That is to say, the boundary values can exceed the range of the original tensor itself.

[0088] As an example, Figure 8A shows a schematic diagram of setting boundary values for the original tensor according to the embodiments of the present disclosure. As Figure 8A shown, the rectangular box represents the range covered by the original tensor, which can be any dimension in the original tensor, such as the C dimension, the W dimension, the H dimension, or the D dimension. Generally, boundary values are not set for the N dimension. In the six subgraphs of Figure 8A the relationship between the left boundary value and the right boundary value (shown as "L" and "R" in Figure 8A ) and the size value of the original tensor in this dimension is respectively shown.

[0089] According to the embodiments of the present disclosure, when setting boundary values, sequentially obtaining data from the original tensor includes: for the data part of the original tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the original tensor defined by the boundary values, for example, referring to the order described in Figure 7A and Figure 7B . In comparison, for the data part of the original tensor not covered by the range defined by the boundary values, it is represented as invalid data. For invalid data, the buffer directly fills the tensor to be processed with a predetermined value, where the predetermined value is equal to 0.

[0090] As an example, in Figure 8AIn the first sub - picture, both the left boundary value L and the right boundary value R are on the left side of the original tensor. That is, the data to be obtained in this dimension are all invalid data. In this case, for the invalid data, for example, this part of the data can be automatically filled by sending a zero - filling instruction to the buffer. Another example is, in Figure 8A In the second sub - picture, the left boundary value L is on the left side of the left boundary of the original tensor data range, and the right boundary value R is on the left side of the right boundary of the original tensor. That is, for the tensor to be processed to be obtained, part of it is invalid data and part of it is valid data in the original tensor. Schematically, in Figure 8A , the data corresponding to the slanted shaded part represents valid data, and the rest are all invalid data. In this case, for the valid data, for example, refer to Figure 7A and Figure 7B for the described order, while for the invalid data, for example, this part of the data can be automatically filled by sending a zero - filling instruction to the buffer.

[0091] In the method according to the embodiments of the present disclosure, by setting boundary values for the original tensor, the range from which the tensor to be processed is to be taken can be further delimited. In practical applications, this implementation method can adapt to the characteristics of operations such as convolution. For example, it is beneficial to reduce the amount of calculation and greatly improve the flexibility of data. As an example, assume that the original tensor corresponds to an intermediate tensor for feature extraction of an entire input picture, and the input picture includes a specific target, such as an object to be recognized, and the object does not cover the entire picture, that is, the picture includes a background part. In this case, by setting boundary values, the range of the tensor to be processed to be obtained can be limited to the part of the original tensor corresponding to the specific target, so as to reduce the amount of calculation of subsequent operations such as convolution and improve the processing efficiency.

[0092] As an example, Figure 8B shows a schematic diagram of continuous data acquisition in the case where boundaries are set for the original tensor. Among them, compared with the situation shown in Figure 6 , Figure 8B can be understood as setting boundary values (bound) for the C - dimension, W - dimension, and H - dimension of the tensor A in Figure 6 , and the boundaries are all within the size ranges of the tensor A in the C - dimension, W - dimension, and H - dimension, which is equivalent to the boundary situation shown in the fourth sub - picture in Figure 8A . Specifically, Figure 8B the outer box as a whole corresponds to the tensor A (corresponding to the tensor A in Figure 6 ). After setting the boundaries therein, according to the data loading method of the embodiments of the present disclosure, data will be sequentially acquired starting from the starting - point coordinates within the set boundary range until the number of acquired data is equal to the number of pixels included in the tensor to be processed. InFigure 8B Among them, the data part composed of squares corresponds to the data range delimited by the boundary values set for the C dimension, W dimension, and H dimension, and data is sequentially acquired from the data range delimited by the boundary. Data located outside the data range delimited by the boundary can be regarded as invalid data. Regarding Figure 8B The process of sequentially acquiring data in Figure 6 can be referred to in combination with

[0093] In some embodiments according to the present disclosure, a step value for acquiring data can also be set. As an example, the set data step value is used to specify the stride for sequentially acquiring data from the original tensor. Among them, for the data of the original tensor, starting from the starting coordinate, sequentially acquiring data from the original tensor includes: for the data of the original tensor, starting from the starting coordinate, sequentially acquiring data from the original tensor according to the data step value.

[0094] Figure 8C FIG. shows a schematic diagram of continuously acquiring a tensor to be processed according to the set stride in an embodiment of the present disclosure. Among them, the stride is equal to 2, that is, one data is taken every other pixel point, that is, only the pixels shown in the shaded part are sequentially acquired, and a stride equal to 2 means skipping one pixel. Similarly, when the stride is equal to 3, it means skipping two pixels, and so on. It can be understood that when the set stride is equal to 1, it corresponds to Figure 7A - Figure 7B the data loading method of each pixel point shown. In practical applications, by setting the stride, the computational amount of subsequent operations such as convolution operations on the data can be further reduced, the processing efficiency can be improved, and in addition, the flexibility of data loading can be further increased.

[0095] In the data loading method provided according to the embodiments of the present disclosure, a new data acquisition mode different from the block-based data acquisition method is provided, that is, pixel-by-pixel data loading can be realized according to the starting point and the number of pixels to be acquired (as shown in Figure 6 ), rather than Figure 4 the block-based acquisition in

[0096] The above in combination with Figure 6 、 Figure 7A - Figure 7B 、 Figure 8A - Figure 8C, which details the implementation process of sequentially obtaining data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels (copy_pixel_num) included in the tensor to be processed in the data loading method according to an embodiment of the present disclosure.

[0097] Next, the implementation process of determining multiple requests for loading the tensor to be processed by combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor in step S102 shown in Figure 5 will be described.

[0098] In an embodiment according to the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. Moreover, in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the data storage format of the original tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent. As an example, when continuous loading cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, requests are divided in the second dimension.

[0099] As an example, when the data storage formats of the original tensor and the tensor to be processed are NDHWC, the first dimension can be the C dimension, the second dimension can be the adjacent W dimension, and the C dimension has priority over the W dimension during storage or loading. Specifically, in response to the size relationship between the tensor to be processed and the original tensor in the C dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the C dimension but can be performed in dimensions lower than the C dimension (in this case, since the C dimension is already the lowest dimension, dimensions lower than the C dimension are no longer considered), requests are divided in the W dimension so that the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. For ease of description, this way of dividing requests can be referred to as the request division rule according to the W dimension (PerW).

[0100] As another example, when the data storage formats of the original tensor and the tensor to be processed are NDHWC, the first dimension can be the W dimension, the second dimension can be the adjacent H dimension, and the W dimension takes precedence over the H dimension during storage or loading. Specifically, in response to the size relationship between the tensor to be processed and the original tensor in the W dimension such that when loading the tensor to be processed, it cannot be continuously loaded in the W dimension but can be continuously loaded in dimensions lower than the W dimension (in this case, the C dimension is lower than the W dimension, that is, the data cannot be continuously loaded in the W dimension but can be continuously loaded in the C dimension lower than the W dimension), the requests are divided in the H dimension so that the data loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor. For ease of description, this way of dividing requests can be called the request division rule according to the H dimension (PerH).

[0101] It can be understood that for the case where the data storage formats of the original tensor and the tensor to be processed are NDHWC, the division rules for other dimensions can be deduced by analogy and will not be repeated here. According to the method described above, the way of dividing requests can also include: the request division rule according to the D dimension (PerD, that is, the data cannot be continuously loaded in the H dimension but can be continuously loaded in the W and C dimensions lower than the H dimension), the request division rule according to the N dimension (PerN, that is, the data cannot be continuously loaded in the D dimension but can be continuously loaded in the H, W, and C dimensions lower than the D dimension), and the request division rule according to the overall dimension (Per1, that is, the data cannot be continuously loaded in the N dimension but can be continuously loaded in the D, H, W, and C dimensions lower than the N dimension). It can be understood that for the request division rule of Per1, for the operation of loading the tensor to be processed in the original tensor into the buffer, it is entirely implemented by one request, that is, these data are loaded into the buffer through one instruction, because in this case, the data to be loaded are continuously stored in each dimension and the data loading process can be executed through one instruction. Generally speaking, for the NDHWC case, the ways of dividing requests can include 5 types: PerW, PerH, PerD, PerN, Per1.

[0102] For the case where the data storage format of the original tensor and the tensor to be processed is N(C / x)DHW(xC), the rules for request partitioning are similar to those described above for the NDHWC case, except that when the data storage format is N(C / x)DHW(xC), the priority of the C dimension is between the N dimension and the D dimension. From this, it can be obtained that for the data storage format of N(C / x)DHW(xC), the request partitioning methods can include the following schemes: according to the request partitioning rule for the H dimension (PerH), that is, the data in the W dimension at the lowest dimension cannot be continuously loaded; according to the request partitioning rule for the D dimension (PerD), that is, the data in the H dimension cannot be continuously loaded but the data in the W dimension lower than the H dimension can be continuously loaded; according to the request partitioning rule for the C dimension (PerC), that is, the data in the D dimension cannot be continuously loaded but the data in the H dimension and the W dimension lower than the D dimension can be continuously loaded; according to the request partitioning rule for the N dimension (PerN), that is, the data in the C dimension cannot be continuously loaded but the data in the D dimension, H dimension, and W dimension lower than the C dimension can be continuously loaded; and according to the request partitioning rule for the overall dimension (Per1), that is, the data in the N dimension cannot be continuously loaded but the data in other dimensions lower than the N dimension can be continuously loaded. In this case, for the C dimension bound to the W dimension, it is required that copy_c = 8 * c, where copy_c represents the size of the tensor to be processed in the C dimension to be obtained.

[0103] As an implementation method, for the data storage format of N(C / x)DHW(xC), there is also a request partitioning rule according to the C dimension, that is, PerC. In this case, the data in the D dimension and the H dimension and W dimension lower than the D dimension can be continuously loaded. However, since xC is bound to the W dimension, when the C dimension is discontinuous or copy_c is greater than xC, the request partitioning rule of PerC is also followed.

[0104] As another implementation method, as described above, for the data storage format of N(C / x)DHW(xC), the xC dimension is bound to the W dimension, that is, the data itself is continuous in the W dimension. According to the above-described scheme, the request partitioning rule of PerH should be followed. However, there is also a case where when the stride set for the W dimension is greater than 1 (stride_x > 1), for example, as Figure 8C shown, the stride of the W dimension is equal to 2, which will make the data to be obtained for the W dimension discontinuous, that is, one data is obtained every other pixel. Therefore, in this case, the request partitioning rule of PerW needs to be used to partition the multiple requests for loading data, rather than following the request partitioning rule of PerH.

[0105] Generally speaking, for the case of N(C / x)DHW(xC), the requested partitioning methods can include six types: PerW, PerH, PerD, PerC, PerN, and Per1.

[0106] It can be understood that, compared with the five requested partitioning methods for the NDHWC case, the requested partitioning rules for the N(C / x)DHW(xC) case are logically similar, and the only difference lies in the priority of the C dimension. This is because for the placement method of N(C / x)DHW(xC), the xC dimension is bound to the W dimension, that is, the data in the lowest xC dimension is continuous.

[0107] According to some embodiments of the present disclosure, the determination of whether continuous loading can be performed in the first dimension is specified as follows: When the first dimension is the channel number dimension, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension. Further, when the first dimension is the channel number dimension and the data storage format of the original tensor is NDHWC, in response to the size (copy_c) of the tensor to be processed in the channel number dimension being equal to the size tensor_c of the original tensor in the channel number dimension, and the starting coordinate (c_coord_b) of the tensor to be processed in the channel number dimension being equal to the starting coordinate of the original tensor in the channel number dimension, it is determined that the size relationship in the first dimension enables continuous loading in the first dimension when loading the tensor to be processed. For other dimensions, the process of determining whether continuous loading can be performed is similar to that of the C dimension and will not be described one by one here.

[0108] The object to be loaded in the present disclosure, that is, the tensor to be processed, is split into multiple different requests according to a certain rule and sent sequentially, and the returned data is written into the buffer area in order, thereby efficiently loading the tensor to be processed into the memory. A principle during splitting is that the data loaded by each request is either all in the memory or all not in the memory. This is because when the requested data is in the memory, a request is sent to the memory, and when the requested data is not in the memory, the request can be converted into an operation such as writing a predetermined value to the buffer area by the hardware. This splitting method can send corresponding loading requests to different hardware more reasonably.

[0109] In addition, during splitting, if the size relationship between the tensor to be processed and the original tensor in the first dimension is such that when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, then a further division of the requests is performed in the second dimension, so that the requests can be split more reasonably, the data stored continuously in memory can be retained as much as possible, the number of requests can be reduced, the tensor to be processed can be loaded efficiently, the bandwidth when loading or storing data from / to memory can be significantly increased, the performance can be improved, and the efficiency of the hardware computing unit can be improved.

[0110] In an implementation solution where boundary values are set for the original tensor, the data loading method according to the embodiments of the present disclosure may further include: obtaining boundary values for the original tensor on at least a part of the five dimensions respectively, and the boundary values are used to define the boundary range for obtaining data from the original tensor. In some implementation manners, the boundary values may be arbitrary values compared with the size values of the original tensor in this dimension. As an example, the boundary value set for the W dimension may be any of the Figure 8A cases shown. Further, for the dimension with boundary values set, the data loaded by each request either all belongs to the valid data range defined by the data boundary of the original tensor itself and the boundary values, or all does not belong to the valid data range defined by both the data boundary of the tensor itself and the set boundary values. As Figure 8A shown, for the first sub - figure, both the left and right boundary values are located on the left side of the data range of the original tensor, that is, both are outside the data range of the original tensor, thereby indicating that all the data to be obtained for this dimension is invalid (for example, represented as out of bound, oob). For Figure 8A the second sub - figure in, the left boundary is located on the left side of the data range of the original tensor, and the right boundary is located within the data range of the original tensor. Thus, only the part of the data that is both within the data range of the original tensor and to the left of the right boundary R is valid data (in Figure 8A , the data corresponding to the diagonal shaded part is represented as valid data), and the rest of the data to be obtained is invalid data. According to the embodiments of the present disclosure, for the dimension with boundary values set, the rule for dividing requests is that the data loaded by each request is either all valid data or all invalid data (i.e., oob), that is to say, valid data and invalid data cannot be loaded by the same request.

[0111] According to an embodiment of the present disclosure, for multiple requests determined according to step S102, in the actual implementation process of the processor, it can be implemented by setting a state machine. For the state machine, multiple parameters required to determine the above requests can be set, and according to the specific values of the parameters and the initial parameters of the original tensor and the tensor to be processed, the specific tensor to be processed in the original tensor is loaded into the buffer area. The implementation scheme related to determining multiple requests for loading the tensor to be processed will be described below in conjunction with the state machine.

[0112] According to some embodiments of the present disclosure, in step S102, in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, multiple requests for loading the tensor to be processed are determined, including: based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, determining the first request among the multiple requests and the initial state of the first request entering the state machine, where the first request indicates the number of pixels to be acquired by the request; and based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, using the state machine to determine each request after the first request among the multiple requests.

[0113] As an example, for the task of loading the tensor to be processed in the original tensor in memory into the buffer area, the first request can be determined first based on the initial parameters. For example, based on the number of pixels copy_pixel_num included in the tensor to be processed, the data storage format of the original tensor (such as NDHWC or N(C / x)DHW(xC)), and the starting coordinates of the tensor to be processed (the starting coordinates indicate the position of the first pixel of the tensor to be processed in the coordinate system of the original tensor, for example, represented as its positions in 5 dimensions, c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b), determining the first request among the multiple requests and the initial state of the first request entering the state machine, where the first request indicates the number of pixels to be acquired by the request (req_size_1).

[0114] Specifically, taking the data storage format as NDHWC, copy_pixel_num = 40, and the starting coordinates are w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0 as an example, in combination with Figure 7A Describe how to determine the partitioning method of the requests and the first request. As Figure 7AAs shown, for the above starting coordinates, the starting pixel corresponding to the starting coordinates in the original tensor is obtained, that is, data is obtained starting from this starting pixel. Then, for the NDHWC data storage format, the C dimension is the lowest dimension. Thus, with the C dimension as the first dimension, it is determined whether continuous loading can be performed in the C dimension. In response to the size relationship between the tensor to be processed and the original tensor in the C dimension such that continuous loading cannot be performed in the C dimension when loading the tensor to be processed, the above PerW request partitioning method is adopted, that is, requests are partitioned one by one for each W. That is to say, each pixel is loaded into the buffer by one request. According to the judgment on whether continuous loading can be performed in the C dimension described above: In response to the size of the tensor to be processed in the C dimension (copy_c) not being equal to the size of the original tensor in the C dimension (tensor_c), and / or the first coordinate value (c_coord_b) of the starting coordinates of the tensor to be processed in the C dimension not being equal to the second coordinate value (for example, 0) of the starting coordinates of the original tensor in the C dimension, it is determined that continuous loading cannot be performed in the C dimension for the tensor to be processed. Summarized as follows:

[0115] It is determined that the data to be obtained is discontinuous in the C dimension when any of the following conditions holds:

[0116] (1) The starting coordinate c_coord_b of the tensor to be processed in the C dimension is not equal to 0;

[0117] (2) The size copy_c of the tensor to be processed in the C dimension is greater than the size tensor_c of the original tensor in the C dimension;

[0118] (3) The size copy_c of the tensor to be processed in the C dimension is less than the size tensor_c of the original tensor in the C dimension;

[0119] For example, the coordinates of the data loaded by each request are different in the C dimension, but the same in the W dimension, H dimension, D dimension, and N dimension. At this time, it can be understood that without considering the depth direction, the original tensor is unfolded into a one-dimensional vector composed of multiple pixels (pixel) in the order of WHDN. The data loaded by each request is the data within one pixel, that is, requests are partitioned in the W dimension, and the next request can be to load the data in the adjacent next pixel, that is, requests are partitioned in a pixel-skipping manner.

[0120] After determining the partitioning method, for example, in the PerW manner, the corresponding first request can be made, referring to Figure 7A , the first request is used to load the starting pixel corresponding to the starting coordinates. Its initial state is equal to the starting coordinates of the tensor to be processed, and the number of pixels to be obtained by this first request is req_size_1 = 1 because subsequent requests are partitioned in a pixel-by-pixel manner. The specific acquisition order is as Figure 7Ain the order indicated by the arrow until the number of pixels to be retrieved is equal to copy_pixel_num = 40 of the tensor to be processed.

[0121] If continuous loading is possible in the C dimension, then it is determined whether continuous loading is possible in the W dimension adjacent to and higher than the C dimension. If continuous loading is not possible in the W dimension and continuous loading is possible in the C dimension lower than the W dimension, the above-mentioned request partitioning method of PerH is adopted, that is, the request is partitioned row by row. That is to say, each row of pixels is loaded into the buffer by one request.

[0122] Next, still taking the data storage format as NDHWC and the second coordinate value as 0 as an example, assuming that continuous loading is possible in the C dimension, then, with the first dimension as the W dimension and the second dimension as the H dimension for judgment, it is determined that the W dimension is discontinuous when any of the following conditions is satisfied:

[0123] (1) The starting coordinate of the tensor to be processed in the W dimension is not equal to 0;

[0124] (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension;

[0125] (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension.

[0126] At this time, when the C dimension is continuous, the coordinates of the data to be loaded by each request are different in the C dimension and the W dimension, but the coordinates in the H dimension, D dimension, and N dimension are the same. At this time, it can be understood that the data loaded by each request belongs to the same row, that is, the requests are partitioned in the H dimension, and the data loaded by different requests is located in different rows.

[0127] Assuming that the first dimension is the H dimension and the second dimension is the D dimension, it is determined that the H dimension is discontinuous when any of the following conditions is satisfied:

[0128] (1) The starting coordinate of the tensor to be processed in the H dimension is not equal to 0;

[0129] (2) The size copy_h of the tensor to be processed in the H dimension is greater than the size tensor_h of the original tensor in the H dimension;

[0130] (3) The size copy_h of the tensor to be processed in the H dimension is less than the size tensor_h of the original tensor in the H dimension.

[0131] At this time, when both the C dimension and the W dimension are continuous, the coordinates of the data to be loaded by each request are different in the C dimension, W dimension, and H dimension, but the coordinates in the D dimension and N dimension are the same.

[0132] Assume that the first dimension is the D dimension and the second dimension is the N dimension. Determine discontinuity in the D dimension when any of the following conditions holds:

[0133] (1) The starting coordinate of the tensor to be processed in the D dimension is not equal to 0;

[0134] (2) The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension;

[0135] (3) The size copy_d of the tensor to be processed in the D dimension is less than the size tensor_d of the original tensor in the D dimension.

[0136] At this time, when the C dimension, W dimension, and H dimension are all continuous, the coordinates of the data to be loaded for each request are different in the C dimension, W dimension, H dimension, and D dimension, but the coordinates in the N dimension are the same. That is, the requests are partitioned in the D dimension.

[0137] Assume that the first dimension is the N dimension. Determine discontinuity in the N dimension when any of the following conditions holds:

[0138] (1) The starting coordinate of the tensor to be processed in the N dimension is not equal to 0;

[0139] (2) The size copy_n of the tensor to be processed in the N dimension is greater than the size tensor_n of the original tensor in the N dimension;

[0140] (3) The size copy_n of the tensor to be processed in the N dimension is less than the size tensor_n of the original tensor in the N dimension.

[0141] At this time, when the C dimension, W dimension, H dimension, and D dimension are all continuous, the coordinates of the data to be loaded for each request are the same in the CWHDN dimensions. For the tensor data located in memory, in essence, at this time, one request can be sent to the memory to load the data in the memory, and other requests can be used to fill the data that is not located in the memory, that is, invalid data.

[0142] The request partitioning logic for the tensor to be processed with the data storage format of N(C / x)DHW(xC) is the same as that of NDHWC. The difference is that the dimension arrangement of N(C / x)DHW(xC) is different from that of NDHWC. For N(C / x)DHW(xC), due to its special interleaved structure, it is considered to be necessarily continuous in (xC) by default. Therefore, the lowest dimension is considered to be the W dimension, followed by the H dimension, then the D dimension, then the C dimension, and the highest is the N dimension.

[0143] For example, taking the data storage format as N(C / x)DHW(xC) and the second coordinate value as 0, assuming that the first dimension is D dimension and the second dimension is C dimension, discontinuity in D dimension is determined when any of the following conditions is met:

[0144] (1) The starting coordinate of the tensor to be processed in dimension D is not equal to 0;

[0145] (2) The size of the tensor to be processed in dimension D, copy_d, is larger than the size of the original tensor in dimension D, tensor_d;

[0146] (3) The size of the tensor to be processed in dimension D, copy_d, is smaller than the size of the original tensor in dimension D, tensor_d.

[0147] At this time, when both the W dimension and the H dimension are continuous, the coordinates of the data to be loaded by each request are different in the W dimension, the H dimension, and the D dimension, but the coordinates in the N dimension are the same, that is, the requests are divided in the C dimension.

[0148] In addition, for N(C / x)DHW(xC), due to the particularity of the interleaving pattern (xC), the pixels are continuous, so the first dimension from low to high is W, H, D, C, N. For the conditions that cannot be loaded continuously when the data storage format is N(C / x)DHW(xC) and the division principle of the request, please refer to the relevant content of NDHWC, which will not be listed here one by one.

[0149] As described above for NDHWC, for the N(C / x)DHW(xC) format, after determining the partitioning method, for example, according to the PerH method, the first request can be similarly determined, referring to Figure 7B , the first request is used to load the row of pixels corresponding to the starting pixel of the starting coordinate. Its initial state is equal to the starting coordinate of the tensor to be processed, and the number of pixels to be obtained by the first request is req_size_1=2, because subsequent requests are divided row by row. The specific acquisition order is as follows Figure 7B The order shown by the arrows in is repeated until the number of pixels to be obtained is equal to copy_pixel_num=40 of the tensor to be processed.

[0150] According to an embodiment of the present disclosure, after determining a request, it is further necessary to determine a data read address of a starting position of reading data from memory for the request, and a data write address for indicating a starting position of writing data to a buffer. Among them, based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, determining the first request among multiple requests and an initial state when the first request enters a state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address of writing the tensor to be processed in the buffer as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0151] According to some embodiments of the present disclosure, in the case where the data storage format of the original tensor is NDHWC, for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes:

[0152] Calculating the data read address Addr1_n of the nth request according to the following formula:

[0153] Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord),

[0154] wherein, u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the original tensor in five dimensions.

[0155] According to some embodiments of the present disclosure, in the case where the data storage format of the original tensor is N(C / x)DHW(xC), for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes:

[0156] Calculate the data read address Addr2_n of the nth request according to the following formula:

[0157] Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0158] c_coord * tensor_d * tensor_h * tensor_w +

[0159] d_coord * tensor_h * tensor_w * xC +

[0160] h_coord * tensor_w * xC +

[0161] w_coord * xC

[0162] Wherein, u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in five dimensions, and wherein, xC represents the size occupied by x channels.

[0163] According to some embodiments of the present disclosure, for multiple requests for loading a tensor to be processed, determining the data write address Addr3_n of the nth request among the multiple requests includes:

[0164] Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1

[0165] Wherein, b_addr_base is the starting address for writing the tensor to be processed in the buffer, and req_size_1, req_size_2, ..., req_size_n-1 represent the number of pixels to be acquired by the first n-1 requests respectively.

[0166] For the first sent request, the request initial coordinates corresponding thereto are the starting coordinates of the tensor to be processed. Thus, the data read address of the first request can be calculated with reference to the above formula. Take the starting address for writing the tensor to be processed in the buffer as the data write address of the first request.

[0167] The length of the data loaded by the first request can be determined according to parameters such as the determined request partitioning method, the starting coordinates of the tensor to be processed, and the size of the original tensor. For example, if the sum of the first coordinate value t_coord_b of the starting coordinates of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is less than the second coordinate value (i.e., the coordinate value of the starting coordinates of the original tensor in the first dimension, for example, 0), the length of the data loaded by the first request is the shape size copy_t of the tensor to be processed in the first dimension. For example, if the sum of the first coordinate value t_coord_b of the starting coordinates of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value (i.e., the coordinate value of the starting coordinates of the original tensor in the first dimension, for example, 0), the length of the data loaded by the first request is the absolute value of the first coordinate value t_coord_b.

[0168] Considering the advantages of the state machine having a clear logical structure, being easy to maintain and expand, and being particularly suitable for processing logical scenarios with multiple conditions and multiple branches, in the present disclosure, the state machine is adopted to automatically update the request initial coordinates and the length of the loaded data corresponding to each request, avoiding complex conditional nesting, with clear logic, easy to maintain, and strong scalability.

[0169] According to some embodiments of the present disclosure, based on the starting coordinates of the tensor to be processed, determining the initial state of the first request entering the state machine includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0170] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be obtained by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, using the state machine to determine each request after the first request among multiple requests includes: based on the initial state, using the state machine to determine the request initial coordinates of the second request after the first request and the number of pixels to be obtained by the second request; according to the shape size of the original tensor, the request initial coordinates of the second request, and the data storage format of the original tensor, determining the data read address of the second request; according to the number of pixels to be obtained by the second request, determining the data write address of the second request; updating the state machine according to the information related to the second request, and sequentially determining the subsequent requests in each request based on the updated state machine.

[0171] Figure 9A Shows a state diagram of a state machine provided according to some embodiments of the present disclosure.

[0172] For example, the states of the state machine include a first state s0, a second state s1, and a third state s2. The initial state of entering the state machine is determined by the starting coordinates c_coord_b, w_coord_b, h_coord_b, d_coord_b, and n_coord_b of the tensor to be processed.

[0173] For example, in response to the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinate of the original tensor in the first dimension, the initial state is determined to be the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, the initial state is determined to be the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, the initial state is determined to be the third state.

[0174] Refer to Figure 9A , the tensor_t direction represents the data in the first dimension. The data with the coordinate in the t dimension less than the second coordinate value and greater than the third coordinate value is not in memory or not within the data range of the original tensor ( Figure 9A the part shown by the diagonal shading, and this part of the data can be called invalid data), and the data with the size of the t dimension between the second coordinate value and the third coordinate value ( Figure 9A the white part) is stored in memory, which is the actual size of the original tensor data in the first dimension.

[0175] The state only switches to itself and adjacent states. For example Figure 9A in, the first state s0 can jump to the first state s0 or the second state s1, the second state s1 can jump to the first state s0, the second state s1, and the third state s2, and the third state s2 can jump to the first state s0, the second state s1, and the third state s2.

[0176] Figure 9A The 6 cases (① to ⑥) in

[0177] For example, Figure 9A in case ①, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the second coordinate value. The state machine always loops and jumps in the first state s0 on the left.

[0178] For example, Figure 9AIn case ②, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value and less than the third coordinate value. The state machine loops and jumps between the first state s0 and the second state s1.

[0179] For example, Figure 9A In case ③, the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value. The state machine always loops and jumps in the second state s1.

[0180] For example, Figure 9A In case ④, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1, and the third state s2.

[0181] For example, Figure 9A In case ⑤, the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps between the second state s1 and the third state s2.

[0182] For example, Figure 9A In case ⑥, the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops and jumps in the third state s2.

[0183] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and combine with the trigger condition to determine the state entered by the second sent request. Then, based on the state entered by the second sent request, determine the request initial coordinates corresponding to the third sent request and the data length loaded by the second sent request, and combine with the trigger condition to determine the state entered by the third sent request, and so on.

[0184] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request in each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state. After determining the request initial coordinates corresponding to the current request, it can be judged whether the data loaded by the current request is in the memory according to the request initial coordinates.

[0185] For example, in response to the coordinate value of the request initial coordinate corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the data to be loaded by the current request is located in the memory. It can be understood that for the data range that all belongs to the original tensor, it will all be within the memory range (this part of the data is actually stored in the memory), and for the data range that all does not belong to the original tensor, it will all not be within the memory range and is not the actually stored data. That is, the data exceeding the original tensor range can be understood as invalid data. For example, this part of the invalid data can be directly filled with zeros in the buffer area. At this time, the data read address and data write address of the current request can be determined with reference to the above formula, and then the current request is sent to the memory to load the corresponding data into the buffer area.

[0186] For example, in response to the coordinate value of the request initial coordinate corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the data to be loaded by the current request is not located in the memory. At this time, in response to all the data to be loaded by the current request not being located in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values into the buffer area, where the number of the multiple predetermined values is determined by the length of the data to be loaded specified by the current request. For example, the predetermined value is 0.

[0187] The process of using the state machine to determine the request initial coordinate and the length of the loaded data corresponding to each request will be specifically described below.

[0188] Figure 9B Shows the state transition for the PerW partitioning method of the NDHWC data storage format. The condition for the PerW partitioning method is that c_coord is not equal to 0, or copy_c is not equal to tensor_c. At this time, the two adjacent pixels to be acquired are not continuous in the memory, and the request needs to be split pixel by pixel to acquire the data.

[0189] At Figure 9B On the left side in Figure 9B only two states, s1 and s2, are shown for the C dimension. This is because in practical applications, generally the coordinate of the C dimension is not less than 0, that is, there is no first state. Specifically, s1 can correspond to the above-mentioned second state, and s2 can correspond to the above-mentioned third state. In addition, in the state loop schematic diagram on the right side of

[0190] Table 1

[0191]

[0192] That is to say, at Figure 9BIn the example, the blank box on the right corresponds to the size of the original tensor in the C dimension (tensor_c), and its left boundary can correspond to, for example, Figure 9A the second coordinate value in, for example, equal to 0, and its right boundary can correspond to, for example, Figure 9B the third coordinate value in, for example, equal to tensor_c, corresponding to state s1 when the size copy_c of the tensor to be processed in the C dimension is between 0 and tensor_c, and corresponding to state s2 when copy_c is greater than tensor_c.

[0193] Figure 9B Jumps can occur between the various states shown in. The following Table 2 shows the conditions for jumps between the various states:

[0194] Table 2

[0195]

[0196] In the above Table 2, the parameter remain_copy_p represents the number of remaining pixels not yet fetched. For example, before the first request, remain_copy_p is equal to the number of pixels copy_pixel_num corresponding to the tensor to be processed. After the first request, remain_copy_p is equal to copy_pixel_num - req_size_1. remain_copy_c represents the size of the remaining data not yet fetched in the C dimension, used to characterize the data fetching information in the C dimension. For example, before the first request, remain_copy_c = copy_c.

[0197] The following describes the information that needs to be updated in each state and how to update it.

[0198] First, in state s1, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table:

[0199] Table 3

[0200]

[0201] It should be noted that in the example of the above Table 3, the situation of setting boundary values and data fetching steps for the original tensor is further considered. Among them, right_bound represents the right boundary value set for the W dimension, down_bound represents the boundary value in the H dimension, and hind_bound represents the boundary value in the D dimension. The setting of boundary values and steps can refer to the above in combination with Figure 8A - Figure 8CDescription. In addition, for the case where no boundary value is set, it can be understood that each boundary is equal to the size of the original tensor in that dimension. In Table 3, stride_x, stride_y, and stride_z respectively represent the stride values in the W dimension, H dimension, and D dimension. As an implementation, the strides can all be set to 1. In addition, it can be understood that in other embodiments according to the present disclosure, boundary values and strides can also be set for the C dimension, for example, which is not limited herein.

[0202] The parameter update rules in Table 3 are described below. First, for the C dimension with the lowest priority, the parameter c_coord represents the coordinate of the current request in the C dimension. When c_coord + remain_copy_c <= tensor_c, that is, the next request for the C dimension does not exceed the boundary tensor_c of the original tensor in the C dimension, the parameter c_coord is updated to the starting coordinate c_coord_b of the C dimension; when c_coord + remain_copy_c > tensor_c, that is, the data acquisition for the C dimension in the next request exceeds the boundary tensor_c of the original tensor in the C dimension, the parameter c_coord is updated to tensor_c.

[0203] The coordinate update logic for the other dimensions in Table 3 is the same as that for the W dimension, which is described separately as follows. For the W dimension, the parameter w_coord represents the coordinate of the current request in the W dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be acquired in the next request is still valid data, when w_coord + stride_x >= right_bound, that is, when the W dimension exceeds the boundary, the coordinate w_coord of the W dimension is updated to w_coord + stride_x - copy_w, where copy_w represents the size of the tensor to be processed in the W dimension; otherwise (w_coord + stride_x < right_bound), that is, when it does not exceed the boundary, the coordinate w_coord of the W dimension is updated to w_coord + stride_x. When stride_x = 1, it means that w_coord is incremented by 1, that is, the data to be acquired in the next request is the next data immediately adjacent to the data to be acquired in the current request. Otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate w_coord of the W dimension remains unchanged.

[0204] For the H dimension, the parameter h_coord represents the coordinate of the current request in the H dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be retrieved by the next request is still valid data, at this time w_coord + stride_x >= right_bound. When h_coord + stride_y >= down_bound, that is, when the H dimension exceeds the boundary, the coordinate h_coord of the H dimension is updated to h_coord + stride_y - copy_h, where copy_h represents the size of the tensor to be processed in the H dimension; otherwise (h_coord + stride_y < down_bound), that is, when it does not exceed the boundary, the coordinate h_coord of the H dimension is updated to h_coord + stride_y. Otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate h_coord of the H dimension remains unchanged.

[0205] For the D dimension, the parameter d_coord represents the coordinate of the current request in the D dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be retrieved by the next request is still valid data, at this time w_coord + stride_x >= right_bound and h_coord + stride_y >= down_bound. When d_coord + stride_z >= hind_bound, that is, when the D dimension exceeds the boundary, the coordinate d_coord of the D dimension is updated to d_coord + stride_z - copy_d, where copy_d represents the size of the tensor to be processed in the D dimension; otherwise (d_coord + stride_z < hind_bound), that is, when it does not exceed the boundary, the coordinate d_coord of the D dimension is updated to d_coord + stride_z. Otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate d_coord of the D dimension remains unchanged.

[0206] For the N dimension, the parameter n_coord represents the coordinates of the current request in the N dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be retrieved by the next request is still valid data, at this time w_coord + stride_x >= right_bound and h_coord + stride_y >= down_bound and d_coord + stride_z >= hind_bound, the coordinate n_coord of the N dimension is incremented by 1, that is, updated to n_coord + 1; otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate n_coord of the N dimension remains unchanged. It can be seen that the update logic of the N dimension is slightly different from that of the W, H, and D dimensions because generally, no boundary values and step sizes are set for the N dimension.

[0207] In addition, in addition to the above dimension coordinate parameters, the state machine also needs to maintain two parameters remain_copy_p and remain_copy_c. Among them, remain_copy_p is the number of remaining pixels that have not been retrieved. For example, before the first request, remain_copy_p is equal to the number of pixels copy_pixel_num corresponding to the tensor to be processed. After the first request, remain_copy_p is equal to copy_pixel_num - req_size_1; remain_copy_c represents the size of the remaining data to be retrieved in the C dimension. For example, before the first request, remain_copy_c = copy_c. After the first request, remain_copy_c is equal to copy_c minus the size corresponding to the C dimension retrieved by the first request.

[0208] Specifically, as shown in Table 3, for remain_copy_c, when c_coord + remain_copy_c <= tensor_c, remain_copy_c is updated to copy_c; otherwise (c_coord + remain_copy_c > tensor_c), remain_copy_c is updated to c_coord + remain_copy_c - tensor_c.

[0209] For remain_copy_p, when c_coord + remain_copy_c <= tensor_c, remain_copy_p is updated to remain_copy_p - 1; otherwise (c_coord + remain_copy_c > tensor_c), remain_copy_p remains unchanged.

[0210] Second, in the s2 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table:

[0211] Table 4

[0212]

[0213] For Table 4, the state updates of each parameter can refer to the description for Table 3 and will not be elaborated here. Based on the above update logic, the update results of Table 4 can be derived similarly.

[0214] The state transitions of the data storage format of NDHWC according to the PerW partitioning method are described above in conjunction with Tables 1 - 4, which involve states s1 and s2, and the update outputs of the state machine in these two states. It can be understood that the principles and update logics of the state machine for other partitioning methods (e.g., PerH) of the data storage format of NDHWC are similar to those described above and will not be described here.

[0215] Of course, as the coordinates continue to increase, when the number of acquired pixels is equal to the number of pixels included in the tensor to be processed, the state machine can enter the idle state, indicating that the data of the current pen instruction has been acquired. Taking the data storage format of NDHWC as an example, for the PerW request partitioning method, when the remaining values of the N, D, H, and W dimensions are equal to 1 and the data of the C dimension has been acquired, the idle state is entered; for the PerH request partitioning method, when the remaining values of the N, D, and H dimensions are equal to 1 and the data of the W dimension has been acquired, the idle state is entered; for the PerD request partitioning method, when the remaining values of the N and D dimensions are equal to 1 and the data of the H dimension has been acquired, the idle state is entered; for the PerN request partitioning method, when the remaining value of the N dimension is equal to 1 and the data of the D dimension has been acquired, the idle state is entered; for the Per1 request partitioning method, when the data of the N dimension has been acquired, the idle state is entered.

[0216] The specific state transition conditions of the state machine and the updated coordinate content can be changed and set according to actual needs, and the logic is similar to the state machine logic described above, and will not be exemplified one by one here.

[0217] It should be noted that the process of using the state machine to determine the request - corresponding initial coordinates and the length of the loaded data is basically the same for NDHWC and N(C / x)DHW(xC). The difference is that in response to the data storage format of the tensor to be processed being NDHWC, the coordinate value of the channel number dimension is incremented by 1 when updated, and in response to the data storage format of the tensor to be processed being N(C / x)DHW(xC), the coordinate value of the channel number dimension is incremented by x when updated.

[0218] In the above embodiments, requests are divided according to whether the tensors can be continuously loaded in the first dimension. Each request is used to load sub-data that is either all located in the memory or all not located in the memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.

[0219] According to some embodiments of the present disclosure, another data loading method is provided. Figure 10 FIG. shows a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 10 shown, the data loading method according to the embodiments of the present disclosure includes step S201 and step S202.

[0220] In step S201: Receive a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in the memory to a buffer area.

[0221] According to the embodiments of the present disclosure, the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer area. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers.

[0222] In step S202: After parsing the data loading instruction, use an execution unit to execute the data loading instruction.

[0223] As Figure 10 shown, wherein, step S202 uses an execution unit to execute the data loading instruction, including:

[0224] S2021: Obtain the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor;

[0225] S2022: Determine a plurality of requests for loading the tensor to be processed, in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and

[0226] S2023: Sequentially send the plurality of requests and sequentially write the data corresponding to each request into the buffer to load the tensor to be processed into the buffer.

[0227] For the related descriptions of the tensor to be processed and the original tensor and the specific implementation processes of steps S2021 - S2023, reference may be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.

[0228] As an example, the data loading instruction may be a machine instruction, or the data loading instruction may also be a micro-instruction. For example, the data loading instruction is implemented in the form of a Load instruction.

[0229] According to some embodiments of the present disclosure, there is also provided a data storage method for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory. It can be understood that the data storage method according to the embodiments of the present disclosure can be understood as the reverse process of the data loading method described above, that is, moving data from the buffer to memory. The implementation principle thereof according to the embodiments of the present disclosure is similar to the above data loading method, and the repeated parts will not be described again, and only the different parts will be described in detail.

[0230] In the data storage method according to the embodiments of the present disclosure, the first tensor is a 5D tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. Among them, for the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and it accumulates step by step to higher dimensions, where the number of channels dimension does not calculate the number of pixels. As an example, the first tensor may refer to the tensor stored in the buffer, and it may be any one of the data storage formats of NDHWC or N(C / x)DHW(xC), and there is no limitation thereto. In terms of understanding the implementation principle, it can be correspondingly understood as the original tensor in the data loading method described above, that is, obtaining part or all of the data from the first tensor and transferring it to memory. This part of the tensor obtained from the first tensor is represented as the second tensor, which can be correspondingly understood as the tensor to be processed in the data loading method described above.

[0231] Figure 11 shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure. As Figure 11 shown, the data storage method includes steps S301 - S303.

[0232] In step S301, obtain the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor. Then, in step S302, combine the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for storing the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor. In step S303, sequentially send the plurality of requests and write the data corresponding to each request into the memory in sequence to store the second tensor into the memory.

[0233] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, divide the requests in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer area, where the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0234] In the data storage method according to an embodiment of the present disclosure, the manner of sequentially obtaining data from the first tensor can be referred to and described in combination with Figure 6 the description shown. Compared with the block - type data storage method in the related art, the data storage method provided by the embodiment of the present disclosure can achieve per - pixel data storage instead of Figure 4 the block - type acquisition in [reference]. This data storage method is more beneficial to calculation processes such as convolution operations and is conducive to improving the operation efficiency. It can be understood that, for example, the processor can reasonably use, according to the type of operation to be performed or data processing characteristics, whether to use a continuous data loading method or a block - type data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer area can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0235] In practical applications, for example, in the case of convolution operations in general computing operations, the operation process is multi-layered. For example, after the first layer of processing is completed, the data result needs to be stored for use in the next layer of processing. As described above, for the data that needs to perform convolution operations, adopting the sequential data acquisition method provided by the embodiments of the present disclosure (which can refer to the description combined with Figure 5 、 Figure 6 、 Figure 7A 、 Figure 7B 、 Figure 8A 、 Figure 8B and Figure 8C ) is more conducive to improving the computing efficiency. As an example, the weight data of 1x3 can only obtain data in the form of a sliding window, so it can only be obtained row by row, unless the shape of the graph to be obtained is very regular, exactly a cube and the next layer does not require operations such as padding.

[0236] According to some embodiments of the present disclosure, a plurality of tensors are stored in the buffer. The first tensor is one of the plurality of tensors. The data storage format of the plurality of tensors in the buffer can be any one of NDHWC or N(C / x)DHW(xC). Among them, for the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the channel number dimension; for the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the channel number dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the channel number of xC is bound to the width dimension. Among them, the data represented by the height dimension and the width dimension represents a pixel, and accumulates gradually to higher dimensions. Among them, the channel number dimension does not calculate the number of pixels. Generally, x can be set to an integer multiple of 4.

[0237] According to some embodiments of the present disclosure, in the case where the first dimension is the channel number dimension, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the first dimension, it is determined that the second tensor cannot be continuously acquired in the first dimension.

[0238] According to some embodiments of the present disclosure, when the first dimension is the channel number dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the channel number dimension being equal to the size of the channel number dimension of the original tensor, and the starting coordinate of the tensor to be processed in the channel number dimension being equal to the starting coordinate of the original tensor in the channel number dimension, determine the size relationship of the first dimension such that when loading the tensor to be processed, it is continuously loaded in the first dimension.

[0239] According to some embodiments of the present disclosure, in combination with the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinate of the second tensor in the coordinate system determined by the first tensor, determine multiple requests for storing the second tensor, including: based on the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinate of the second tensor, determine the first request among the multiple requests and the initial state when the first request enters the state machine, where the first request indicates the number of pixels to be acquired by the request; and based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the second tensor, and the data storage format of the first tensor, use the state machine to determine each request among the multiple requests after the first request.

[0240] According to some embodiments of the present disclosure, each request includes a data read address for indicating the starting position for reading data from the buffer, a data write address for indicating the starting position for writing data to the memory, and the number of pixels to be acquired by the request. Among them, determining the first request among the multiple requests and the initial state when the first request enters the state machine based on the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinate of the second tensor includes: based on the starting coordinate of the second tensor, determine the initial state when the first request enters the state machine; use the starting coordinate of the second tensor as the request initial coordinate corresponding to the first request; according to the shape and size of the first tensor, the request initial coordinate corresponding to the first request, and the data storage format of the second tensor, determine the data read address of the first request; determine the starting address for writing the second tensor in the memory as the data write address of the first request; and according to the number of pixels included in the second tensor, the data storage format of the first tensor, the starting coordinate of the second tensor, and the shape and size of the first tensor, determine the number of pixels to be acquired by the first request.

[0241] According to some embodiments of the present disclosure, determining an initial state in a state machine for a first request based on a starting coordinate of a second tensor includes: determining the initial state as a first state in response to a first coordinate value of the starting coordinate of the second tensor in a first dimension being less than a second coordinate value of the starting coordinate of a first tensor in the first dimension; determining the initial state as a second state in response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, where a difference between the second coordinate value and the third coordinate value is equal to a size of the first tensor in the first dimension; and determining the initial state as a third state in response to the first coordinate value being greater than or equal to the third coordinate value.

[0242] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the second tensor, and a data storage format of the first tensor, determining each request after the first request among a plurality of requests by using a state machine includes: based on the initial state, using the state machine to determine a request initial coordinate corresponding to a second request after the first request and the number of pixels to be acquired by the second request; determining a data read address of the second request according to a shape size of the first tensor, the request initial coordinate corresponding to the second request, and the data storage format of the first tensor; determining a data write address of the second request according to the number of pixels to be acquired by the second request; and updating the state machine according to information related to the second request, and sequentially determining subsequent requests among each request based on the updated state machine.

[0243] According to some embodiments of the present disclosure, when the data storage format of the first tensor is NDHWC, for a plurality of requests for storing the second tensor, determining a data read address of an nth request among the plurality of requests includes:

[0244] Calculating a data read address Addr1_n of the nth request according to the following formula:

[0245] Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord),

[0246] where u_addr_base represents a storage address in a buffer of a pixel at a starting coordinate position of the first tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent shape sizes of the first tensor in five dimensions.

[0247] According to some embodiments of the present disclosure, when the data storage format of the first tensor is N(C / x)DHW(xC), for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes:

[0248] Calculating the data read address Addr2_n of the nth request according to the following formula:

[0249] Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0250] c_coord * tensor_d * tensor_h * tensor_w +

[0251] d_coord * tensor_h * tensor_w * xC +

[0252] h_coord * tensor_w * xC +

[0253] w_coord * xC

[0254] wherein, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the first tensor in five dimensions.

[0255] According to some embodiments of the present disclosure, for multiple requests for storing the second tensor, determining the data write address Addr3_n of the nth request among the multiple requests includes:

[0256] Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1

[0257] wherein, b_addr_base is the starting address for writing the second tensor in the memory, and req_size_1, req_size_2,..., req_size_n-1 represent the number of pixels to be acquired by the first n-1 requests respectively.

[0258] The data storage method according to an embodiment of the present disclosure may further include: obtaining boundary values for the first tensor in at least a part of five dimensions, where the boundary values are used to define the boundary range for obtaining data from the first tensor. For example, the boundary values may be any values compared to the size values of the first tensor in that dimension. As an example, the boundary value set for the W dimension may be any of the cases shown as Figure 8A shown. Further, for the dimension with boundary values set, all the data requested to be loaded either belongs to the data boundary of the first tensor itself and within the valid data range specified by the boundary values, or does not belong to the valid data range specified by both the data boundary of the tensor itself and the set boundary values.

[0259] It can be understood that the data storage method according to an embodiment of the present disclosure can achieve similar technical effects to the data loading method according to an embodiment of the present disclosure.

[0260] According to some embodiments of the present disclosure, another data storage method is provided. Figure 12 The schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown. As Figure 12 shown, the data storage method according to an embodiment of the present disclosure includes step S401 and step S402.

[0261] In step S401: Receive a data storage instruction indicating to execute obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into the memory.

[0262] According to an embodiment of the present disclosure, the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape size of the second tensor is represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in five dimensions and are all positive integers. The five dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The shape size of the first tensor is represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in five dimensions and are all positive integers.

[0263] In step S402: After parsing the data storage instruction, use the execution unit to execute the data storage instruction.

[0264] As Figure 12 shown, where step S402 uses the execution unit to execute the data storage instruction, including:

[0265] S4021: Obtain the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor;

[0266] S4022: Combine the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine multiple requests for storing the second tensor, where the multiple requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and

[0267] S4023: Sequentially send the multiple requests and write the data corresponding to each request into the memory in sequence to store the second tensor into the memory.

[0268] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer area, where the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0269] For the related descriptions of the first tensor and the second tensor and the specific implementation processes of steps S4021 - S4023, reference can be made to the related descriptions of the foregoing data storage method, and the repeated parts will not be elaborated.

[0270] As an example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction is implemented in the form of a Store instruction.

[0271] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, where the data loading instruction instructs to load a tensor to be processed from an original tensor in a memory into a buffer. The data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in five dimensions and are all positive integers. And the execution unit is configured to: execute the data loading instruction.

[0272] According to an embodiment of the present disclosure, the execution unit executes the data loading instruction, including: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to write the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer.

[0273] According to an embodiment of the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension. The data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0274] For example, multiple tensors are stored in memory, and the original tensor is one of the multiple tensors. The data storage format of the multiple tensors in memory can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensors in the storage component. Among them, for the NDHWC data storage format, N corresponds to b1, representing the batch dimension; D corresponds to b2, representing the depth dimension; H corresponds to b3, representing the height dimension; W corresponds to b4, representing the width dimension; C corresponds to b5, representing the number of channels dimension. For the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension; (C / x) corresponds to b2, representing the number of channels dimension; D corresponds to b3, representing the depth dimension; H corresponds to b4, representing the height dimension; W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0275] For example, the tensor to be processed is used for convolution operations in the computing units within the processor. It can be understood that the above-mentioned processor can be reasonably used according to the type of operation to be performed or the data processing characteristics, whether it is a continuous data loading method or a block data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0276] For example, when the first dimension is the number of channels dimension, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0277] For example, when the first dimension is the number of channels dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the number of channels dimension being equal to the size of the original tensor in the number of channels dimension, and the starting coordinate of the tensor to be processed in the number of channels dimension being equal to the starting coordinate of the original tensor in the number of channels dimension, it is determined that the size relationship in the first dimension enables continuous loading of the tensor to be processed in the first dimension.

[0278] For example, when the data storage format of the original tensor is NDHWC, for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes:

[0279] Calculating the data read address Addr1_n of the nth request according to the following formula:

[0280] Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord),

[0281] where u_addr_base represents the memory storage address of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in five dimensions.

[0282] For example, in the case where the data storage format of the original tensor is N(C / x)DHW(xC), for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes:

[0283] Calculating the data read address Addr2_n of the nth request according to the following formula:

[0284] Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0285] c_coord * tensor_d * tensor_h * tensor_w +

[0286] d_coord * tensor_h * tensor_w * xC +

[0287] h_coord * tensor_w * xC +

[0288] w_coord * xC

[0289] where u_addr_base represents the memory storage address of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in five dimensions.

[0290] For example, for multiple requests for loading the tensor to be processed, determining the data write address Addr3_n of the nth request among the multiple requests includes:

[0291] Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1

[0292] Wherein, b_addr_base is the starting address for writing the tensor to be processed in the buffer, and req_size_1, req_size_2, …, req_size_n-1 represent the number of pixels to be obtained for the first n-1 requests respectively.

[0293] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, wherein the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor to memory, wherein the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in memory, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the second tensor are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape and size of the first tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers; and the execution unit is configured to: execute the data storage instruction.

[0294] According to an embodiment of the present disclosure, the execution unit executes the data storage instruction, including: obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request to memory to store the second tensor to memory.

[0295] In accordance with an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, it cannot be continuously fetched in the first dimension but can be continuously fetched in dimensions lower than the first dimension, a division of requests is performed in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. Moreover, in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer area, where the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0296] For example, multiple tensors are stored in the buffer area, and the first tensor is one of the multiple tensors. The data storage format of the multiple tensors in the buffer area can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensors in the storage component. Among them, for the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension; for the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0297] For example, in the case where the first dimension is the number of channels dimension, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the first dimension, it is determined that the second tensor cannot be continuously fetched in the first dimension.

[0298] For example, in the case where the first dimension is the number of channels dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the number of channels dimension being equal to the size of the original tensor in the number of channels dimension, and the starting coordinate of the tensor to be processed in the number of channels dimension being equal to the starting coordinate of the original tensor in the number of channels dimension, it is determined that the size relationship in the first dimension enables continuous loading of the tensor to be processed in the first dimension.

[0299] For example, in the case where the data storage format of the first tensor is NDHWC, for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes:

[0300] Calculate the data read address Addr1_n of the nth request according to the following formula:

[0301] Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord),

[0302] where u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the first tensor in five dimensions.

[0303] For example, when the data storage format of the first tensor is N(C / x)DHW(xC), for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes:

[0304] Calculate the data read address Addr2_n of the nth request according to the following formula:

[0305] Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w +

[0306] c_coord * tensor_d * tensor_h * tensor_w +

[0307] d_coord * tensor_h * tensor_w * xC +

[0308] h_coord * tensor_w * xC +

[0309] w_coord * xC

[0310] where u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the first tensor in five dimensions.

[0311] For example, for multiple requests for storing a second tensor, determining the data write address Addr3_n of the nth request among the multiple requests includes:

[0312] Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1

[0313] where b_addr_base is the starting address for writing the second tensor in the memory, and req_size_1, req_size_2, …, req_size_n-1 represent the number of pixels to be acquired by the first n-1 requests respectively.

[0314] As an example, Figure 13 shows a schematic block diagram of a processor according to some embodiments of the present disclosure. As Figure 13 shown, the processor 1000 may include an instruction parsing unit 1010 and an execution unit 1020. It can be understood that the processor 1000 may be implemented to execute the data loading method according to the embodiments of the present disclosure to load the tensor to be processed from the original tensor in the memory to the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain the second tensor based on the first tensor in the buffer and write the second tensor to the memory.

[0315] Regarding the specific implementation processes of the data storage method and the data loading method, reference may be made to the above description and will not be repeated here. The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0316] According to some embodiments of the present disclosure, an electronic device is further provided. Figure 14 shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure. As Figure 14 shown, the electronic device 2000 may include a processor 2010 and a memory 2020 connected to the processor 2010. In addition, the processor 2010 may further include a buffer. According to the embodiments of the present disclosure, the memory 2020 may be implemented in the form of a high-bandwidth memory HBM, which is not limited thereto. Specifically, according to the embodiments of the present disclosure, the processor 2010 is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor 2010, the data loading method according to the embodiments of the present disclosure is implemented to load the tensor to be processed from the original tensor in the memory to the buffer, or the data storage method according to the embodiments of the present disclosure is implemented to obtain the second tensor based on the first tensor in the buffer and write the second tensor to the memory.

[0317] The processor 2010 can perform various actions and processes according to a program stored in a non-transitory memory such as a non-volatile memory. Specifically, the processor 2010 can refer to a processor chip capable of performing parallel computing. For example, it can be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In addition, the processor 2010 can also be implemented as other conventional types of processors, which are not limited herein.

[0318] Regarding the specific implementation processes of the data storage method and the data loading method, reference can be made to the above description, and details will not be repeated here. The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0319] Figure 15 A block diagram showing an example computing device implementing some embodiments according to the present disclosure is shown. As Figure 15 shown, the computing device 3000 is, for example, suitable for implementing the data loading method or the data storage method provided by the embodiments of the present disclosure. It should be noted that Figure 15 the components of the computing device 3000 shown are exemplary and not restrictive. According to actual application needs, the computing device 3000 may also have other components.

[0320] As Figure 15 shown, the computing device 3000 may include a processing device 3010 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.

[0321] For example, when the computer-readable instructions are executed by the processing device 3010, one or more steps of the data loading method according to any of the above embodiments, or one or more steps of the data storage method according to any of the above embodiments can be performed. It should be noted that for a detailed description of the processing process of the data loading method, reference can be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference can be made to the relevant descriptions in the embodiments of the data storage method above.

[0322] For example, the processing device 3010, the read-only memory (ROM) 3020, and the random access memory (RAM) 3030 are connected to each other via the bus 3040. The input / output (I / O) interface 3050 is also connected to the bus 3040.

[0323] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 3030 and / or cache memory, etc. For example, computer-readable instructions can be loaded from the storage device 3080 into the random access memory (RAM) 3030 to execute the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 3020, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, etc. Various application programs and various data can also be stored in the computer-readable storage medium, as well as various data used and / or generated by the application programs, etc.

[0324] Generally, the following devices can be connected to the I / O interface 3050: an input device 3060 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 3070 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 3080 including, for example, magnetic tape, a hard disk, a flash memory, etc.; and a communication device 3090. The communication device 3090 can allow the computing device 3000 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 15A computing device 3000 with various devices is shown. However, it should be understood that it is not required to implement or have all the shown devices, and the computing device 3000 may alternatively implement or have more or fewer devices. For example, the processing device 3010 may control other components in the computing device 3000 to perform desired functions. The processing device 3010 may be a Central Processing Unit (CPU), a Tensor Processing Unit (TPU), or a Graphics Processing Unit (GPU), etc., which has data processing capabilities and / or program execution capabilities. The GPU may be directly integrated into a System on Chip (SOC), directly integrated onto the motherboard, or built into the North Bridge chip of the motherboard.

[0325] According to some embodiments of the present disclosure, a non-transitory computer-readable storage medium is also provided. The non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they implement the data loading method according to the embodiments of the present disclosure, or implement the data storage method according to the embodiments of the present disclosure.

[0326] Figure 16 It is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 16 shown, the computer-readable storage medium 4000 may be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 4010 may be non-temporarily stored on the storage medium 4000. For example, when the computer-readable instructions 4010 are executed by a processor, one or more steps of the data loading method described in any of the above embodiments may be executed, or one or more steps of the data storage method described in any of the above embodiments may be executed. It should be noted that for a detailed description of the processing procedure of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing procedure of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0327] As an example, the storage medium 4000 may be applied to the electronic device 2000 and / or the computing device 3000. For example, the storage medium 4000 may be implemented as the storage device 3080 in the computing device 3000.

[0328] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0329] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases.

[0330] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on. The above description is only for the preferred embodiments of the present disclosure and the illustration of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.

[0331] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0332] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

[0333] For the present disclosure, the following points also need to be noted:

[0334] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.

[0335] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.

[0336] The above are only the specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the appended claims.

Claims

1. A data loading method for loading a tensor to be processed from an original tensor in memory into a buffer, where The data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape of the tensor to be processed is represented by a1, a2, a3, a4, and a5. a1, a2, a3, a4, and a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The shape of the original tensor is represented by b1, b2, b3, b4, and b5. b1, b2, b3, b4, and b5 respectively indicate the dimensions of the original tensor in the 5 dimensions and are all positive integers. The data loading method includes: Obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, to determine a plurality of requests for loading the tensor to be processed. Among them, the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and Sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the buffer to load the tensor to be processed into the buffer. Wherein, in response to the dimension relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in the dimensions lower than the first dimension, requests are divided in the second dimension. The data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory. Wherein, the data storage format of the original tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

2. The method according to claim 1, wherein, Multiple tensors are stored in the memory, and the original tensor is one of the multiple tensors. The data storage format of the multiple tensors in the memory is any one of NDHWC or N(C / x)DHW(xC). Among them, for the data storage format of NDHWC, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension; for the data storage format of N(C / x)DHW(xC), N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension. Among them, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions. Among them, the number of channels dimension does not calculate the number of pixels.

3. The method according to claim 1, wherein, The tensor to be processed is used for convolution operations in the computing units within the processor.

4. The method according to claim 1, wherein, When the first dimension is the number of channels dimension, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

5. The method according to claim 4, wherein, When the first dimension is the number of channels dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the number of channels dimension being equal to the size of the number of channels dimension of the original tensor, and the starting coordinate of the tensor to be processed in the number of channels dimension being equal to the starting coordinate of the original tensor in the number of channels dimension, it is determined that the size relationship of the first dimension enables continuous loading in the first dimension when loading the tensor to be processed.

6. The method according to claim 1, wherein, Combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinate of the tensor to be processed in the coordinate system determined by the original tensor, multiple requests for loading the tensor to be processed are determined, including: Based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinate of the tensor to be processed, the first request among the multiple requests and the initial state when the first request enters the state machine are determined, where the first request indicates the number of pixels to be acquired by this request; and Based on the initial state, combining the number of pixels to be acquired by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, the state machine is used to determine each request among the multiple requests after the first request.

7. The method according to claim 6, wherein, Each request includes a data read address for indicating a starting position to read data from the memory, a data write address for indicating a starting position to write data to the buffer, and the number of pixels to be acquired by the request. Among them, determining the first request among the multiple requests and the initial state when the first request enters the state machine based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed includes: Determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; Taking the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; Determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; Determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and Determining the number of pixels to be acquired by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

8. The method according to claim 7, wherein Determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed includes: In response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; In response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and In response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

9. The method according to claim 6, wherein Based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, using the state machine to determine each request after the first request among the multiple requests includes: Based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired by the second request; Determining the data read address of the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; Determining the data write address of the second request according to the number of pixels to be acquired by the second request; Updating the state machine according to the information related to the second request, and sequentially determining the subsequent requests among the respective requests based on the updated state machine.

10. The method according to claim 7, wherein, When the data storage format of the original tensor is NDHWC, for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord), where u_addr_base represents the storage address in the memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in the five dimensions.

11. The method according to claim 7, wherein, When the data storage format of the original tensor is N(C / x)DHW(xC), for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr2_n of the nth request according to the following formula: Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC where u_addr_base represents the storage address in the memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in the five dimensions.

12. The method according to claim 10 or 11, wherein, For multiple requests for loading the tensor to be processed, determining the data write address Addr3_n of the nth request among the multiple requests includes: Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 Among them, b_addr_base is the starting address in the buffer area where the tensor to be processed is written, and req_size_1, req_size_2, ..., req_size_n-1 represent the number of pixels to be obtained by the first n-1 requests respectively.

13. The method according to claim 1, further comprising: Obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the original tensor, Among them, for the dimensions with boundary values set, the data loaded by each request belongs to the valid data range defined by the original tensor and the boundary values or does not belong to the valid data range.

14. A data loading method, comprising: Receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in memory into a buffer area, where the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer area, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers; and After parsing the data loading instruction, using an execution unit to execute the data loading instruction, Among them, using the execution unit to execute the data loading instruction includes: Obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, determining a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and Sequentially sending the plurality of requests, and writing the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area. Among them, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension. The requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory. Among them, the data storage format of the original tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

15. A data storage method for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, wherein, The data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the second tensor are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The shape and size of the first tensor are represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in the 5 dimensions and are all positive integers. The data storage method includes: Obtain the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Combine the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine multiple requests for storing the second tensor. Among them, the multiple requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and sequentially write the data corresponding to each request into the memory to store the second tensor to the memory. Among them, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, it cannot be continuously acquired in the first dimension but can be continuously acquired in dimensions lower than the first dimension, a request is partitioned in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer. Among them, the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

16. A data storage method, comprising: Receiving a data storage instruction indicating to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, where the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape dimensions of the second tensor are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the dimensions of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape dimensions of the first tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the dimensions of the first tensor in the 5 dimensions and are all positive integers; and After parsing the data storage instruction, using an execution unit to execute the data storage instruction, wherein, using the execution unit to execute the data storage instruction includes: Obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, determining a plurality of requests for storing the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and Sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the memory to store the second tensor into the memory. Among them, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be achieved in the first dimension but can be achieved in dimensions lower than the first dimension, a request is divided in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer. Among them, the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

17. A processor, comprising an instruction parsing unit and an execution unit, wherein, The instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the data loading instruction instructs to load a tensor to be processed from an original tensor in a memory into a buffer, wherein the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in the 5 dimensions and are all positive integers; and The execution unit is configured to: execute the data loading instruction, wherein, when the execution unit executes the data loading instruction, it includes: Obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; Combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor, to determine a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and Sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the buffer to load the tensor to be processed into the buffer. Wherein, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, requests are partitioned in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory. Wherein, the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

18. A processor, comprising an instruction parsing unit and an execution unit, wherein, The instruction parsing unit is configured to: receive and parse a data storage instruction, wherein the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into the memory, wherein the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the second tensor are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape and size of the first tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in the 5 dimensions and are all positive integers; and The execution unit is configured to: execute the data storage instruction, Wherein, the execution unit executing the data storage instruction includes: Obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, to determine a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and Sequentially sending the plurality of requests and writing the data corresponding to each request into the memory in sequence to store the second tensor into the memory, Among them, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer. Among them, the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

19. An electronic device, including a processor and a memory connected to the processor, wherein, The processor includes a buffer, where The processor is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor, the data loading method according to any one of claims 1-14 is implemented to load a tensor to be processed from the original tensor in the memory to the buffer, or the data storage method according to claim 15 or 16 is implemented to obtain a second tensor based on the first tensor in the buffer and write the second tensor to the memory.

20. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer-executable instructions are executed by a processor, the data loading method according to any one of claims 1-14 is implemented, or the data storage method according to claim 15 or 16 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN116822612A

  • Data processing method and device, electronic equipment and computer readable storage medium

    CN118152713A