Data loading method, data storage method, processor, electronic equipment and medium

By sequentially obtaining data from the original 5-dimensional tensors of memory and loading it into the cache area or storing it into memory, the problem of low data loading/storing efficiency in the prior art is solved, fast and efficient data processing is achieved, and the performance of computing devices is improved.

CN120216401AActive Publication Date: 2025-06-27SHANGHAI BIREN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510607198.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-27
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to provide fast and efficient data loading/storage methods, which affects the computing performance of parallel processors.

Method used

A data loading method and data storage method are proposed, and the data processing is achieved by obtaining data sequentially from the original 5-dimensional tensors of memory, loading them into a cache area or storing them into memory.

Benefits of technology

It improves memory access bandwidth, improves hardware utilization of computing devices, and is especially suitable for computing operations such as convolutional operations, reduces memory access time and improves the overall performance of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216401A_ABST
    Figure CN120216401A_ABST
Patent Text Reader

Abstract

The invention provides a data loading method, a data storage method, a processor, electronic equipment and a medium. The data loading method is used for loading a to-be-processed tensor to a cache region from an original tensor of a memory, and comprises the following steps: acquiring the number of pixels included in the to-be-processed tensor and an initial coordinate of the to-be-processed tensor in a coordinate system determined by the original tensor; for the data of the original tensor, sequentially acquiring the data from the original tensor by taking the starting coordinate as a starting point until the number of the acquired data is equal to the number of pixels included in the tensor to be processed; and loading the acquired data to a cache region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a medium. Background Art

[0002] A tensor is a data structure of a multi-dimensional array. Tensor operations are widely used in processors such as parallel processors. For example, in the field of deep learning, the dimensions of the input data, the intermediate data processed during the deep learning process, and the output data are flexible and not exact. Therefore, a flexible data form is needed to describe various types of data, and thus the concept of a tensor is generated. In the field of deep learning, all data to be operated on can be stored and exist in the form of a tensor. If the data is not in the form of a tensor, it needs to be first converted into the data structure form of a tensor. As an example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor itself can have any number of dimensions. For example, it can be represented as a 5-dimensional array.

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processing devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of a parallel processor need to process a large amount of data, and this data is generally stored in the storage component of the parallel processor. For example, the storage component can be a high-speed memory. Through a data loading instruction, this data can be loaded from the storage component to the buffer for calculation, and through a data storage instruction, the data in the buffer can be stored to the memory.

[0004] How to provide a fast and efficient data loading / storing method is crucial for the computing performance of the device. Summary of the Invention

[0005] Embodiments of the present disclosure provide a data loading method, a data storage method, a processor, an electronic device, and a medium, which are used to provide a fast and efficient data loading / storing method, improve the memory access bandwidth, and increase the hardware utilization rate of the computing device.

[0006] According to a first aspect of the present disclosure, there is provided a data loading method for loading a tensor to be processed from an original tensor in memory into a buffer. The original tensor is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. For the 5D tensor of the original tensor, the data represented by the height dimension and the width dimension is a pixel, and it accumulates step by step to higher dimensions. The number of pixels is not calculated for the number of channels dimension. The data loading method includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; for the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and loading the obtained data into the buffer.

[0007] According to some embodiments of the present disclosure, multiple tensors are stored in memory, and the original tensor is one of the multiple tensors. The data storage format of the multiple tensors in memory can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. For the NDHWC data storage format, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the number of channels dimension. For the N(C / x)DHW(xC) data storage format, N corresponds to b1, indicating the batch dimension, (C / x) corresponds to b2, indicating the number of channels dimension, D corresponds to b3, indicating the depth dimension, H corresponds to b4, indicating the height dimension, and W(xC) corresponds to b5, indicating the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0008] According to some embodiments of the present disclosure, the tensor to be processed is used for convolution operations in a computing unit within a processor.

[0009] According to some embodiments of the present disclosure, multiple tensors are stored in memory, and the data loading method further includes: obtaining indication information of the storage location of the original tensor in memory. The indication information includes the starting coordinates of the original tensor in memory and the size values in five dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5.

[0010] According to some embodiments of the present disclosure, loading the acquired data into the buffer includes: loading the acquired data into the buffer from the memory according to multiple instructions.

[0011] According to some embodiments of the present disclosure, loading the acquired data into the buffer further includes: determining a first data storage format of the tensor to be processed in the buffer; placing the tensor to be processed in the buffer according to the first data storage format, where the first data storage format of the tensor to be processed in the buffer is the same as or different from the second data storage format of the original tensor in the memory.

[0012] According to some embodiments of the present disclosure, the data loading method further includes: obtaining boundary values for at least a part of five dimensions of the original tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the original tensor, and the boundary values are arbitrary values compared to the size values of the original tensor in this dimension.

[0013] According to some embodiments of the present disclosure, sequentially obtaining data from the original tensor includes: for the data part of the original tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the original tensor defined by the boundary values; and for the data part of the original tensor not covered by the range defined by the boundary values, which is represented as invalid data, for the invalid data, the buffer directly fills the tensor to be processed with a predetermined value, where the predetermined value is equal to 0.

[0014] According to some embodiments of the present disclosure, the data loading method further includes: obtaining a data step value, where the data step value is used to specify the step for sequentially obtaining data from the original tensor. For the data of the original tensor, starting from the starting coordinate, sequentially obtaining data from the original tensor includes: for the data of the original tensor, starting from the starting coordinate, sequentially obtaining data from the original tensor according to the data step value.

[0015] According to a second aspect of the present disclosure, a data loading method is provided, including: receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in memory into a buffer area, where the original tensor is a 5D tensor, and its shape dimensions are represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension; wherein, for the 5D tensor of the original tensor, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step to higher dimensions, where the number of channels dimension does not count the number of pixels; and after parsing the data loading instruction, using an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: obtaining the number of pixels included in the tensor to be processed, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; for the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and loading the obtained data into the buffer area.

[0016] According to a third aspect of the present disclosure, a data storage method is provided, for obtaining a second tensor based on a first tensor in a buffer area and writing the second tensor into memory, where the first tensor is a 5D tensor, and its shape dimensions are represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the dimensions of the first tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension; wherein, for the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step to higher dimensions, where the number of channels dimension does not count the number of pixels. The data storage method according to an embodiment of the present disclosure includes: obtaining the number of pixels included in the second tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; for the data of the first tensor, starting from the starting coordinates, sequentially obtaining data from the first tensor until the number of obtained data is equal to the number of pixels included in the second tensor; and storing the obtained data in memory.

[0017] According to some embodiments of the present disclosure, a plurality of tensors are stored in a buffer. The first tensor is one of the plurality of tensors. The data storage format of the plurality of tensors in the buffer can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensors in the storage component. Among them, for the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension; for the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0018] According to some embodiments of the present disclosure, a plurality of tensors are stored in the buffer. The above data storage method may further include: obtaining indication information of the storage position of the first tensor in the buffer, where the indication information includes the starting coordinates of the first tensor in the buffer and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5.

[0019] According to some embodiments of the present disclosure, storing the obtained data in the memory includes: storing the obtained data from the buffer to the memory according to multiple instructions.

[0020] According to some embodiments of the present disclosure, storing the obtained data in the memory further includes: determining the first data storage format of the second tensor in the memory; placing the second tensor in the memory according to the first data storage format, where the first data storage format of the second tensor in the memory is the same as or different from the second data storage format of the first tensor in the buffer.

[0021] According to some embodiments of the present disclosure, the data storage method further includes: obtaining boundary values for at least a part of the 5 dimensions of the first tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the first tensor, and the boundary values are arbitrary values compared to the size values of the first tensor in this dimension.

[0022] According to some embodiments of the present disclosure, sequentially obtaining data from the first tensor includes: for a data portion of the first tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the first tensor defined by the boundary values; and for a data portion of the first tensor not covered by the range defined by the boundary values, which is represented as invalid data, the invalid data is not stored in the memory.

[0023] According to some embodiments of the present disclosure, the data storage method further includes: obtaining a data step value, where the data step value is used to specify the step size for sequentially obtaining data from the first tensor. For the data of the first tensor, starting from the starting coordinates, sequentially obtaining data from the first tensor includes: for the data of the first tensor, starting from the starting coordinates, sequentially obtaining data from the first tensor according to the data step value.

[0024] According to a fourth aspect of the present disclosure, a data storage method is provided, including: receiving a data storage instruction indicating to execute obtaining a second tensor based on the first tensor in a buffer and writing the second tensor into the memory, where the first tensor is a 5-dimensional tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension; where for the 5-dimensional tensor of the first tensor, the data determined by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, and the number of pixels is not calculated for the number of channels dimension; and after parsing the data storage instruction, using an execution unit to execute the data storage instruction, where using the execution unit to execute the data storage instruction includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; for the data of the first tensor, starting from the starting coordinates, sequentially obtaining data from the first tensor until the number of obtained data is equal to the number of pixels included in the second tensor; and storing the obtained data in the memory.

[0025] According to a fifth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, where the data loading instruction instructs to execute loading a tensor to be processed from an original tensor in a memory to a buffer area. The original tensor is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. For the 5D tensor of the original tensor, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step to higher dimensions. The number of channels dimension does not count the number of pixels. And the execution unit is configured to: execute the data loading instruction, where the execution unit executing the data loading instruction includes: obtaining the number of pixels included in the tensor to be processed, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; for the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and loading the obtained data to the buffer area.

[0026] According to a sixth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer area and writing the second tensor to the memory. The first tensor is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. For the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step to higher dimensions. The number of channels dimension does not count the number of pixels. And the execution unit is configured to: execute the data storage instruction, where the execution unit executing the data storage instruction includes: obtaining the number of pixels included in the second tensor, and the starting coordinates of the second tensor in a coordinate system determined by the first tensor; for the data of the first tensor, starting from the starting coordinates, sequentially obtaining data from the first tensor until the number of obtained data is equal to the number of pixels included in the second tensor; and storing the obtained data to the memory.

[0027] According to a seventh aspect of the present disclosure, an electronic device is provided, including a processor and a memory connected to the processor. The processor includes a buffer area. The processor is configured to run computer-executable instructions, which, when run by the processor, implement a data loading method according to an embodiment of the present disclosure to load a tensor to be processed from an original tensor in the memory into the buffer area, or implement a data storage method according to an embodiment of the present disclosure to obtain a second tensor based on a first tensor in the buffer area and write the second tensor into the memory.

[0028] According to an eighth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement a data loading method according to an embodiment of the present disclosure, or implement a data storage method according to an embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 FIG. shows a schematic structural diagram of a General-Purpose computing on Graphics Processing Unit (GPGPU); Figure 2 FIG. shows a schematic structure of a tensor; Figure 3A FIG. shows a schematic diagram of the data storage format of NDHWC; Figure 3B FIG. shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32; Figure 4 FIG. shows a schematic diagram of obtaining a tensor in the related art; Figure 5 FIG. shows a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure; Figure 6 FIG. shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to an embodiment of the present disclosure; Figure 7A FIG. shows a schematic diagram of obtaining a tensor to be processed in a continuous manner for the data storage format of NDHWC according to an embodiment of the present disclosure; Figure 7BShows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to the data storage format of N(C / x)DHW(xC) in an embodiment of the present disclosure; Figure 8A Shows a schematic diagram of setting boundary values for an original tensor according to an embodiment of the present disclosure; Figure 8B Shows a schematic diagram of continuous data acquisition in the case of setting boundaries for an original tensor according to an embodiment of the present disclosure; Figure 9 Shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to the set step size according to an embodiment of the present disclosure; Figure 10 Shows a schematic flowchart of a data loading method provided by at least one embodiment of the present disclosure; Figure 11 Shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure; Figure 12 Shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure; Figure 13 Shows a schematic block diagram of a processor according to some embodiments of the present disclosure; Figure 14 Shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure; Figure 15 Shows a block diagram of an example computing device implementing some embodiments of the present disclosure; Figure 16 Shows a schematic block diagram of a computer-readable storage medium according to some embodiments of the present disclosure. Detailed implementation manners

[0031] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, rather than all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0032] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Words such as "upper", "lower", "left", "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of this disclosure clear and concise, some detailed descriptions of known functions and known components are omitted in this disclosure.

[0033] Figure 1 A schematic structural diagram of a GPGPU is shown. As Figure 1 shown, a GPGPU is actually an array of programmable multi-processors. For example, the programmable multi-processor can be a Streaming Processor Cluster (SPC for short), for example, including Figure 1 the streaming processor cluster 1 shown, ..., the streaming processor cluster M, where M is a positive integer greater than 1. In a general-purpose graphics processor, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. As an example, data sharing between multiple streaming processor clusters is performed through a global cache or High Bandwidth Memory (HBM).

[0034] As Figure 1 shown, taking the streaming processor cluster 1 as an example, 1 streaming processor cluster may include multiple Compute Units (CUs for short), for example Figure 1 the compute unit 1, compute unit 2, ..., compute unit K in it, where K is a positive integer. Each compute unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A compute unit may include multiple cores (also called computing cores or calculation cores), and each computing core includes an Arithmetic Logic Unit (ALU), a floating-point computing unit, etc., and the computing core is used to perform specific computing tasks. In addition, the compute unit also includes registers (for example Figure 1The register bank) and shared memory are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in one computing unit is used to share data among the cores of this computing unit. In addition, the buffer can be understood as being used to share data among the computing units within the streaming processor cluster.

[0035] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processing unit (or parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0036] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates the thread bundles so that multiple computing cores in this computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the middle-level cache or global cache or high-bandwidth memory for read / write operations, etc.

[0037] As Figure 1 shown, general computing operations, such as matrix computing operations in the field of artificial intelligence, usually require a large amount of data. These data are usually stored in a memory, such as in a high-bandwidth memory HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing results, data needs to be stored in the memory (Store operation). The storage method of data in the memory will affect the memory access bandwidth, and thus affect the hardware utilization rate of the computing unit.

[0038] For example, general computing operations include General Matrix Multiplication (abbreviated as GEMM). As an example, the data required for general matrix multiplication is represented as two 5D arrays, such as the matrix multiplication calculation of two tensors A and tensor B. In addition, general computing operations also include convolution operations, which are manifested as data dot products. It can be understood that in the field of artificial intelligence, there are also other computing operations, which will not be listed one by one here. The data involved in these computations usually appears in the form of tensors.

[0039] For example, for a certain tensor A in the buffer, its shape and size can be represented by a1, a2, a3, a4, and a5. a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor data in 5 dimensions, and a1, a2, a3, a4, and a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, N represents the batch dimension, which is the number of data samples captured in one training. The D dimension represents the depth dimension, the H dimension represents the height dimension of the input data, the W dimension represents the width dimension of the input data, and the C dimension represents the number of channels dimension. For example, taking tensor A as an example, a1 can be the size of the N dimension, a2 can be the size of the D dimension, a3 can be the size of the H dimension, a4 can be the size of the W dimension, and a5 can be the size of the C dimension. Of course, the present disclosure does not make specific limitations on this.

[0040] As an example, Figure 2 shows a schematic structure of a tensor. In Figure 2 the shown tensor, a1 is the size of the N dimension and is equal to 1, a2 is the size of the D dimension and is equal to 1, a3 is the size of the H dimension and is equal to 5, a4 is the size of the W dimension and is equal to 4, and a5 is the size of the C dimension and is equal to 64. For example, Figure 2 the pixel elements of the tensor in

[0041] are represented as 0, 1, 2, 3,... and so on. Figure 2 The placement of the tensor in the memory (such as memory or buffer) can have multiple formats, called data storage format (layout). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The following uses the

[0042] shown tensor to describe different data storage formats. Figure 3A shows a schematic diagram of the NDHWC data storage format.

[0043] For example, for the NDHWC linear mode, as Figure 3A shown, starting from the first element (element 0 in Figure 3A ) of the first channel (a5 = 0, c0 in Figure 3A ), then storing the first element (element 20 in Figure 3A ) of the second channel (a5 = 1, c1 in Figure 3A ), and so on, until all the first elements of all channels are laid out, for example, until the first element of the 64th channel (a5 = 63, c63 in Figure 3A ) (element in Figure 3AAfter the element 1260) in, select the first channel (a5 = 0, Figure 3A the second element of c0) in Figure 3A element 1) in, and then store the second element of the second channel (a5 = 1, Figure 3A c1) in Figure 3A element 21) in, and so on until the second elements of all channels are laid out, and so on.

[0044] In the related art, the data storage format may also include N(C / x)DHW(xC), also known as the Interleave mode, where x can be set to 8, 16, 32, etc. as needed.

[0045] The N(C / x)DHW(xC) data storage format is similar to the NDHWC data storage format, but a key difference is that in the layout of N(C / x)DHW(xC), a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.

[0046] Figure 3B Shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32.

[0047] As Figure 3B shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 3B c0) in to a5 = 31 ( Figure 3B c31) in, and the second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format.

[0048] In memory, for example, a certain tensor B, similar to a certain tensor A in the buffer, this tensor B can be stored in memory according to one of the two data storage formats described above, and the shape dimensions of the tensor can be similarly expressed as b1×b2×b3×b4×b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the tensor B in these 5 dimensions and are all positive integers.

[0049] It can be understood that in the related art and future possible developments, the data storage format of tensors is not limited to the above-described N(C / x)DHW(xC) data storage format and NDHWC data storage format, and the method described in this disclosure does not limit this. Further, for multiple tensors that can be stored in memory and the buffer during the calculation process, usually the storage space of memory is much larger than that of the buffer, but it is farther from the calculation unit, and the data transfer efficiency is lower than that of the buffer.

[0050] In the related art, during the computing process of a processing device, a large amount of computing data will be generated, for example, in the form of tensors, which can be temporarily stored in a buffer. For example, the buffer here can refer to Figure 1 the buffer in the streaming processor cluster shown in []. Further, this data can also be transferred from the buffer or directly stored in the memory. For example, the memory can be Figure 1 the high-bandwidth memory HBM shown in []. The storage forms of tensors in the memory and the buffer can both be any one of the N(C / x)DHW(xC) data storage format and the NDHWC data storage format described above. Thus, during the computing process, according to factors such as technical requirements and the respective storage characteristics of the buffer and the memory, a large amount of data transfer processes need to be performed between the two. For example, tensors in the buffer are stored in the memory through storage instructions, or tensors in the memory are loaded into the buffer through load instructions.

[0051] It can be understood that in this article, the data storage process of storing tensors in the buffer into the memory and the data loading process of loading tensors in the memory into the buffer can be implemented in a similar manner. Therefore, for the sake of convenience in description, in some embodiments or examples, only the data loading process is described as an example, and those skilled in the art can apply it similarly to the data storage process. For the differences between the two, additional descriptions will be made.

[0052] As an example, Figure 4 shows a schematic diagram of obtaining a tensor in the related art. As Figure 4 shown, tensor A is stored in the memory, and it can be placed in the memory in the above N(C / x)DHW(xC) or NDHWC data storage format. That is, tensor A is a 5D array. In Figure 4 only three dimensions, W, H, and C, are schematically shown. Schematically, on the left side of Figure 4 is a dimensional coordinate system of tensor A, where the W dimension, the H dimension, and the C dimension are respectively shown. During the data loading process, all or part of the data in tensor A can be loaded into the buffer through a load instruction. For example, tensor B in Figure 4 is loaded into the buffer as a whole. Specifically, the load instruction can indicate the first starting point (C = 0, W = 0, H = 0) of tensor A to be loaded, the second starting point (C = 0, W = 3, H = 0) of tensor B in the coordinate system of tensor A, and the sizes of tensor B in each dimension. Schematically, in Figure 4In the 3D schematic diagram shown, the tensor B is a cuboid determined by the above-mentioned second starting point and each dimension size. Based on the information about the second starting point and each dimension size of the tensor to be loaded, the memory can load the tensor B into the buffer area.

[0053] In the above related technologies, the block-based data loading method can be applied to general matrix multiplication (GEMM) calculations in general computing operations. However, the above block-based data loading method is not applicable to convolution operations, whose computational characteristics are data dot products. The above block-based data loading method will limit the computational efficiency of such per-pixel type convolution operations, increase data access time, and reduce the overall performance of the processor, which limits the further development space of efficient and general-purpose processors.

[0054] To address the above technical problems in the related technologies, the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. In the present disclosure, a new data transfer mode is proposed. The tensor is no longer obtained and transferred in a block form, but is obtained and transferred sequentially in units of pixels to be applicable to computational operations such as convolution operations, improving the data transfer efficiency for such operations, reducing memory access time, and enhancing the overall performance of the processor.

[0055] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It can be understood that the present disclosure is not limited to these specific embodiments.

[0056] The data loading method provided by at least one embodiment of the present disclosure is used to load a tensor to be processed from the original tensor in the memory into the buffer area. The memory according to the embodiments of the present disclosure can be, for example, a high-bandwidth memory (HBM, High Bandwidth Memory), and the buffer area is, for example, the buffer in the streaming processor cluster, which is not limited herein.

[0057] As an example, the original tensor in the memory is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in five dimensions and are all positive integers. The five dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. As an example, the data storage format of the original tensor in the memory can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. For example, the storage component refers to the above-mentioned HBM or the buffer area.

[0058] For the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension.

[0059] For the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension. In general cases, x can be set to an integer multiple of 4.

[0060] Regarding the features of the above two data storage formats, reference can be made to the description above in combination with Figure 3A - Figure 3B and will not be repeated here.

[0061] Furthermore, for the 5D data of the original tensor, the data determined by its height dimension and width dimension is represented as a pixel (or, can also be called an element), and accumulates step by step towards the higher dimension. Among them, the number of pixels is not calculated for the number of channels dimension. As an example, as Figure 4 shown in, the values of each W dimension and H dimension can determine a pixel. For example, W = 0 and H = 0 correspond to the first pixel (or element) in tensor A, and W = 1 and H = 0 correspond to the second pixel in tensor A. In the tensor, the C dimension does not affect the number of pixels.

[0062] Figure 5 is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 5 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S101 - S103.

[0063] In step S101, obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Then, in step S102, for the data of the original tensor, starting from the starting coordinates, sequentially obtain the data from the original tensor until the number of obtained data is equal to the number of pixels included in the tensor to be processed. In step S103, load the obtained data into the buffer area. According to the embodiment of the present disclosure, the tensor to be processed can be used for convolution operations in the computing units within the processor.

[0064] In the data loading method according to the embodiment of the present disclosure, in order to obtain the tensor to be processed from the original tensor in the memory, it is necessary to indicate the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. For example, the starting coordinates can be combined withFigure 4 The described second starting point (C = 0, W = 3, H = 0). Further, in the data loading method according to an embodiment of the present disclosure, it is also necessary to indicate the number of pixels included in the tensor to be processed (for example, denoted as copy_pixel_num), that is, the total number of pixels in the tensor to be loaded into the buffer. This is different from the block-based data loading method in the related art above (where it is necessary to indicate the sizes of the tensor to be processed in each dimension). According to the data loading method of the embodiment of the present disclosure, the range of the tensor to be processed is determined by the number of pixels to be obtained. Further, instead of obtaining the data in a block from the original tensor, the data is sequentially obtained from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed. The specific implementation details of realizing data loading based on copy_pixel_num and the starting coordinates will be elaborated in detail below.

[0065] As an example, Figure 6 shows a schematic diagram of obtaining the tensor to be processed in a continuous manner according to an embodiment of the present disclosure. In Figure 6 , tensor A represents the original tensor in memory. In Figure 6 , each pixel point is shown as a square. Tensor A covers all the squares, that is, both the white squares and the gray-shaded squares. The starting point of this original tensor is denoted as (0, 0, 0), that is, it represents C = 0, W = 0, H = 0, for example, used to indicate the position of this original tensor in memory. Then, in Figure 6 , tensor B represents the tensor to be obtained. Its starting point is denoted as (C = 0, W = 3, H = 0), that is, data is obtained starting from the 3rd pixel point in the first row of tensor A. Then, in Figure 6 schematically shows the process of obtaining the tensor to be processed in a sequential (or also called continuous) manner according to an embodiment of the present disclosure, that is, starting from the starting point (C = 0, W = 3, H = 0) as shown by the dashed arrow, data is obtained from tensor A pixel by pixel until the number of obtained pixels reaches copy_pixel_num. In the example of Figure 6 , only the three dimensions of H, W, and C are shown. It can be understood that if the data in these 3 dimensions still does not reach copy_pixel_num, data in higher dimensions can be further obtained, which is not limited here. In addition, as shown in Figure 6 , the C dimension itself does not affect the number of pixels. Assume that Figure 6 the data storage format of tensor A in memory is the NDHWC format. In the example shown in Figure 6 , the total number of pixels of the tensor to be processed is copy_pixel_num = 27.

[0066] In an embodiment according to the present disclosure, in step S102, for the data of the original tensor, sequentially obtaining data from the original tensor starting from the starting coordinates includes: excluding the channel number dimension from the five dimensions of the original tensor, and in the order of dimensions from low to high, using the starting coordinates as the starting point for data acquisition, and sequentially obtaining data until the number of pixels obtained reaches copy_pixel_num. The reason for excluding the channel number dimension from the five dimensions is that for a tensor, the channel number dimension does not count the number of pixels. That is to say, S102 may include obtaining pixel data from the original tensor with the starting coordinates as the starting point for data acquisition in the order of the width dimension, height dimension, depth dimension, and batch dimension until the number of pixels obtained reaches copy_pixel_num.

[0067] Compare Figure 4 and Figure 6 with the two ways of obtaining the tensors to be processed shown, the data loading method provided by the embodiments of the present disclosure can achieve pixel-by-pixel data loading, rather than Figure 4 the block-by-block acquisition in, this data loading method is more conducive to the calculation process such as convolution operations and is conducive to improving the operation efficiency. Thus, based on the method provided by the embodiments of the present disclosure, for the data to be subjected to convolution operations next, the processor can, for example, indicate in the form of an instruction to fetch this part of the data from the memory to the buffer area according to the above continuous data loading method for use in convolution operations. It can be understood that the above processor can be reasonably used according to the type of operation to be performed or the data processing characteristics, whether it is the continuous data loading method or the block-by-block data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer area may also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0068] According to some embodiments of the present disclosure, there are multiple tensors stored in the memory, and the data loading method further includes: obtaining indication information of the storage location of the original tensor in the memory, where the indication information includes the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. It can be understood that the storage space of the memory is generally much larger than the buffer, and various tensor data generated during the calculation process can be stored therein. These tensor data can be placed in the memory according to any one of the above N(C / x)DHW(xC) or NDHWC data storage formats. Based on step S101 of the embodiments of the present disclosure, in order to enable the memory to know the specific location of the tensor to be obtained in the memory, indication information of the storage location of the original tensor in the memory can also be obtained. The indication information includes the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. As an example, in the case where the original tensor is in the NDHWC data storage format, tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5 respectively represent the values of the original tensor in the 5 dimensions of N, D, H, W, and C.

[0069] According to some embodiments of the present disclosure, loading the obtained data into the buffer includes: loading the obtained data into the buffer from the memory according to multiple instructions. It can be understood that for the tensor B shown in Figure 6 it can be loaded into the buffer through multiple instructions. For example, it can be split into multiple requests, which is not limited herein.

[0070] According to some embodiments of the present disclosure, loading the acquired data into the buffer further includes: determining a first data storage format of the tensor to be processed in the buffer; and placing the tensor to be processed in the buffer according to the first data storage format, wherein the first data storage format of the tensor to be processed in the buffer is the same as or different from the second data storage format of the original tensor in the memory. Before storing tensor B into the buffer, the method according to the embodiments of the present disclosure can also indicate the data placement manner thereof in the buffer, that is, whether it is placed according to the NDHWC data storage format or according to the N(C / x)DHW(xC) data storage format. In the method according to the embodiments of the present disclosure, there is no limitation on the placement manner of the tensor to be processed in the buffer. It can be NDHWC or N(C / x)DHW(xC). It can be the same as or different from the placement manner of the original tensor in the memory. The data loading method of the embodiments of the present disclosure does not require the data placement manner. The difference is only that different placement manners may affect the pixel reading order of tensor B and the specific content corresponding to multiple instructions required when loading tensor B into the buffer.

[0071] Regarding the influence of the data storage format on the loading process of the tensor to be processed (“tensor B”), it will be described below in conjunction with Figure 7A - Figure 7B which will be Figure 7A and Figure 7B In the example shown, the data storage manner of the tensor to be processed in the buffer is the same as the data placement manner of the original tensor in the memory. It can be understood that the method according to the embodiments of the present disclosure can also be applicable to the case where the two are different.

[0072] Figure 7A FIG. shows a schematic diagram of continuously acquiring the tensor to be processed according to the NDHWC data storage format according to the embodiments of the present disclosure. For the Figure 7A tensor shown, its data storage format in the memory is in the form of NDHWC, or is placed according to the NDHWC placement manner.

[0073] In Figure 7A the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 8 (the C dimension is not shown in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Taking the N dimension as an example, N_dim = 2 corresponds to Figure 7A N0 and N1 in

[0074] Next, in Figure 7AIn the example, the coordinates of the starting point of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixels included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7A The pixel point where W = 2 and H = 2 in the dimensions of N = 0 and D = 0 shown in Figure 7A , starting from this pixel point, according to the data arrangement method, in the order from low dimension to high dimension (that is, in the order of W, H, D, N, as shown by the data acquisition schematic arrow in Figure 7A ), the data is sequentially acquired until the number of acquired data is equal to the number of pixels included in the tensor to be processed. In the example shown in

[0075] Figure 7B shows a schematic diagram of continuously acquiring a tensor to be processed for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure. For Figure 7B the shown tensor, its data storage format in memory is in the form of N(C / x)DHW(xC), or is arranged in the manner of N(C / x)DHW(xC). In the example shown in Figure 7B , x = 8, and every 8 C channels are grouped and bound to the W dimension. Compared with the NDHWC data storage format shown in Figure 7A , for the data storage format of N(C / x)DHW(xC), the level of the C dimension in the tensor is higher and is located between the N dimension and the D dimension.

[0076] In Figure 7B the shown example, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 16 (shown as two 8*C dimensions in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Thus, it can be obtained that the total number of pixel points included in the original tensor is 64.

[0077] Next, in Figure 7BIn the example, the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixels included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7B In the pixel point where W = 2 and H = 2 under N = 0 and D = 0 and the first 8*C dimension shown in, starting from this pixel point, according to the data arrangement method, in the order from low dimension to high dimension (i.e., in the order of W, H, D, C, N, as shown by the data acquisition schematic arrow in Figure 7B , data is sequentially acquired until the number of acquired data is equal to the number of pixels included in the tensor to be processed. Compared with Figure 7A the acquisition order shown in, since the rank of the C dimension is higher, in Figure 7B the example of, first, data of W, H, D, and the first 8*C dimension is acquired according to the starting point, then, data of the second 8*C dimension is acquired in the order of W, H, D, and finally, data of the N dimension is acquired. It can be understood that in Figure 7A and Figure 7B the examples of, the number of pixels of the tensor to be processed acquired is 40, and the difference is only caused by the different arrangement methods of the C dimension. In addition, it should be noted that for Figure 7B the second 8*C dimension shown in, its starting point at the W and H dimensions should be aligned with the starting point of the first 8*C dimension.

[0078] In Figure 7B the example shown in, starting from the determined starting point, data is acquired pixel by pixel in the arrow order, so as to load this part of data in the original tensor in the memory into the buffer area as the tensor to be processed for subsequent operations such as convolution calculation.

[0079] According to some embodiments of the present disclosure, the data loading method further includes: obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the original tensor, and the boundary values are arbitrary values compared with the size values of the original tensor in this dimension. As an example, left and right boundary values can be set for, for example, the C dimension, W dimension, H temperature, and D dimension, for example, represented as (L-tensor_C, R-tensor_C), (L-tensor_W, R-tensor_W), (L-tensor_H, R-tensor_H), (L-tensor_D, R-tensor_D).

[0080] In the method described in combination with Figure 7A and Figure 7B , the range (i.e., boundary values) of the original tensor is not limited, that is, the tensor to be processed is obtained from the complete data range of the original tensor. In the method according to an embodiment of the present disclosure, it is also proposed that the range for obtaining the tensor to be processed from the original tensor can be delimited by setting boundary values. Further, in the implementation process, the boundary value can be set to any value compared to the size value of the original tensor in this dimension. That is to say, the boundary value can exceed the range of the original tensor itself.

[0081] As an example, Figure 8A shows a schematic diagram of setting boundary values for the original tensor according to an embodiment of the present disclosure. As Figure 8A shown, the rectangular box represents the range covered by the original tensor, which can be any dimension in the original tensor, such as the C dimension, the W dimension, the H dimension, or the D dimension. Generally, boundary values are not set for the N dimension. In Figure 8A 's six sub - pictures, the relationship between the left boundary value and the right boundary value ( Figure 8A shown as "L" and "R" in

[0082] and the size value of the original tensor in this dimension is shown respectively. According to an embodiment of the present disclosure, when setting boundary values, obtaining data sequentially from the original tensor includes: for the range defined by the boundary values that covers the data part of the original tensor, obtaining data sequentially from the range of the original tensor defined by the boundary values, for example, referring to the order described in Figure 7A and Figure 7B . In comparison, for the data part of the original tensor not covered by the range defined by the boundary values, it is represented as invalid data. For invalid data, the buffer directly fills the tensor to be processed with a predetermined value, where the predetermined value is equal to 0.

[0083] For example, in Figure 8A 's first sub - picture, both the left boundary value L and the right boundary value R are on the left side of the original tensor, that is, all the data to be obtained in this dimension are invalid data. In this case, for invalid data, this part of the data can be automatically filled, for example, by sending a zero - filling instruction to the buffer. Again, for example, in Figure 8A 's second sub - picture, the left boundary value L is on the left side of the left boundary of the original tensor data range, and the right boundary value R is on the left side of the right boundary of the original tensor, that is, for the tensor to be processed to be obtained, part is invalid data and part is valid data in the original tensor. Schematically, in Figure 8A , the data corresponding to the slanted shaded part is represented as valid data, and the rest are all invalid data. In this case, for valid data, for example, referring to Figure 7A and Figure 7BPerform in the described order. For invalid data, for example, this part of the data can be automatically filled by sending a zero-padding instruction to the buffer area.

[0084] In the method according to an embodiment of the present disclosure, by setting boundary values for the original tensor, the range from which the tensor to be processed is to be taken can be further delimited. In practical applications, this implementation method can adapt to the characteristics of operations such as convolution. For example, it is beneficial to reduce the computational amount and greatly improve the flexibility of data. As an example, assume that the original tensor corresponds to an intermediate tensor for feature extraction of an entire input image, and the input image includes a specific target, such as an object to be recognized, and the object does not cover the entire image, that is, the image includes a background part. In this case, by setting boundary values, the range of the tensor to be processed to be obtained can be limited to the part of the original tensor corresponding to the specific target, so as to reduce the computational amount of subsequent operations such as convolution and improve the processing efficiency.

[0085] As an example, Figure 8B shows a schematic diagram of continuous data acquisition in the case where boundaries are set for the original tensor. Among them, compared with Figure 6 the situation shown in Figure 8B it can be understood that boundary values (bound) are respectively set for the C dimension, W dimension, and H dimension of the tensor A in Figure 6 , and the boundaries are all within the size ranges of the tensor A in the C dimension, W dimension, and H dimension, which is equivalent to the boundary situation shown in the fourth sub-image in Figure 8A . Specifically, Figure 8B the outer square in Figure 6 corresponds to the tensor A (corresponding to the tensor A in Figure 8B . After setting boundaries therein, according to the data loading method of the embodiment of the present disclosure, data will be sequentially acquired starting from the starting point coordinates within the set boundary range until the number of acquired data is equal to the number of pixels included in the tensor to be processed. In Figure 8B , the data part composed of squares corresponds to the data range delimited by the boundary values set for the C dimension, W dimension, and H dimension, and data is sequentially acquired from within the data range delimited by the boundary. Data located outside the data range delimited by the boundary can be regarded as invalid data. For the process of sequentially acquiring data in Figure 6 , reference can be made to the description in combination with

[0086] According to some embodiments of the present disclosure, the data loading method further includes: obtaining a data stride value, where the data stride value is used to define the stride for sequentially obtaining data from the original tensor. For the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor includes: for the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor according to the data stride value.

[0087] Figure 9 FIG. shows a schematic diagram of continuously obtaining a tensor to be processed according to the set stride according to an embodiment of the present disclosure. Wherein, the stride is equal to 2, that is, one data is taken every other pixel point, that is, only the pixels shown in the shaded part are sequentially obtained. The stride equal to 2 means skipping one pixel. Similarly, when the stride is equal to 3, it means skipping two pixels, and so on. It can be understood that when the set stride is equal to 1, it corresponds to Figure 7A - Figure 7B the data loading method of each pixel point shown. In practical applications, by setting the stride, the computational amount of subsequent operations such as convolution operations on the data can be further reduced, and the processing efficiency can be improved. In addition, the flexibility of data loading can be further improved.

[0088] According to the data loading method provided by the embodiments of the present disclosure, a new data acquisition mode different from the block-based data acquisition method is provided, that is, data loading of each pixel point can be realized according to the starting point and the number of pixels to be acquired (as Figure 6 shown), rather than Figure 4 the block-based acquisition in. This data loading method is more conducive to the calculation process of operations such as convolution operations and is beneficial to improving the operation efficiency. Specifically, the tensor is no longer acquired and transported in a block form, but is sequentially acquired and transported in units of pixel points to be applicable to calculation operations such as convolution operations, improving the data transportation efficiency for such operations, reducing the memory access time, and improving the overall performance of the processor.

[0089] According to some embodiments of the present disclosure, another data loading method is provided. Figure 10 FIG. shows a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 10 shown, the data loading method according to the embodiments of the present disclosure includes step S201 and step S202.

[0090] In step S201: Receive a data loading instruction indicating to execute loading a tensor to be processed from the original tensor in the memory to the buffer area.

[0091] According to an embodiment of the present disclosure, the original tensor is a 5D tensor, and its shape dimensions are represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in five dimensions and are all positive integers. The five dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. Among them, for the 5D tensor of the original tensor, the data determined by the height dimension and the width dimension represents one pixel, and accumulates gradually to higher dimensions. Among them, the number of channels dimension does not calculate the number of pixels.

[0092] In step S202: After parsing the data loading instruction, the execution unit is used to execute the data loading instruction.

[0093] As Figure 10 shown, where step S202 uses the execution unit to execute the data loading instruction, including: S2021: Obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; S2022: For the data of the original tensor, starting from the starting coordinates, sequentially obtain data from the original tensor until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and S2023: Load the obtained data into the buffer.

[0094] Regarding the related descriptions of the tensor to be processed and the original tensor and the specific implementation processes of steps S2021 - S2023, reference can be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.

[0095] As an example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction is implemented in the form of a Load instruction.

[0096] According to some embodiments of the present disclosure, a data storage method is further provided for obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into the memory. It can be understood that the data storage method according to the embodiments of the present disclosure can be understood as the reverse process of the data loading method described above, that is, moving the data from the buffer to the memory. The implementation principle according to the embodiments of the present disclosure is similar to the above data loading method, and the repeated parts will not be described again, and only the different parts will be described in detail.

[0097] In the data storage method according to an embodiment of the present disclosure, the first tensor is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, and b5. b1, b2, b3, b4, and b5 respectively indicate the sizes of the first tensor in five dimensions and are all positive integers. The five dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. Among them, for the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and it accumulates gradually towards higher dimensions. Among them, the number of pixels is not calculated in the number of channels dimension. As an example, the first tensor may refer to the tensor stored in the buffer area, and it may be any one of the data storage formats of NDHWC or N(C / x)DHW(xC), and there is no limitation on this. In terms of understanding the implementation principle, it can be correspondingly understood as the original tensor in the data loading method described above, that is, part or all of the data is obtained from the first tensor and transferred and stored in the memory. This part of the tensor obtained from the first tensor is represented as the second tensor, which can be correspondingly understood as the tensor to be processed in the data loading method described above.

[0098] Figure 11 FIG. shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure, as Figure 11 shown, the data storage method includes steps S301-S303.

[0099] In step S301, obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor. Then, in step S302, for the data of the first tensor, starting from the starting coordinates, sequentially obtain data from the first tensor until the number of obtained data is equal to the number of pixels included in the second tensor. In step S303, store the obtained data in the memory.

[0100] In the data storage method according to an embodiment of the present disclosure, the manner of sequentially obtaining data from the first tensor can be referred to and combined with Figure 6 the description manner shown. Compared with the block-based data storage method in the related art, the data storage method provided by the embodiment of the present disclosure can achieve pixel-by-pixel data storage, rather than Figure 4 the block-based acquisition in. This data storage method is more conducive to the calculation process such as convolution operation and is beneficial to improving the operation efficiency. It can be understood that, for example, the processor can reasonably use according to the type of operation to be performed or the data processing characteristics, whether it is a continuous data loading method or a block-based data loading method, that is, it can support the adaptive switching between these two loading methods, and no further expansion is made here. In addition, the memory or the buffer area may also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0101] In practical applications, for example, in the convolution operation of general computing operations, the operation process is multi-layered. For example, after the first layer is processed, the data result needs to be stored for use in the next layer processing. As described above, for the data that needs to perform convolution operations, adopting the sequential data acquisition method provided by the embodiments of the present disclosure (which can refer to the description made in conjunction with Figure 5 、 Figure 6 、 Figure 7A 、 Figure 7B 、 Figure 8A and Figure 8B ) is more conducive to improving the computing efficiency. As an example, the weight data of 1x3 can only obtain data in the form of a sliding window, so it can only be obtained row by row, unless the shape of the graph to be obtained is very regular, exactly a cube and the next layer does not require operations such as padding.

[0102] According to some embodiments of the present disclosure, multiple tensors are stored in the buffer. The first tensor is one of the multiple tensors. The data storage format of the multiple tensors in the buffer can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensors in the storage component. Among them, for the NDHWC data storage format, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the channel number dimension; for the N(C / x)DHW(xC) data storage format, N corresponds to b1, indicating the batch dimension, (C / x) corresponds to b2, indicating the channel number dimension, D corresponds to b3, indicating the depth dimension, H corresponds to b4, indicating the height dimension, and W(xC) corresponds to b5, indicating the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the channel number of xC is bound to the width dimension.

[0103] According to some embodiments of the present disclosure, multiple tensors are stored in the buffer. The above data storage method may further include: obtaining indication information of the storage position of the first tensor in the buffer. The indication information includes the starting coordinates of the first tensor in the buffer and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5.

[0104] According to some embodiments of the present disclosure, storing the obtained data in the memory includes: storing the obtained data in the memory from the buffer according to multiple instructions.

[0105] According to some embodiments of the present disclosure, storing the acquired data in the memory further includes: determining a first data storage format of the second tensor in the memory; placing the second tensor in the memory according to the first data storage format, wherein the first data storage format of the second tensor in the memory is the same as or different from the second data storage format of the first tensor in the buffer.

[0106] According to some embodiments of the present disclosure, the data storage method further includes: obtaining boundary values for at least a part of the five dimensions of the first tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the first tensor, and where the boundary values are arbitrary values compared to the size values of the first tensor in that dimension.

[0107] According to some embodiments of the present disclosure, sequentially obtaining data from the first tensor includes: for the data part of the first tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the first tensor defined by the boundary values; and for the data part of the first tensor not covered by the range defined by the boundary values, which is represented as invalid data, the invalid data is not stored in the memory.

[0108] For invalid data beyond the range, the processing method of the data storage method is different from the above-mentioned data loading method according to the embodiments of the present disclosure. For the data beyond the range of the first tensor in the buffer, it can be understood that it is meaningless for storage in the memory, and thus, this part of the data can be directly not stored. In contrast, in the data loading method according to the embodiments of the present disclosure, for invalid data beyond the boundary of the original tensor, it can be directly filled by the buffer, for example, filled with a predetermined value. As an example, the predetermined value can be the numerical value 0.

[0109] According to some embodiments of the present disclosure, the data storage method further includes: obtaining a data step value, where the data step value is used to specify the step for sequentially obtaining data from the first tensor, and where for the data of the first tensor, starting from the starting coordinate, sequentially obtaining data from the first tensor includes: for the data of the first tensor, starting from the starting coordinate, sequentially obtaining data from the first tensor according to the data step value.

[0110] It can be understood that the data storage method according to the embodiments of the present disclosure can achieve similar technical effects to the data loading method according to the embodiments of the present disclosure.

[0111] According to some embodiments of the present disclosure, another data storage method is provided. Figure 12 The schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown. As Figure 12 shown, the data storage method according to the embodiments of the present disclosure includes step S401 and step S402.

[0112] In step S401: Receive a data storage instruction instructing to execute obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into memory.

[0113] According to an embodiment of the present disclosure, the first tensor is a 5D tensor, and its shape dimensions are represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the dimensions of the first tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. Among them, for the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step to higher dimensions. Among them, the number of channels dimension does not calculate the number of pixels.

[0114] In step S402: After parsing the data storage instruction, use an execution unit to execute the data storage instruction.

[0115] As Figure 12 shown, wherein step S402 uses an execution unit to execute the data storage instruction, including: S4021: Obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; S4022: For the data of the first tensor, starting from the starting coordinates, sequentially obtain data from the first tensor until the number of obtained data is equal to the number of pixels included in the second tensor; and S4023: Store the obtained data into memory.

[0116] Regarding the related descriptions of the first tensor and the second tensor and the specific implementation processes of steps S4021 - S4023, reference can be made to the related descriptions of the foregoing data storage method, and repeated parts will not be elaborated.

[0117] As an example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction is implemented in the form of a Store instruction.

[0118] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, where the data loading instruction instructs to load a tensor to be processed from an original tensor in a memory to a buffer area. The original tensor is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. For the 5D tensor of the original tensor, the data represented by the height dimension and the width dimension is a pixel, and it accumulates step by step to higher dimensions. The number of channels dimension does not calculate the number of pixels. The execution unit is configured to: execute the data loading instruction, where the execution unit executing the data loading instruction includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; for the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and loading the obtained data to the buffer area.

[0119] For example, there are multiple tensors stored in the memory, and the original tensor is one of the multiple tensors. The data storage format of the multiple tensors in the memory can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. For the NDHWC data storage format, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the number of channels dimension. For the N(C / x)DHW(xC) data storage format, N corresponds to b1, indicating the batch dimension, (C / x) corresponds to b2, indicating the number of channels dimension, D corresponds to b3, indicating the depth dimension, H corresponds to b4, indicating the height dimension, and W(xC) corresponds to b5, indicating the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension. Generally, x can be set to an integer multiple of 4.

[0120] For example, the tensor to be processed is used for convolution operations in a computing unit within the processor. It can be understood that the above processor can be reasonably used according to the type of operation to be performed or the data processing characteristics, whether it is a continuous data loading method or a block data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or the buffer area can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0121] For example, when multiple tensors are stored in memory, the execution unit executing the data loading instruction further includes: obtaining indication information of the storage location of the original tensor in memory, where the indication information includes the starting coordinates of the original tensor in memory and the size values in five dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, and tensor_b5.

[0122] For example, the execution unit loading the obtained data into the buffer includes: loading the obtained data from memory into the buffer according to multiple instructions.

[0123] For example, the execution unit loading the obtained data into the buffer further includes: determining a first data storage format of the tensor to be processed in the buffer; and placing the tensor to be processed in the buffer according to the first data storage format, where the first data storage format of the tensor to be processed in the buffer is the same as or different from the second data storage format of the original tensor in memory.

[0124] For example, the execution unit executing the data loading instruction further includes: obtaining boundary values for at least a part of the five dimensions of the original tensor, where the boundary values are used to define the boundary range for obtaining data from the original tensor, and the boundary values are arbitrary values compared to the size values of the original tensor in that dimension.

[0125] For example, the execution unit sequentially obtaining data from the original tensor includes: for the data part of the original tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the original tensor defined by the boundary values; and for the data part of the original tensor not covered by the range defined by the boundary values, which is represented as invalid data, for the invalid data, the buffer directly fills the tensor to be processed with a predetermined value, where the predetermined value is equal to 0.

[0126] For example, the execution unit executing the data loading instruction further includes: obtaining a data step value, where the data step value is used to specify the step for sequentially obtaining data from the original tensor. Specifically, the execution unit executing the data for the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor includes: for the data of the original tensor, starting from the starting coordinates, sequentially obtaining data from the original tensor according to the data step value.

[0127] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory. The first tensor is a 5D tensor, and its shape size is represented by five parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. For the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and accumulates step by step to higher dimensions. The number of channels dimension does not calculate the number of pixels. The execution unit is configured to: execute the data storage instruction. When the execution unit executes the data storage instruction, it includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; for the data of the first tensor, starting from the starting coordinates, sequentially obtaining data from the first tensor until the number of obtained data is equal to the number of pixels included in the second tensor; and storing the obtained data into memory.

[0128] For example, multiple tensors are stored in the buffer. The first tensor is one of the multiple tensors. The data storage format of the multiple tensors in the buffer can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. For the NDHWC data storage format, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the number of channels dimension. For the N(C / x)DHW(xC) data storage format, N corresponds to b1, indicating the batch dimension, (C / x) corresponds to b2, indicating the number of channels dimension, D corresponds to b3, indicating the depth dimension, H corresponds to b4, indicating the height dimension, and W(xC) corresponds to b5, indicating the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0129] For example, when multiple tensors are stored in the buffer, the execution unit executing the data storage instruction further includes: obtaining indication information of the storage position of the first tensor in the buffer. The indication information includes the starting coordinates of the first tensor in the buffer and the size values in five dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5.

[0130] For example, the execution unit storing the acquired data into the memory includes: storing the acquired data into the memory from the buffer area according to multiple instructions.

[0131] For example, the execution unit storing the acquired data into the memory further includes: determining a first data storage format of the second tensor in the memory; and placing the second tensor in the memory according to the first data storage format, where the first data storage format of the second tensor in the memory is the same as or different from the second data storage format of the first tensor in the buffer area.

[0132] For example, the execution unit executing a data storage instruction further includes: obtaining boundary values for at least a part of five dimensions of the first tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the first tensor, and where the boundary values are arbitrary values compared to the size values of the first tensor in that dimension.

[0133] For example, the execution unit sequentially obtaining data from the first tensor includes: for the data part of the first tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the first tensor defined by the boundary values; and for the data part of the first tensor not covered by the range defined by the boundary values, which is represented as invalid data, and for the invalid data, not storing it into the memory.

[0134] For example, the execution unit executing a data storage instruction further includes: obtaining a data stride value, where the data stride value is used to specify the stride for sequentially obtaining data from the first tensor. Specifically, the execution unit executing the data of the first tensor, starting from the starting coordinate, sequentially obtaining data from the first tensor includes: for the data of the first tensor, starting from the starting coordinate, sequentially obtaining data from the first tensor according to the data stride value.

[0135] As an example, Figure 13 FIG. shows a schematic block diagram of a processor according to some embodiments of the present disclosure. As Figure 13 shown, the processor 1000 may include an instruction parsing unit 1010 and an execution unit 1020. It can be understood that the processor 1000 may be implemented to execute a data loading method according to an embodiment of the present disclosure to load a tensor to be processed from an original tensor in the memory to the buffer area, or implement a data storage method according to an embodiment of the present disclosure to obtain a second tensor based on the first tensor in the buffer area and write the second tensor into the memory.

[0136] Regarding the specific implementation processes of the data storage method and the data loading method, reference may be made to the above description and will not be repeated here. The processor provided by at least one embodiment of the present disclosure can achieve technical effects similar to those of the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0137] According to some embodiments of the present disclosure, an electronic device is further provided. Figure 14 A schematic block diagram of an electronic device according to some embodiments of the present disclosure is shown. As Figure 14 shown, the electronic device 2000 may include a processor 2010 and a memory 2020 connected to the processor 2010. In addition, the processor 2010 may further include a buffer. According to an embodiment of the present disclosure, the memory 2020 may be implemented in the form of a high-bandwidth memory HBM, which is not limited thereto. Specifically, according to an embodiment of the present disclosure, the processor 2010 is configured to run computer-executable instructions, which, when run by the processor 2010, implement a data loading method according to an embodiment of the present disclosure to load a tensor to be processed from an original tensor in the memory to the buffer, or implement a data storage method according to an embodiment of the present disclosure to obtain a second tensor based on a first tensor in the buffer and write the second tensor to the memory.

[0138] The processor 2010 may perform various actions and processes according to a program stored in a non-transitory memory, for example. Specifically, the processor 2010 may refer to a processor chip capable of performing parallel computing. For example, it may be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In addition, the processor 2010 may also be implemented as other conventional types of processors, which are not limited herein.

[0139] Regarding the specific implementation processes of the data storage method and the data loading method, reference may be made to the above description, which will not be repeated herein. The processor provided by at least one embodiment of the present disclosure may achieve similar technical effects to the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0140] Figure 15 A block diagram of an example computing device implementing some embodiments of the present disclosure is shown. As Figure 15 shown, the computing device 3000 is, for example, suitable for implementing the data loading method or the data storage method provided by an embodiment of the present disclosure. It should be noted that Figure 15The components of the computing device 3000 shown are merely exemplary and not restrictive. According to actual application requirements, the computing device 3000 may also have other components.

[0141] As Figure 15 shown, the computing device 3000 may include a processing device 3010 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.

[0142] For example, when the computer-readable instructions are run by the processing device 3010, one or more steps of the data loading method described in any of the above embodiments, or one or more steps of the data storage method described in any of the above embodiments may be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0143] For example, the processing device 3010, the read-only memory (ROM) 3020, and the random access memory (RAM) 3030 are connected to each other via a bus 3040. The input / output (I / O) interface 3050 is also connected to the bus 3040.

[0144] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 3030 and / or cache, etc. For example, the computer-readable instructions may be loaded from the storage device 3080 into the random access memory (RAM) 3030 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 3020, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage medium, as well as various data used and / or generated by the application programs, etc.

[0145] Typically, the following devices can be connected to the I / O interface 3050: input devices 3060 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 3070 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 3080 including, for example, magnetic tape, a hard disk, a flash memory, etc.; and a communication device 3090. The communication device 3090 can allow the computing device 3000 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 15 the computing device 3000 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the computing device 3000 can alternatively implement or have more or fewer devices. For example, the processing device 3010 can control other components in the computing device 3000 to perform desired functions. The processing device 3010 can be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit (GPU), etc., which has data processing capabilities and / or program execution capabilities. The GPU can be directly integrated into a system on chip (SOC), directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.

[0146] According to some embodiments of the present disclosure, a non-transitory computer-readable storage medium is further provided, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data loading method according to the embodiments of the present disclosure is implemented, or the data storage method according to the embodiments of the present disclosure is implemented.

[0147] Figure 16 A schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 16 shown, the computer-readable storage medium 4000 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 4010 can be non-temporarily stored on the storage medium 4000. For example, when the computer-readable instructions 4010 are executed by a processor, one or more steps of the data loading method according to any of the above embodiments can be executed, or one or more steps of the data storage method according to any of the above embodiments can be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference can be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference can be made to the relevant descriptions in the embodiments of the data storage method above.

[0148] As an example, the storage medium 4000 can be applied to the electronic device 2000 and / or the computing device 3000. For example, the storage medium 4000 can be implemented as the storage device 3080 in the computing device 3000.

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0150] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.

[0151] The functions described above in this document can be at least partially performed by one or more hardware logic components. For example, without limitation, the exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0152] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0153] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0154] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

[0155] Regarding the present disclosure, the following points also need to be noted: (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0156] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0157] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the appended claims.

Claims

1. A data loading method for loading a to-be-processed tensor from an original tensor in memory into a buffer, wherein: The original tensor is a 5-dimensional tensor, and its shape and size are represented by 5 parameters b1, b2, b3, b4, and b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers, and the 5 dimensions include batch dimension, depth dimension, height dimension, width dimension, and channel number dimension; wherein, for the 5-dimensional tensor of the original tensor, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to higher dimensions, wherein the channel number dimension does not calculate the number of pixels, The data loading method comprises: Acquire the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; For the data of the original tensor, starting from the starting coordinate, sequentially acquire data from the original tensor until the number of acquired data is equal to the number of pixels included in the tensor to be processed; and The acquired data is loaded into the buffer area.

2. The method according to claim 1, wherein: The memory stores multiple tensors, the original tensor is one of the multiple tensors, and the data storage format of the multiple tensors in the memory can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimensional arrangement of the tensors in the storage component, wherein, for the data storage format of NDHWC, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the number of channels. For the data storage format of N(C / x)DHW(xC), N corresponds to b1, indicating the batch dimension, (C / x) corresponds to b2, indicating the number of channels, D corresponds to b3, indicating the depth dimension, H corresponds to b4, indicating the height dimension, and W(xC) corresponds to b5, indicating the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension.

3. The method according to claim 1, wherein: The tensor to be processed is used for convolution operation in a computing unit within the processor.

4. The method according to claim 1, wherein: The memory stores a plurality of tensors, and the data loading method further includes: Obtain indication information of the storage location of the original tensor in the memory, wherein the indication information includes the starting coordinates of the original tensor in the memory and size values ​​in five dimensions, which are respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, and tensor_b5.

5. The method according to claim 1, wherein: The loading of the acquired data into the buffer area comprises: The acquired data is loaded from the memory into the cache area according to a plurality of instructions.

6. The method according to claim 5, wherein: The step of loading the acquired data into the cache area further comprises: Determine a first data storage format of the tensor to be processed in the buffer area; and The tensor to be processed is placed in the cache area according to the first data storage format, wherein the first data storage format of the tensor to be processed in the cache area is the same as or different from the second data storage format of the original tensor in the memory.

7. The method according to claim 4, further comprising: Obtain boundary values ​​for the original tensor in at least a portion of the five dimensions, respectively, wherein the boundary values ​​are used to limit a boundary range for obtaining data from the original tensor, wherein the boundary values ​​are arbitrary values ​​compared to the size values ​​of the original tensor in this dimension.

8. The method according to claim 7, wherein: The sequentially acquiring data from the original tensor comprises: For a data portion of the original tensor whose range defined by the boundary value covers the original tensor, sequentially acquiring data from the range of the original tensor defined by the boundary value; and The data portion of the original tensor that is not covered by the range defined by the boundary value is represented as invalid data. The invalid data is directly filled into the tensor to be processed by the buffer area according to a predetermined value, wherein the predetermined value is equal to 0.

9. The method according to claim 1, further comprising: Obtaining a data step value, wherein the data step value is used to specify a step size for sequentially obtaining data from the original tensor, wherein, for the data of the original tensor, taking the starting coordinate as the starting point, sequentially obtaining data from the original tensor includes: For the data of the original tensor, taking the starting coordinate as the starting point, data is sequentially acquired from the original tensor according to the data step value.

10. A data loading method, comprising: Receive a data loading instruction instructing execution of loading a to-be-processed tensor from an original tensor in memory into a cache area, wherein the original tensor is a 5-dimensional tensor, and its shape and size are represented by 5 parameters b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers, and the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension; wherein, for the 5-dimensional tensor of the original tensor, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and After parsing the data loading instruction, using the execution unit to execute the data loading instruction, Wherein, using the execution unit to execute the data loading instruction includes: Acquire the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; For the data of the original tensor, starting from the starting coordinate, sequentially acquire data from the original tensor until the number of acquired data is equal to the number of pixels included in the tensor to be processed; and The acquired data is loaded into the buffer area.

11. A data storage method for acquiring a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, wherein: The first tensor is a 5-dimensional tensor, and its shape and size are represented by 5 parameters b1, b2, b3, b4, and b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers, and the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension; wherein, for the 5-dimensional tensor of the first tensor, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to higher dimensions, wherein the channel number dimension does not calculate the number of pixels, The data storage method comprises: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; For the data of the first tensor, starting from the starting coordinate, sequentially acquire data from the first tensor until the number of acquired data is equal to the number of pixels included in the second tensor; and The acquired data is stored in the memory.

12. The method according to claim 11, wherein: The buffer area stores multiple tensors, the first tensor is one of the multiple tensors, and the data storage format of the multiple tensors in the buffer area can be any one of NDHWC or N(C / x)DHW(xC), and the data storage format is used to indicate the storage order and dimensional arrangement of the tensors in the storage component, wherein, for the data storage format of NDHWC, N corresponds to b1, indicating the batch dimension, D corresponds to b2, indicating the depth dimension, H corresponds to b3, indicating the height dimension, W corresponds to b4, indicating the width dimension, and C corresponds to b5, indicating the number of channels dimension; for the data storage format of N(C / x)DHW(xC), N corresponds to b1, indicating the batch dimension, (C / x) corresponds to b2, indicating the number of channels dimension, D corresponds to b3, indicating the depth dimension, H corresponds to b4, indicating the height dimension, and W(xC) corresponds to b5, indicating the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension.

13. The method according to claim 11, wherein: The buffer area stores a plurality of tensors, and the data storage method further comprises: Obtain indication information of a storage location of the first tensor in the cache area, the indication information including a starting coordinate of the first tensor in the cache area and size values ​​in five dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, and tensor_b5.

14. The method according to claim 11, wherein: The storing the acquired data into the memory comprises: The acquired data is stored from the buffer area to the memory according to a plurality of instructions.

15. The method according to claim 14, wherein: The storing the acquired data into the memory further comprises: Determine a first data storage format of the second tensor in the memory; The second tensor is placed in the memory according to the first data storage format, wherein the first data storage format of the second tensor in the memory is the same as or different from the second data storage format of the first tensor in the cache area.

16. The method according to claim 13, further comprising: Obtain boundary values ​​for the first tensor in at least a portion of the five dimensions, wherein the boundary values ​​are used to limit a boundary range for obtaining data from the first tensor, wherein the boundary values ​​are arbitrary values ​​compared to the size values ​​of the first tensor in this dimension.

17. The method according to claim 16, wherein: The sequentially acquiring data from the first tensor includes: For a data portion of the first tensor whose range defined by the boundary value covers the data portion of the first tensor, sequentially acquiring data from the range of the first tensor defined by the boundary value; and The data portion of the first tensor that is not covered by the range defined by the boundary value is represented as invalid data, and the invalid data is not stored in the memory.

18. The method according to claim 11, further comprising: Obtaining a data step value, wherein the data step value is used to specify a step size for sequentially obtaining data from the first tensor, wherein, for the data of the first tensor, taking the starting coordinate as a starting point, sequentially obtaining data from the first tensor includes: For the data of the first tensor, taking the starting coordinate as the starting point, sequentially obtain data from the first tensor according to the data step value.

19. A data storage method, comprising: Receive a data storage instruction for executing a first tensor in a buffer to obtain a second tensor and write the second tensor into a memory, wherein the first tensor is a 5-dimensional tensor, and its shape and size are represented by 5 parameters b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers, and the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension; wherein, for the 5-dimensional tensor of the first tensor, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and After parsing the data storage instruction, use the execution unit to execute the data storage instruction, The step of using the execution unit to execute the data storage instruction includes: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; For the data of the first tensor, starting from the starting coordinate, sequentially acquire data from the first tensor until the number of acquired data is equal to the number of pixels included in the second tensor; and The acquired data is stored in the memory.

20. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the data loading instruction instructs execution to load a to-be-processed tensor from an original tensor in a memory into a cache area, wherein the original tensor is a 5-dimensional tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers, and the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension; wherein, for the 5-dimensional tensor of the original tensor, the data determined by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and The execution unit is configured to: execute the data loading instruction, The execution unit executes the data loading instruction, including: Acquire the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; For the data of the original tensor, starting from the starting coordinate, sequentially acquire data from the original tensor until the number of acquired data is equal to the number of pixels included in the tensor to be processed; and The acquired data is loaded into the buffer area.

21. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is configured to: receive and parse a data storage instruction, wherein the data storage instruction instructs execution to obtain a second tensor based on a first tensor in a cache area and write the second tensor into a memory, wherein the first tensor is a 5-dimensional tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, and b5, b1, b2, b3, b4, and b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers, and the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension; wherein, for the 5-dimensional tensor of the first tensor, the data determined by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and The execution unit is configured to: execute the data storage instruction, The execution unit executes the data storage instruction, including: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; For the data of the first tensor, starting from the starting coordinate, sequentially acquire data from the first tensor until the number of acquired data is equal to the number of pixels included in the second tensor; and The acquired data is stored in the memory.

22. An electronic device comprising a processor and a memory connected to the processor, wherein: The processor includes a cache area, wherein The processor is configured to run computer-executable instructions, which, when run by the processor, implement the data loading method according to any one of claims 1-10 to load the tensor to be processed from the original tensor in the memory to the cache area, or implement the data storage method according to any one of claims 11-19 to obtain the second tensor based on the first tensor in the cache area and write the second tensor to the memory.

23. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the data loading method according to any one of claims 1-10 is implemented, or the data storage method according to any one of claims 11-19 is implemented.

Citation Information

Patent Citations

  • Multi-operator operation method and device for neural network model

    CN116134446A

  • Tensor data processing method and device and storage medium

    CN119105798A

  • Tensor convolution calculation method and device, equipment, storage medium and program product

    CN119849570A

  • Flexible compute array utilization in a tensor processor

    US11972349B1

  • Compilation optimization method and apparatus, computer device and storage medium

    WO2023030507A1