Data loading method, data storage method, processor, electronic equipment and medium

By loading the pending tensor from the original tensor of memory into the cache area in a parallel processor, and data is obtained and stored sequentially using multiple requests, the problem of low data loading/storage efficiency in the prior art is solved, and efficient memory access and hardware utilization are achieved.

CN120216402AActive Publication Date: 2025-06-27SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510607276.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-27
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to provide fast and efficient data loading/storage methods, which affects the computing performance of parallel processors.

Method used

By loading the pending tensor from the original tensor in memory to the cache, obtaining data in sequence using multiple requests and writing sequentially in the cache, optimizes the loading and stored procedures of data.

Benefits of technology

Improves data loading/storage efficiency, improves memory access bandwidth and hardware utilization of computing devices, and enhances the overall performance of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216402A_ABST
    Figure CN120216402A_ABST
Patent Text Reader

Abstract

The invention provides a data loading method, a data storage method, a processor, electronic equipment and a medium. The data loading method is used for loading a to-be-processed tensor from an original tensor of a memory to a cache region, and comprises the following steps: determining a plurality of requests for loading the to-be-processed tensor in combination with the number of pixels included in the to-be-processed tensor, a data storage format of the original tensor and an initial coordinate of the to-be-processed tensor in a coordinate system determined by the original tensor, the plurality of requests are used for sequentially acquiring data of the original tensor from the original tensor by taking the starting coordinate as a starting point until the number of the acquired data is equal to the number of pixels included in the tensor to be processed, and the data loaded by each request belong to the data range of the original tensor or do not belong to the data range of the original tensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a medium. Background Art

[0002] A tensor is a data structure of a multi-dimensional array. Tensor operations are widely used in processors such as parallel processors. For example, in the field of deep learning, the dimensions of the input data, the intermediate data processed during the deep learning process, and the output data are elastic and not exact. Therefore, an elastic data form is needed to describe various types of data, and thus the concept of a tensor is generated. In the field of deep learning, all data to be operated on can be stored and exist in the form of a tensor. If the data is not in the form of a tensor, it needs to be first converted into the data structure form of a tensor. As an example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor itself can have any number of dimensions. For example, it can be represented as a 5-dimensional array.

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processing devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of a parallel processor need to process a large amount of data, and these data are generally stored in the storage components of the parallel processor. For example, the storage component can be a high-speed memory. Through data loading instructions, these data can be loaded from the storage component to the buffer for calculation, and through data storage instructions, the data in the buffer can be stored in the memory.

[0004] How to provide a fast and efficient data loading / storing method is crucial for the computing performance of the device. Summary of the Invention

[0005] Embodiments of the present disclosure provide a data loading method, a data storage method, a processor, an electronic device, and a medium, which are used to provide a fast and efficient data loading / storing method, improve the memory access bandwidth, and improve the hardware utilization rate of the computing device.

[0006] According to a first aspect of the present disclosure, a data loading method is provided for loading a tensor to be processed from an original tensor in memory into a buffer area, wherein the data storage format of the original tensor in memory is the same as the data storage format of the tensor to be processed in the buffer area, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in 5 dimensions and are all positive integers. The data loading method includes: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into the buffer area to load the tensor to be processed into the buffer area, where in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, the requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the data storage format of the original tensor indicates that the first dimension is prior to the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0007] According to a second aspect of the present disclosure, a data loading method is provided, including: receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in a memory to a buffer, where the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component; the shape dimensions of the tensor to be processed are represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the dimensions of the tensor to be processed in five dimensions and are all positive integers, and the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension; the shape dimensions of the original tensor are represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the original tensor in five dimensions and are all positive integers; and after parsing the data loading instruction, using an execution unit to execute the data loading instruction, where using the execution unit to execute the data loading instruction includes: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; determining a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests and writing the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer, where in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, partitioning requests are made in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension in storage or loading, and the first dimension and the second dimension are adjacent.

[0008] According to a third aspect of the present disclosure, a data storage method is provided for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, wherein the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in memory, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape size of the second tensor is represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape size of the first tensor is represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers. The data storage method includes: obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests and writing the data corresponding to each request into memory in sequence to store the second tensor into memory, wherein in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data obtained by each request either all belong to the data range of the first tensor or all do not belong to the data range of the first tensor, and in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, wherein the data storage format of the first tensor indicates that the first dimension takes precedence over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0009] According to a fourth aspect of the present disclosure, a data storage method is provided, including: receiving a data storage instruction indicating to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, wherein a data storage format of the first tensor in the buffer is the same as a data storage format of the second tensor in the memory, the data storage format is used to indicate a storage order and a dimension arrangement of a tensor in a storage component, a shape size of the second tensor is represented by a1, a2, a3, a4, a5, a1, a2, a3, a4, a5 respectively indicate sizes of the second tensor in 5 dimensions and are all positive integers, the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, a shape size of the first tensor is represented by b1, b2, b3, b4, b5, b1, b2, b3, b4, b5 respectively indicate sizes of the first tensor in 5 dimensions and are all positive integers; and after parsing the data storage instruction, executing the data storage instruction using an execution unit, wherein executing the data storage instruction using the execution unit includes: obtaining a number of pixels included in the second tensor, a data storage format of the first tensor, and a starting coordinate of the second tensor in a coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinate of the second tensor in the coordinate system determined by the first tensor, determining a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinate until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests, and sequentially writing data corresponding to each request into the memory to store the second tensor into the memory, wherein in response to a size relationship between the second tensor and the first tensor in a first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but continuous acquisition can be performed in dimensions lower than the first dimension, division of requests is performed in a second dimension, data obtained by each request either all belongs to a data range of the first tensor or all does not belong to the data range of the first tensor, and in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, wherein the data storage format of the first tensor indicates that the first dimension is prior to the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0010] According to a fifth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, where the data loading instruction instructs to execute loading a tensor to be processed from an original tensor in a memory to a buffer, where the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in a storage component, the shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5, a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in five dimensions and are all positive integers, the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, and the shape and size of the original tensor are represented by b1, b2, b3, b4, b5, b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in five dimensions and are all positive integers; and the execution unit is configured to: execute the data loading instruction. Wherein, the execution unit executes the data loading instruction, including: obtaining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor, determining a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests, writing the data corresponding to each request sequentially into the buffer to load the tensor to be processed into the buffer, where, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension in storage or loading, and the first dimension and the second dimension are adjacent.

[0011] According to a sixth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, where the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory, the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in a storage component, the shape size of the second tensor is represented by a1, a2, a3, a4, a5, a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers, the 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension, the shape size of the first tensor is represented by b1, b2, b3, b4, b5, b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers; and the execution unit is configured to: execute the data storage instruction. Wherein, the execution unit executes the data storage instruction, including: obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in a coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in a coordinate system determined by the first tensor, determining a plurality of requests for storing the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the memory to store the second tensor into the memory, where, in response to the dimensional relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but continuous acquisition can be performed in dimensions lower than the first dimension, partitioning requests are performed in the second dimension, the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor, and in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, where the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0012] According to a seventh aspect of the present disclosure, an electronic device is provided, including a processor and a memory connected to the processor, where the processor includes a buffer, and the processor is configured to run computer-executable instructions, and the computer-executable instructions, when run by the processor, implement the data loading method according to the embodiments of the present disclosure to load a tensor to be processed from an original tensor in the memory into the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain a second tensor based on a first tensor in the buffer and write the second tensor into the memory.

[0013] According to an eighth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, a data loading method according to an embodiment of the present disclosure is implemented, or a data storage method according to an embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 A schematic structural diagram of a General-Purpose computing on Graphics Processing Unit (GPGPU) is shown; Figure 2 A schematic structure of a tensor is shown; Figure 3A A schematic diagram of the data storage format of NDHWC is shown; Figure 3B A schematic diagram of the data storage format of N(C / x)DHW(xC) is shown, where x = 32; Figure 4 A schematic diagram of obtaining a tensor in the related art is shown; Figure 5 A schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure is shown; Figure 6 A schematic diagram of obtaining a tensor to be processed in a continuous manner according to an embodiment of the present disclosure is shown; Figure 7A A schematic diagram of obtaining a tensor to be processed in a continuous manner for the data storage format of NDHWC according to an embodiment of the present disclosure is shown; Figure 7B A schematic diagram of obtaining a tensor to be processed in a continuous manner for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure is shown; Figure 8A A schematic diagram of setting boundary values for an original tensor according to an embodiment of the present disclosure is shown; Figure 8B A schematic diagram of continuous data acquisition in the case of setting boundaries for an original tensor according to an embodiment of the present disclosure is shown; Figure 8C A schematic diagram showing the continuous acquisition of a tensor to be processed according to a set step size according to an embodiment of the present disclosure; Figure 9A A state diagram of a state machine provided according to some embodiments of the present disclosure; Figure 9B Shows the state transition according to the PerW partitioning method for the NDHWC data storage format; Figure 10 A schematic flowchart of a data loading method provided according to at least one embodiment of the present disclosure; Figure 11 A schematic flowchart of a data storage method provided according to at least one embodiment of the present disclosure; Figure 12 A schematic flowchart of a data storage method provided according to at least one embodiment of the present disclosure; Figure 13 A schematic block diagram of a processor according to some embodiments of the present disclosure; Figure 14 A schematic block diagram of an electronic device according to some embodiments of the present disclosure; Figure 15 A block diagram showing an example computing device implementing according to some embodiments of the present disclosure; and Figure 16 A schematic block diagram of a computer-readable storage medium according to some embodiments of the present disclosure. Detailed implementation manners

[0016] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0017] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Words such as "upper", "lower", "left", "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of some known functions and known components are omitted in this disclosure.

[0018] Figure 1 A structural schematic diagram of a GPGPU is shown. As Figure 1 shown, a GPGPU is actually an array of programmable multiprocessors. For example, the programmable multiprocessors can be Streaming Processor Clusters (SPCs), such as including Figure 1 the streaming processor cluster 1 shown, ..., the streaming processor cluster M, where M is a positive integer. In a general-purpose graphics processor, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. As an example, data sharing between multiple streaming processor clusters is performed through a global cache or High Bandwidth Memory (HBM).

[0019] As Figure 1 shown, taking the streaming processor cluster 1 as an example, 1 streaming processor cluster can include multiple Compute Units (CUs), such as Figure 1 the compute unit 1, compute unit 2, ..., compute unit K in it, where K is a positive integer. Each compute unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A compute unit can include multiple cores (also called computing cores or compute cores), and each compute core includes an Arithmetic Logic Unit (ALU), a floating-point computing unit, etc., and the compute core is used to perform specific computing tasks. In addition, the compute unit also includes registers (such as Figure 1The register bank) and shared memory are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit. In addition, the buffer can be understood as being used to share data among the computing units within the streaming processor cluster.

[0020] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processing unit (or parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0021] In each computing unit, the thread bundle scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates the thread bundles so that multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-shared. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the middle-level cache or global cache or high-bandwidth memory for read / write operations, etc.

[0022] As Figure 1 shown, general computing operations, such as matrix computing operations in the field of artificial intelligence, usually require a large amount of data. These data are usually stored in a memory, such as in a high-bandwidth memory HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing result, data needs to be stored to the memory (Store operation). The storage method of data in the memory will affect the memory access bandwidth, and thus affect the hardware utilization rate of the computing unit.

[0023] For example, general computing operations include General Matrix Multiplication (abbreviated as GEMM). As an example, the data required for general matrix multiplication is represented as two 5D arrays, such as the matrix multiplication calculation of two tensors A and tensor B. In addition, general computing operations also include convolution operations, which are manifested as data dot products. It can be understood that in the field of artificial intelligence, other computing operations are also involved, which will not be listed one by one here. The data involved in these calculations is usually embodied in the form of tensors.

[0024] For example, for a certain tensor A in the buffer, its shape and size can be represented by a1, a2, a3, a4, and a5. a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor data in 5 dimensions, and a1, a2, a3, a4, and a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, N represents the batch dimension, which is the number of data samples captured in one training. The D dimension represents the depth dimension, the H dimension represents the height dimension of the input data, the W dimension represents the width dimension of the input data, and the C dimension represents the number of channels dimension. For example, taking tensor A as an example, a1 can be the size of the N dimension, a2 can be the size of the D dimension, a3 can be the size of the H dimension, a4 can be the size of the W dimension, and a5 can be the size of the C dimension. Of course, the present disclosure does not make specific limitations on this.

[0025] As an example, Figure 2 shows a schematic structure of a tensor. In Figure 2 the shown tensor, a1 is the size of the N dimension and is equal to 1, a2 is the size of the D dimension and is equal to 1, a3 is the size of the H dimension and is equal to 5, a4 is the size of the W dimension and is equal to 4, and a5 is the size of the C dimension and is equal to 64. For example, Figure 2 the pixel elements of the tensor in

[0026] are represented as 0, 1, 2, 3,... and so on. Figure 2 The placement of the tensor in the memory (such as memory or buffer) can have multiple formats, called data storage format (layout). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The following describes different data storage formats with the

[0027] shown tensor. Figure 3A In the related art, the data storage format can include NDHWC, also known as the Linear mode.

[0028] Figure 3A shows a schematic diagram of the NDHWC data storage format. For example, for the NDHWC linear mode, as Figure 3A shown, starting from the first element (element 0 in Figure 3A Figure 3A Figure 3A Figure 3A Figure 3A ​​​​After the element 1260) in, select the first channel (a5 = 0, Figure 3A the second element of c0) in Figure 3A element 1) in, and then store the second element of the second channel (a5 = 1, Figure 3A c1) in Figure 3A element 21) in, and so on until the second elements of all channels are laid out, and so on.

[0029] In the related art, the data storage format may also include N(C / x)DHW(xC), also known as the Interleave mode, where x can be set to 8, 16, 32, etc. as needed.

[0030] The N(C / x)DHW(xC) data storage format is similar to the NDHWC data storage format, but there is a key difference. In the layout of N(C / x)DHW(xC), a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.

[0031] Figure 3B Shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32.

[0032] As Figure 3B shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 3B c0) in to a5 = 31 ( Figure 3B c31) in, and the second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format.

[0033] In memory, for example, a certain tensor B, similar to a certain tensor A in the buffer, this tensor B can be stored in memory in one of the two data storage formats described above. The shape dimensions of the tensor can be similarly expressed as b1×b2×b3×b4×b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the tensor B in these 5 dimensions and are all positive integers.

[0034] It can be understood that in the related art and possible future developments, the data storage format of tensors is not limited to the above-described N(C / x)DHW(xC) data storage format and NDHWC data storage format. The method described in this disclosure does not limit this. Further, for multiple tensors that can be stored in memory and the buffer during the calculation process, usually the storage space of memory is much larger than that of the buffer, but it is farther from the calculation unit, and the data transfer efficiency is lower than that of the buffer.

[0035] In the related art, during the computing process of a processing device, a large amount of computing data will be generated, for example, in the form of tensors, which can be temporarily stored in a buffer. For example, the buffer here can refer to Figure 1 the buffer in the streaming processor cluster shown in []. Further, this data can also be transferred from the buffer or directly stored in the memory. For example, the memory can be Figure 1 the high-bandwidth memory HBM shown in []. The storage forms of tensors in the memory and the buffer can both be any one of the N(C / x)DHW(xC) data storage format and the NDHWC data storage format described above. Thus, during the computing process, according to factors such as technical requirements and the respective storage characteristics of the buffer and the memory, a large amount of data transfer processes need to be performed between the two. For example, tensors in the buffer are stored in the memory through storage instructions, or tensors in the memory are loaded into the buffer through load instructions.

[0036] It can be understood that in this article, the data storage process of storing tensors in the buffer into the memory and the data loading process of loading tensors in the memory into the buffer can be implemented in a similar manner. Therefore, for the sake of convenience in description, in some embodiments or examples, only the data loading process is described as an example, and those skilled in the art can apply it similarly to the data storage process. For the differences between the two, additional descriptions will be made.

[0037] As an example, Figure 4 shows a schematic diagram of obtaining a tensor in the related art. As Figure 4 shown, there is a tensor A stored in the memory, which can be placed in the memory in the above N(C / x)DHW(xC) or NDHWC data storage format. That is, tensor A is a 5D array. In Figure 4 only three dimensions, W, H, and C, are schematically shown. Schematically, on the left side of Figure 4 a dimension coordinate system of tensor A is shown, where they are the W dimension, the H dimension, and the C dimension respectively. During the data loading process, all or part of the data in tensor A can be loaded into the buffer through a load instruction. For example, Figure 4 tensor B in [] is loaded into the buffer as a whole. Specifically, the load instruction can indicate the first starting point (C = 0, W = 0, H = 0) of tensor A to be loaded, the second starting point (C = 0, W = 3, H = 0) of this tensor B in the coordinate system of tensor A, and the size of this tensor B in each dimension. Schematically, in Figure 4In the 3D schematic diagram shown, the tensor B is a cuboid determined by the above-mentioned second starting point and each dimension size. Based on the information about the second starting point and the dimension sizes of the tensor to be loaded, the memory can load the tensor B into the buffer area.

[0038] In the above-related technologies, the block-based data loading method can be applied to general matrix multiplication (GEMM) calculations in general computing operations. However, the above block-based data loading method is not applicable to convolution operations, whose computational characteristics are data dot products. The above block-based data loading method will limit the computational efficiency of such per-pixel type convolution operations, increase data access time, and reduce the overall performance of the processor, which limits the further development space of efficient and general-purpose processors.

[0039] Furthermore, in the memory, the original tensor is continuously stored in the memory in the data storage format as described above. The shape of the original tensor is represented as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in these 5 dimensions and are all positive integers. When extracting or storing a partial tensor in the original tensor, since it is a partial tensor inside the original tensor, each dimension may not be continuously extractable during extraction, and the exact data amount and data location cannot be obtained during data loading or storage to optimally achieve data loading or storage, greatly reducing the data bandwidth during data loading and the hardware computational efficiency.

[0040] In view of the above technical problems in the related technologies, the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. In the present disclosure, first, the tensor is no longer obtained and transported in a block form, but is sequentially obtained and transported in units of pixels to be applicable to computational operations such as convolution operations. Further, for the operation of loading the tensor to be processed from the original tensor in the memory into the buffer area, multiple requests for loading the tensor to be processed are determined according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, so as to sequentially obtain data from the original tensor starting from the starting coordinates using the determined multiple requests until the number of obtained data is equal to the number of pixels included in the tensor to be processed, thereby improving the data transportation efficiency for data loading operations, reducing the memory access time, and improving the overall performance of the processor.

[0041] In the data loading method provided by at least one embodiment of the present disclosure, requests are divided according to whether the tensors can be continuously loaded in the first dimension. The data to be loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby enhancing the hardware utilization rate of the computing unit and improving the hardware performance.

[0042] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0043] The data loading method provided by at least one embodiment of the present disclosure is used to load the tensor to be processed from the original tensor in the memory to the buffer area. The memory according to the embodiments of the present disclosure can be, for example, a high-bandwidth memory (HBM), and the buffer area can be, for example, the buffer (buffer) in the streaming processor cluster, which is not limited herein.

[0044] In the method according to the embodiments of the present disclosure, the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer area. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. For example, the storage component refers to the above-mentioned HBM or buffer area. That is to say, in the data loading method according to the embodiments of the present disclosure, the placement manner of the original tensor in the memory is the same as the placement manner of the tensor to be processed to be loaded in the buffer area next. For example, both are arranged in the NDHWC linear manner or both are arranged in the N(C / x)DHW(xC) interleaved manner, which is not limited herein. The method proposed by the present disclosure is applicable to the above two data storage formats.

[0045] The original tensor in the memory is a 5D tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, depth dimension, height dimension, width dimension, and number of channels dimension. Similarly, the shape size of the tensor to be processed is represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers.

[0046] For the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension.

[0047] For the data storage format of N(C / x)DHW(xC), N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension. Generally, x can be set to an integer multiple of 4.

[0048] Regarding the features of the above two data storage formats, reference can be made to the above description in combination with Figure 3A - Figure 3B and will not be repeated here.

[0049] Furthermore, for the 5D data of the original tensor, the data determined by the height dimension and the width dimension represents a pixel (or, can also be called an element), and accumulates step by step towards the higher dimensions. Among them, the number of pixels is not calculated for the number of channels dimension. As an example, as Figure 4 shown in, the values of each W dimension and H dimension can determine a pixel. For example, W = 0 and H = 0 correspond to the first pixel (or element) in tensor A, and W = 1 and H = 0 correspond to the second pixel in tensor A. In the tensor, the C dimension does not affect the number of pixels.

[0050] Figure 5 is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 5 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S101 - S103.

[0051] In step S101, obtain the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Then, in step S102, combine the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine multiple requests for loading the tensor to be processed, where the multiple requests are used to sequentially obtain data from the original tensor with the starting coordinates as the starting point until the number of obtained data is equal to the number of pixels included in the tensor to be processed. In step S103, sequentially send the multiple requests and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area. According to the embodiment of the present disclosure, the tensor to be processed can be used for convolution operations in the computing units within the processor. Herein, the number of data can be expressed as the number of pixels.

[0052] In the data loading method according to an embodiment of the present disclosure, in order to obtain a tensor to be processed from the original tensor in the memory, it is necessary to indicate the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. For example, the starting coordinates may be the second starting point (C = 0, W = 3, H = 0) described in conjunction with Figure 4 Furthermore, in the data loading method according to an embodiment of the present disclosure, it is also necessary to indicate the number of pixels included in the tensor to be processed (e.g., denoted as copy_pixel_num), that is, the total number of pixels in the tensor that is desired to be loaded into the buffer. This continuous data copying method is different from the block-based data loading method adopted in the related art above (wherein it is necessary to indicate the sizes of the tensor to be processed in each dimension). According to the data loading method of the embodiment of the present disclosure, the range of the tensor to be processed is determined by the number of pixels to be obtained. Further, instead of obtaining the tensor from the original tensor in a block manner, data is sequentially obtained from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed.

[0053] Next, first, the implementation process of implementing the above continuous data loading based on copy_pixel_num and the starting coordinates will be described, and then, the implementation process of determining multiple requests for loading the tensor to be processed will be described later.

[0054] As an example, Figure 6 shows a schematic diagram of obtaining a tensor to be processed in a continuous manner according to an embodiment of the present disclosure. In Figure 6 , tensor A represents the original tensor in the memory. In Figure 6 , each pixel point is shown as a square, and tensor A covers all the squares, that is, it includes both white squares and gray-shaded squares. The starting point of this original tensor is denoted as (C = 0, W = 0, H = 0), for example, to indicate the position of this original tensor in the memory. Then, in Figure 6 , tensor B represents the tensor to be obtained, and its starting point is denoted as (C = 0, W = 3, H = 0), that is, data is obtained starting from the 3rd pixel point in the first row of tensor A. Then, in Figure 6 , the process of obtaining the tensor to be processed in a sequential (or also referred to as continuous) manner according to an embodiment of the present disclosure is schematically shown. That is, starting from the starting point (C = 0, W = 3, H = 0) as shown by the dashed arrow, data is obtained from tensor A pixel by pixel until the number of obtained pixels reaches copy_pixel_num. In the example of Figure 6 , only the three dimensions of W, H, and C are shown. It can be understood that if the data in these 3 dimensions still does not reach copy_pixel_num, data in higher dimensions can be further obtained, which is not limited herein. In addition, asFigure 6 As shown in, the C dimension itself does not affect the number of pixels. Suppose, Figure 6 the data storage format of tensor A in in memory is NDHWC. In Figure 6 the example shown, the total number of pixels of the tensor to be processed is copy_pixel_num = 27.

[0055] In an embodiment according to the present disclosure, the step of sequentially obtaining data from the original tensor starting from the starting coordinates for the data of the original tensor includes: removing the channel number dimension from the 5 dimensions of the original tensor, and in the order of dimensions from low to high, using the starting coordinates as the starting point for data acquisition, and sequentially obtaining data until the number of pixels obtained reaches copy_pixel_num. The reason for removing the channel number dimension from the 5 dimensions is that for a tensor, the channel number dimension does not count the number of pixels. That is to say, S102 may include obtaining pixel data from the original tensor starting from the starting coordinates as the starting point for data acquisition in the order of the width dimension, height dimension, depth dimension, and batch dimension until the number of pixels obtained reaches copy_pixel_num.

[0056] Compare Figure 4 and Figure 6 the two ways of obtaining the tensors to be processed shown, the data loading method provided by the embodiments of the present disclosure can achieve per-pixel data loading, rather than Figure 4 the block-by-block acquisition in, this data loading method is more beneficial to the calculation process such as convolution operation and is beneficial to improving the operation efficiency. Thus, based on the method provided by the embodiments of the present disclosure, for the data to be subjected to convolution operation next, the processor can, for example, indicate in the form of an instruction to fetch this part of the data from the memory to the buffer area according to the above continuous data loading method for use in the convolution operation. It can be understood that the above processor can reasonably use whether to use the continuous data loading method or the block-by-block data loading method according to the type of operation to be performed or the data processing characteristics, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer area may also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0057] According to some embodiments of the present disclosure, there are multiple tensors stored in the memory, and the data loading method may further include: obtaining indication information of the storage location of the original tensor in the memory. As an example, the indication information may include the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. It can be understood that the storage space of the memory is generally much larger than the buffer, and various tensor data generated during the calculation process can be stored therein. These tensor data can be arranged in the memory in any one of the above-mentioned N(C / x)DHW(xC) or NDHWC data storage formats. In the embodiments according to the present disclosure, in order to enable the memory to know the specific location of the tensor to be obtained in the memory, indication information of the storage location of the original tensor in the memory may also be obtained. The indication information includes the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. As an example, in the case where the original tensor is in the NDHWC data storage format, tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5 respectively represent the values of the original tensor in the 5 dimensions of N, D, H, W, and C.

[0058] Regarding the influence of the data storage format on the loading process of the tensor to be processed ("tensor B"), it will be described below in conjunction with Figure 7A - Figure 7B will be described. In conjunction with Figure 7A and Figure 7B In the example shown, the data storage method of the tensor to be processed in the buffer is the same as the data arrangement method of the original tensor in the memory.

[0059] Figure 7A FIG. shows a schematic diagram of continuously obtaining the tensor to be processed for the NDHWC data storage format according to an embodiment of the present disclosure. For the Figure 7A tensor shown, its data storage format in the memory is in the form of NDHWC, or is arranged in the NDHWC manner.

[0060] In Figure 7A In the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 8 (the C dimension is not shown in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Taking the N dimension as an example, N_dim = 2 corresponds to Figure 7AFor N0 and N1 in it, based on the above parameters, the total number of pixel points included in the original tensor can be obtained as 64.

[0061] Next, in Figure 7A In the example of, the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixel points included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7A In the N = 0 and D = 0 dimensions shown in, the pixel points where W = 2 and H = 2 are located. Taking this pixel point as the starting point, according to the data arrangement method, in the order from low dimension to high dimension (that is, in the order of W, H, D, N, as shown by the data acquisition schematic arrow in Figure 7A ), data is sequentially acquired until the number of acquired data is equal to the number of pixel points included in the tensor to be processed. In Figure 7A In the example shown in, taking the determined starting point as the starting point, data is acquired pixel by pixel in the order of the arrow, so as to load this part of the data in the original tensor in memory into the buffer area as the tensor to be processed for subsequent operations such as convolution calculations.

[0062] Figure 7B shows a schematic diagram of continuously acquiring the tensor to be processed for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure. For Figure 7B the tensor shown in, its data storage format in memory is in the form of N(C / x)DHW(xC), or is arranged in the manner of N(C / x)DHW(xC). In Figure 7B In the example of, x = 8, that is, every 8 C channels are grouped and bound to the W dimension. Compared with Figure 7A the NDHWC data storage format shown in, for the N(C / x)DHW(xC) data storage format, the level of the C dimension in the tensor is higher and is located between the N dimension and the D dimension.

[0063] In Figure 7B In the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 16 (shown as two 8*C dimensions in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Thus, the total number of pixel points included in the original tensor can be obtained as 64.

[0064] Next, in Figure 7B the example of, the coordinates of the starting point of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixels included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7B the pixel point where W = 2 and H = 2 under N = 0 and D = 0 and the first 8*C dimension shown in, starting from this pixel point, according to the data arrangement method, in the order from low dimension to high dimension (that is, in the order of W, H, D, C, N, as shown by the data acquisition indication arrow in Figure 7B ), data is sequentially acquired until the number of acquired data is equal to the number of pixels included in the tensor to be processed. Compared with the acquisition order shown in Figure 7A , since the rank of the C dimension is higher, in the example of Figure 7B , first, data of W, H, D, and the first 8*C dimension is acquired according to the starting point, then, data of the second 8*C dimension is acquired in the order of W, H, D, and finally, data of the N dimension is acquired. It can be understood that in the examples of Figure 7A and Figure 7B , the number of pixels of the tensor to be processed acquired is 40, and the difference is only caused by the different arrangement methods of the C dimension. In addition, it should be noted that for the second 8*C dimension shown in Figure 7B , its starting point at the W and H dimensions should be aligned with the starting point of the first 8*C dimension.

[0065] In the example shown in Figure 7B , starting from the determined starting coordinates, data is acquired pixel by pixel in the arrow order, so as to load this part of the data in the original tensor in the memory into the buffer area as the tensor to be processed for subsequent operations such as convolution calculations.

[0066] In some embodiments according to the present disclosure, boundary values can also be defined for at least a part of the five dimensions of the original tensor respectively. As an example, left and right boundary values can be set for, for example, the C dimension, the W dimension, the H dimension, and the D dimension, which are expressed as (L-tensor_C, R-tensor_C), (L-tensor_W, R-tensor_W), (L-tensor_H, R-tensor_H), (L-tensor_D, R-tensor_D) respectively. The above boundary values are used to define the boundary range for obtaining data from the original tensor, where the boundary values are arbitrary values compared to the size values of the original tensor in this dimension.

[0067] In the method combined Figure 7A with Figure 7B described, the range (i.e., boundary values) of the original tensor is not defined, that is, the tensor to be processed is obtained from the complete data of the original tensor. In the method according to the embodiments of the present disclosure, it is also proposed that the range for obtaining the tensor to be processed from the original tensor can be delimited by setting boundary values. Further, in the implementation process, the boundary values can be set to arbitrary values compared to the size values of the original tensor in this dimension. That is to say, the boundary values can exceed the range of the original tensor itself.

[0068] As an example, Figure 8A shows a schematic diagram of setting boundary values for the original tensor according to the embodiments of the present disclosure. As Figure 8A shown, the rectangular box represents the range covered by the original tensor, which can be any dimension in the original tensor, such as the C dimension, the W dimension, the H dimension, or the D dimension. Generally, boundary values are not set for the N dimension. In the Figure 8A six subgraphs of, the left boundary value and the right boundary value ( Figure 8A shown as "L" and "R" in) are respectively shown in relation to the size values of the original tensor in this dimension.

[0069] According to the embodiments of the present disclosure, when setting boundary values, sequentially obtaining data from the original tensor includes: for the data part of the original tensor covered by the range defined by the boundary values, sequentially obtaining data from the range of the original tensor defined by the boundary values, for example, referring to the Figure 7A and Figure 7B described order. In comparison, for the data part of the original tensor not covered by the range defined by the boundary values, it is represented as invalid data. For the invalid data, the buffer directly fills the tensor to be processed with a predetermined value, where the predetermined value is equal to 0.

[0070] As an example, in Figure 8AIn the first sub - picture, both the left boundary value L and the right boundary value R are on the left side of the original tensor. That is, the data to be obtained in this dimension are all invalid data. In this case, for the invalid data, for example, this part of the data can be automatically filled by sending a zero - filling instruction to the buffer area. Another example is that in Figure 8A In the second sub - picture, the left boundary value L is on the left side of the left boundary of the original tensor data range, and the right boundary value R is on the left side of the right boundary of the original tensor. That is, for the tensor to be processed to be obtained, part of it is invalid data and part of it is valid data in the original tensor. Schematically, in Figure 8A , the data corresponding to the slanted shaded part represents valid data, and the rest are all invalid data. In this case, for the valid data, for example, refer to Figure 7A and Figure 7B for the described order, while for the invalid data, for example, this part of the data can be automatically filled by sending a zero - filling instruction to the buffer area.

[0071] In the method according to the embodiments of the present disclosure, by setting boundary values for the original tensor, the range from which the tensor to be processed is to be taken can be further delimited. In practical applications, this implementation method can adapt to the characteristics of operations such as convolution. For example, it is beneficial to reduce the amount of calculation and greatly improve the flexibility of data. As an example, assume that the original tensor corresponds to an intermediate tensor for feature extraction of an entire input picture, and the input picture includes a specific target, such as an object to be recognized, and the object does not cover the entire picture, that is, the picture includes a background part. In this case, by setting boundary values, the range of the tensor to be processed to be obtained can be limited to the part of the original tensor corresponding to the specific target, so as to reduce the amount of calculation of subsequent operations such as convolution and improve the processing efficiency.

[0072] As an example, Figure 8B shows a schematic diagram of continuous data acquisition in the case where boundaries are set for the original tensor. Among them, compared with the situation shown in Figure 6 , Figure 8B can be understood as setting boundary values (bound) for the C - dimension, W - dimension, and H - dimension of tensor A in Figure 6 , and the boundaries are all within the size ranges of tensor A in the C - dimension, W - dimension, and H - dimension, which is equivalent to the boundary situation shown in the fourth sub - picture in Figure 8A . Specifically, Figure 8B the outer square as a whole corresponds to tensor A (corresponding to tensor A in Figure 6 ). After setting the boundaries therein, according to the data loading method of the embodiments of the present disclosure, data will be sequentially obtained starting from the starting point coordinates within the set boundary range until the number of obtained data is equal to the number of pixels included in the tensor to be processed. InFigure 8B Among them, the data part composed of squares corresponds to the data range framed by the boundary values set for the C dimension, W dimension, and H dimension, and data is sequentially acquired from the data range framed by the boundary. Data located outside the data range framed by the boundary can be regarded as invalid data. Regarding Figure 8B The process of sequentially acquiring data in Figure 6 can be referred to in combination with

[0073] In some embodiments according to the present disclosure, a step value for acquiring data can also be set. As an example, the set data step value is used to specify the stride for sequentially acquiring data from the original tensor. Among them, for the data of the original tensor, starting from the starting coordinate, sequentially acquiring data from the original tensor includes: for the data of the original tensor, starting from the starting coordinate, sequentially acquiring data from the original tensor according to the data step value.

[0074] Figure 8C FIG. shows a schematic diagram of continuously acquiring a tensor to be processed according to the set step in an embodiment of the present disclosure. Among them, the step is equal to 2, that is, one data is taken every other pixel point, that is, only the pixels shown in the shaded part are sequentially acquired, and a step equal to 2 means skipping one pixel. Similarly, when the step is equal to 3, it means skipping two pixels, and so on. It can be understood that when the set step is equal to 1, it corresponds to Figure 7A - Figure 7B the data loading method of each pixel point shown in. In practical applications, by setting the step, the computational amount of subsequent operations such as convolution operations on the data can be further reduced, the processing efficiency can be improved. In addition, the flexibility of data loading can be further increased.

[0075] In the data loading method provided according to the embodiments of the present disclosure, a new data acquisition mode different from the block-based data acquisition method is provided, that is, pixel-by-pixel data loading can be realized according to the starting point and the number of pixels to be acquired (as Figure 6 shown), rather than Figure 4 the block-based acquisition in, this data loading method is more conducive to the calculation process of operations such as convolution operations, and is conducive to improving the operation efficiency. Specifically, the tensor is no longer acquired and transported in a block form, but is sequentially acquired and transported in units of pixel points to be applicable to calculation operations such as convolution operations, improving the data transportation efficiency for such operations, reducing the memory access time, and improving the overall performance of the processor.

[0076] As described above in combination with Figure 6 、 Figure 7A - Figure 7B 、 Figure 8A - Figure 8C, which details the implementation process of the data loading method according to the embodiments of the present disclosure. In this process, data of the original tensor is sequentially retrieved from the original tensor starting from the starting coordinates until the number of retrieved data is equal to the number of pixels (copy_pixel_num) included in the tensor to be processed.

[0077] Next, the implementation process of determining multiple requests for loading the tensor to be processed by combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor in step S102 shown in Figure 5 will be described.

[0078] In an embodiment according to the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension. Data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the data storage format of the original tensor indicates that the first dimension has priority over the second dimension in storage or loading, and the first dimension and the second dimension are adjacent. As an example, when continuous loading cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, requests are divided in the second dimension.

[0079] As an example, when the data storage formats of the original tensor and the tensor to be processed are NDHWC, the first dimension can be the C dimension, the second dimension can be the adjacent W dimension, and the C dimension has priority over the W dimension in storage or loading. Specifically, in response to the size relationship between the tensor to be processed and the original tensor in the C dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the C dimension but can be performed in dimensions lower than the C dimension (in this case, since the C dimension is already the lowest dimension, dimensions lower than the C dimension are no longer considered), requests are divided in the W dimension so that data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. For ease of description, this way of dividing requests can be referred to as the request division rule according to the W dimension (PerW).

[0080] As another example, when the data storage formats of the original tensor and the tensor to be processed are NDHWC, the first dimension can be the W dimension, the second dimension can be the adjacent H dimension, and the W dimension takes precedence over the H dimension during storage or loading. Specifically, in response to the size relationship between the tensor to be processed and the original tensor in the W dimension such that when loading the tensor to be processed, continuous loading cannot be achieved in the W dimension but can be achieved in dimensions lower than the W dimension (in this case, the C dimension is lower than the W dimension, that is, data cannot be continuously loaded in the W dimension but can be continuously loaded in the C dimension lower than the W dimension), the requests are divided in the H dimension so that the data loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor. For ease of description, this way of dividing requests can be referred to as the request division rule according to the H dimension (PerH).

[0081] It can be understood that for the case where the data storage formats of the original tensor and the tensor to be processed are NDHWC, the division rules for other dimensions can be deduced by analogy and will not be repeated here. According to the above-described method, the way of dividing requests can also include: the request division rule according to the D dimension (PerD, that is, continuous loading cannot be achieved in the H dimension but can be achieved in the W and C dimensions lower than the H dimension), the request division rule according to the N dimension (PerN, that is, continuous loading cannot be achieved in the D dimension but can be achieved in the H, W, and C dimensions lower than the D dimension), and the request division rule according to the overall dimension (Per1, that is, continuous loading cannot be achieved in the N dimension but can be achieved in the D, H, W, and C dimensions lower than the N dimension). It can be understood that for the request division rule of Per1, for the operation of loading the tensor to be processed in the original tensor into the buffer, the whole is implemented by one request, that is, these data are loaded into the buffer by one instruction. This is because in this case, the data to be loaded is continuously stored in each dimension and the data loading process can be executed by one instruction. Generally speaking, for the NDHWC case, the way of dividing requests can include 5 types: PerW, PerH, PerD, PerN, Per1.

[0082] For the case where the data storage format of the original tensor and the tensor to be processed is N(C / x)DHW(xC), the rules for request partitioning are similar to those described above for the NDHWC case, except that when the data storage format is N(C / x)DHW(xC), the priority of the C dimension is between the N dimension and the D dimension. From this, it can be obtained that for the data storage format of N(C / x)DHW(xC), the request partitioning methods can include the following schemes: According to the request partitioning rule for the H dimension (PerH), that is, the data in the W dimension at the lowest dimension cannot be continuously loaded; according to the request partitioning rule for the D dimension (PerD), that is, the data in the H dimension cannot be continuously loaded but the data in the W dimension lower than the H dimension can be continuously loaded; according to the request partitioning rule for the C dimension (PerC), that is, the data in the D dimension cannot be continuously loaded but the data in the H dimension and W dimension lower than the D dimension can be continuously loaded; according to the request partitioning rule for the N dimension (PerN), that is, the data in the C dimension cannot be continuously loaded but the data in the D dimension, H dimension, and W dimension lower than the C dimension can be continuously loaded; and according to the request partitioning rule for the overall dimension (Per1), that is, the data in the N dimension cannot be continuously loaded but the data in other dimensions lower than the N dimension can be continuously loaded. In this case, for the C dimension bound to the W dimension, it is required that copy_c = 8 * c, where copy_c represents the size of the tensor to be processed in the C dimension to be obtained.

[0083] As an implementation method, for the data storage format of N(C / x)DHW(xC), there is also a request partitioning rule according to the C dimension, that is, PerC. In this case, the data in the D dimension and the H dimension and W dimension lower than the D dimension can be continuously loaded. However, since xC is bound to the W dimension, when the C dimension is discontinuous or copy_c is greater than xC, the request partitioning rule of PerC is also adopted.

[0084] As another implementation method, as described above, for the data storage format of N(C / x)DHW(xC), the xC dimension is bound to the W dimension, that is, the data itself is continuous in the W dimension. According to the above-described scheme, the request partitioning rule of PerH should be followed. However, there is also a case where when the stride set for the W dimension is greater than 1 (stride_x > 1), for example, as Figure 8C shown, the stride of the W dimension is equal to 2, which will make the data to be obtained for the W dimension discontinuous, that is, one data is obtained every other pixel. Therefore, in this case, the request partitioning rule of PerW needs to be adopted to partition the multiple requests for loading data, rather than following the request partitioning rule of PerH.

[0085] Generally speaking, for the case of N(C / x)DHW(xC), the requested partitioning methods can include six types: PerW, PerH, PerD, PerC, PerN, and Per1.

[0086] It can be understood that, compared with the five requested partitioning methods for the NDHWC case, the requested partitioning rules for the N(C / x)DHW(xC) case are logically similar, and the only difference lies in the priority of the C dimension. This is because for the placement method of N(C / x)DHW(xC), the xC dimension is bound to the W dimension, that is, the data in the lowest xC dimension is continuous.

[0087] According to some embodiments of the present disclosure, the determination of whether continuous loading can be performed in the first dimension is specified as follows: when the first dimension is the channel number dimension, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension. Further, when the first dimension is the channel number dimension and the data storage format of the original tensor is NDHWC, in response to the size (copy_c) of the tensor to be processed in the channel number dimension being equal to the size tensor_c of the original tensor in the channel number dimension, and the starting coordinate (c_coord_b) of the tensor to be processed in the channel number dimension being equal to the starting coordinate of the original tensor in the channel number dimension, it is determined that the size relationship in the first dimension enables continuous loading in the first dimension when loading the tensor to be processed. For other dimensions, the process of determining whether continuous loading can be performed is similar to that of the C dimension, and will not be described one by one here.

[0088] The object to be loaded in the present disclosure, that is, the tensor to be processed, is split into multiple different requests according to a certain rule and sent sequentially, and the returned data is written into the buffer area in order, thereby efficiently loading the tensor to be processed into the memory. A principle during splitting is that the data loaded by each request is either all in the memory or all not in the memory. This is because when the requested data is in the memory, a request is sent to the memory, and when the requested data is not in the memory, the request can be converted into an operation such as writing a predetermined value to the buffer area by the hardware. This splitting method can send corresponding loading requests to different hardware more reasonably.

[0089] In addition, during splitting, if the size relationship between the tensor to be processed and the original tensor in the first dimension causes the tensor to be processed to not be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, then a further division of the requests is performed in the second dimension, so that the requests can be split more reasonably, the data stored continuously in memory can be retained as much as possible, the number of requests can be reduced, the tensor to be processed can be efficiently loaded, the bandwidth when loading or storing data from memory can be significantly increased, the performance can be improved, and the efficiency of the hardware computing unit can be improved.

[0090] In an implementation solution where boundary values are set for the original tensor, the data loading method according to the embodiments of the present disclosure may further include: obtaining boundary values for the original tensor in at least some of the five dimensions respectively, and the boundary values are used to define the boundary range for obtaining data from the original tensor. In some implementation manners, the boundary values may be arbitrary values compared with the size values of the original tensor in this dimension. As an example, the boundary value set for the W dimension may be any of the Figure 8A cases shown. Further, for the dimension where the boundary value is set, the data loaded by each request either all belongs to the data boundary of the original tensor itself and the valid data range defined by the boundary value, or all does not belong to the valid data range defined by both the data boundary of the tensor itself and the set boundary value. As Figure 8A shown, for the first sub-graph, both the left and right boundary values are located on the left side of the data range of the original tensor, that is, both are outside the data range of the original tensor, indicating that the data to be obtained for this dimension is all invalid (for example, represented as out of bound, oob). For Figure 8A the second sub-graph in, the left boundary is located on the left side of the data range of the original tensor, and the right boundary is located within the data range of the original tensor. Thus, only the part of the data that is both within the data range of the original tensor and to the left of the right boundary R is valid data (in Figure 8A , the data corresponding to the diagonal shaded part is represented as valid data), and the rest of the data to be obtained is all invalid data. According to the embodiments of the present disclosure, for the dimension where the boundary value is set, the rule for dividing requests is that the data loaded by each request is either all valid data or all invalid data (i.e., oob), that is, valid data and invalid data cannot be loaded by the same request.

[0091] According to an embodiment of the present disclosure, for multiple requests determined according to step S102, in the actual implementation process of the processor, it can be implemented by setting a state machine. For the state machine, multiple parameters required to determine the above requests can be set, and according to the specific values of the parameters and the initial parameters of the original tensor and the tensor to be processed, the specific tensor to be processed in the original tensor is loaded into the buffer area. The implementation scheme related to determining multiple requests for loading the tensor to be processed will be described below in conjunction with the state machine.

[0092] According to some embodiments of the present disclosure, in step S102, in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, multiple requests for loading the tensor to be processed are determined, including: determining the first request among the multiple requests and the initial state of the first request entering the state machine based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, where the first request indicates the number of pixels to be acquired by the request; and based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, using the state machine to determine each request among the multiple requests after the first request.

[0093] As an example, for the task of loading the tensor to be processed in the original tensor in memory into the buffer area, the first request can be determined first based on the initial parameters. For example, based on the number of pixels copy_pixel_num included in the tensor to be processed, the data storage format of the original tensor (such as NDHWC or N(C / x)DHW(xC)), and the starting coordinates of the tensor to be processed (the starting coordinates indicate the position of the first pixel of the tensor to be processed in the coordinate system of the original tensor, for example, represented as its positions in 5 dimensions, c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b), determine the first request among the multiple requests and the initial state of the first request entering the state machine, where the first request indicates the number of pixels to be acquired by the request (req_size_1).

[0094] Specifically, taking the data storage format as NDHWC, copy_pixel_num = 40, and the starting coordinates are w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0 as an example, in combination with Figure 7A Describe how to determine the partitioning method of the request and the first request. As Figure 7AAs shown, for the above starting coordinates, the starting pixel corresponding to the starting coordinates in the original tensor is obtained, that is, data is obtained starting from this starting pixel. Then, for the NDHWC data storage format, the C dimension is the lowest dimension. Thus, with the C dimension as the first dimension, it is determined whether continuous loading can be performed in the C dimension. In response to the size relationship between the tensor to be processed and the original tensor in the C dimension such that continuous loading cannot be performed in the C dimension when loading the tensor to be processed, the above PerW request partitioning method is adopted, that is, requests are partitioned one by one for each W. That is to say, each pixel is loaded into the buffer by one request. According to the judgment on whether continuous loading can be performed in the C dimension described above: In response to the size of the tensor to be processed in the C dimension (copy_c) not being equal to the size of the original tensor in the C dimension (tensor_c), and / or the first coordinate value (c_coord_b) of the starting coordinates of the tensor to be processed in the C dimension not being equal to the second coordinate value (for example, 0) of the starting coordinates of the original tensor in the C dimension, it is determined that continuous loading cannot be performed in the C dimension for the tensor to be processed. Summarized as follows: It is determined that the data to be obtained is discontinuous in the C dimension when any of the following conditions is satisfied: (1)The starting coordinate c_coord_b of the tensor to be processed in the C dimension is not equal to 0; (2)The size copy_c of the tensor to be processed in the C dimension is greater than the size tensor_c of the original tensor in the C dimension; (3)The size copy_c of the tensor to be processed in the C dimension is less than the size tensor_c of the original tensor in the C dimension; For example, the coordinates of the data loaded by each request are different in the C dimension, but the same in the W dimension, H dimension, D dimension, and N dimension. At this time, it can be understood that without considering the depth direction, the original tensor is unfolded into a one-dimensional vector composed of multiple pixels (pixel) in the order of WHDN. The data loaded by each request is the data within one pixel, that is, requests are partitioned in the W dimension, and the next request can be to load the data in the adjacent next pixel, that is, requests are partitioned in a pixel-skipping manner.

[0095] After determining the partitioning method, for example, in the PerW manner, the corresponding first request can be made with reference to Figure 7A , the first request is used to load the starting pixel corresponding to the starting coordinates, its initial state is equal to the starting coordinates of the tensor to be processed, and the number of pixels to be obtained by this first request is req_size_1 = 1 because subsequent requests are partitioned in a pixel-by-pixel manner. The specific acquisition order is as shown by the arrow in Figure 7A until the number of pixels to be obtained is equal to copy_pixel_num = 40 of the tensor to be processed.

[0096] If continuous loading is possible in the C dimension, then it is determined whether continuous loading is possible in the W dimension adjacent to and higher than the C dimension. If continuous loading is not possible in the W dimension and continuous loading is possible in the C dimension lower than the W dimension, the above-mentioned request partitioning method of PerH is adopted, that is, the requests are partitioned row by row. That is to say, each row of pixels is loaded into the buffer by one request.

[0097] Next, still taking the data storage format as NDHWC and the second coordinate value as 0 as an example, assuming that continuous loading is possible in the C dimension, then, with the first dimension as the W dimension and the second dimension as the H dimension for judgment, it is determined that the W dimension is discontinuous when any of the following conditions is satisfied: (1) The starting coordinate of the tensor to be processed in the W dimension is not equal to 0; (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension; (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension.

[0098] At this time, when the C dimension is continuous, the coordinates of the data to be loaded by each request are different in the C dimension and the W dimension, but the coordinates in the H dimension, D dimension, and N dimension are the same. At this time, it can be understood that the data loaded by each request belongs to the same row, that is, the requests are partitioned in the H dimension, and the data loaded by different requests is located in different rows.

[0099] Assuming that the first dimension is the H dimension and the second dimension is the D dimension, it is determined that the H dimension is discontinuous when any of the following conditions is satisfied: (1) The starting coordinate of the tensor to be processed in the H dimension is not equal to 0; (2) The size copy_h of the tensor to be processed in the H dimension is greater than the size tensor_h of the original tensor in the H dimension; (3) The size copy_h of the tensor to be processed in the H dimension is less than the size tensor_h of the original tensor in the H dimension.

[0100] At this time, when both the C dimension and the W dimension are continuous, the coordinates of the data to be loaded by each request are different in the C dimension, W dimension, and H dimension, but the coordinates in the D dimension and N dimension are the same.

[0101] Assuming that the first dimension is the D dimension and the second dimension is the N dimension, it is determined that the D dimension is discontinuous when any of the following conditions is satisfied: (1) The starting coordinate of the tensor to be processed in the D dimension is not equal to 0; (2) The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension; (3) The size copy_d of the tensor to be processed in dimension D is less than the size tensor_d of the original tensor in dimension D.

[0102] At this time, when dimensions C, W, H, and D are all continuous, the coordinates of the data to be loaded for each request are different in dimensions C, W, H, and D, but the coordinates in dimension N are the same, that is, the requests are partitioned in dimension D.

[0103] Assume that the first dimension is dimension N. Determine that it is discontinuous in dimension N when any of the following conditions holds: (1) The starting coordinate of the tensor to be processed in dimension N is not equal to 0; (2) The size copy_n of the tensor to be processed in dimension N is greater than the size tensor_n of the original tensor in dimension N; (3) The size copy_n of the tensor to be processed in dimension N is less than the size tensor_n of the original tensor in dimension N.

[0104] At this time, when dimensions C, W, H, and D are all continuous, the coordinates of the data to be loaded for each request are the same in dimensions C, W, H, D, and N. For the tensor data located in memory, in essence, at this time, one request can be sent to the memory to load the data in the memory, and other requests can be used to fill the data that is not located in the memory, that is, invalid data.

[0105] The request partitioning logic for the tensor to be processed with the data storage format of N(C / x)DHW(xC) is the same as that of NDHWC. The difference is that the dimension arrangement of N(C / x)DHW(xC) is different from that of NDHWC. For N(C / x)DHW(xC), due to its special interleaved structure, it is considered to be necessarily continuous in (xC) by default. Therefore, the lowest dimension is considered to be dimension W, followed by dimension H, then dimension D, then dimension C, and the highest is dimension N.

[0106] For example, taking the data storage format of N(C / x)DHW(xC) and the second coordinate value being 0 as an example, assume that the first dimension is dimension D and the second dimension is dimension C. Determine that it is discontinuous in dimension D when any of the following conditions holds: (1) The starting coordinate of the tensor to be processed in dimension D is not equal to 0; (2) The size copy_d of the tensor to be processed in dimension D is greater than the size tensor_d of the original tensor in dimension D; (3) The size copy_d of the tensor to be processed in dimension D is less than the size tensor_d of the original tensor in dimension D.

[0107] At this time, when both the W dimension and the H dimension are continuous, the coordinates of the data to be loaded for each request are different in the W dimension, the H dimension, and the D dimension, but the coordinates in the N dimension are the same. That is, the requests are partitioned in the C dimension.

[0108] In addition, for N(C / x)DHW(xC), due to the particularity of the interleaving pattern (xC), the pixels are continuous. Therefore, the first dimension is W, H, D, C, N in ascending order. For the conditions under which data cannot be continuously loaded and the request partitioning principle when the data storage format is N(C / x)DHW(xC), reference can be made to the relevant content of NDHWC, which will not be specifically listed here one by one.

[0109] As described above for NDHWC, for the N(C / x)DHW(xC) format, after determining the partitioning method, for example, in the PerH manner, the first request can be determined similarly. Refer to Figure 7B , the first request is used to load the row of pixels where the starting pixel corresponding to the starting coordinates is located. Its initial state is equal to the starting coordinates of the tensor to be processed, and the number of pixels to be obtained by this first request is req_size_1 = 2 because the subsequent requests are partitioned row by row. The specific acquisition order is as shown by the arrows in Figure 7B until the number of pixels to be obtained is equal to copy_pixel_num = 40 of the tensor to be processed.

[0110] According to an embodiment of the present disclosure, after determining the requests, it is further necessary to determine the data read address of the starting position for the requests to read data from the memory, which is used to indicate the data write address of the starting position for writing data to the buffer. Among them, based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, determining the first request among multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; taking the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and determining the number of pixels to be obtained by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0111] According to some embodiments of the present disclosure, when the data storage format of the original tensor is NDHWC, for multiple requests for loading the tensor to be processed, determining the data reading address of the nth request among the multiple requests includes: Calculating the data reading address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord), where u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in five dimensions.

[0112] According to some embodiments of the present disclosure, when the data storage format of the original tensor is N(C / x)DHW(xC), for multiple requests for loading the tensor to be processed, determining the data reading address of the nth request among the multiple requests includes: Calculating the data reading address Addr2_n of the nth request according to the following formula: Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC where u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the original tensor in five dimensions, and xC represents the size occupied by x channels.

[0113] According to some embodiments of the present disclosure, for multiple requests for loading tensors to be processed, determining the data write address Addr3_n of the nth request among the multiple requests includes: Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 Wherein, b_addr_base is the starting address for writing the tensor to be processed in the buffer, and req_size_1, req_size_2, …, req_size_n-1 represent the number of pixels to be acquired by the first n-1 requests respectively.

[0114] For the first request sent, its corresponding request initial coordinate is the starting coordinate of the tensor to be processed. Thus, the data read address of the first request can be calculated with reference to the above formula. The starting address for writing the tensor to be processed in the buffer is used as the data write address of the first request.

[0115] The data length loaded by the first request can be determined according to parameters such as the determined request partitioning method, the starting coordinate of the tensor to be processed, and the size of the original tensor. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is less than the second coordinate value (i.e., the coordinate value of the starting coordinate of the original tensor in the first dimension, for example, is 0), the data length loaded by the first request is the shape size copy_t of the tensor to be processed in the first dimension. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value (i.e., the coordinate value of the starting coordinate of the original tensor in the first dimension, for example, is 0), the data length loaded by the first request is the absolute value of the first coordinate value t_coord_b.

[0116] Considering that the state machine has the advantages of a clear logical structure, being easy to maintain and expand, and being particularly suitable for processing logical scenarios with multiple conditions and multiple branches, in the present disclosure, the state machine is adopted to automatically update the request initial coordinates and the data length loaded corresponding to each request, avoiding complex conditional nesting, with clear logic, being easy to maintain, and having strong scalability.

[0117] According to some embodiments of the present disclosure, determining an initial state in a state machine for a first request based on a starting coordinate of a tensor to be processed includes: in response to a first coordinate value of the starting coordinate of the tensor to be processed in a first dimension being less than a second coordinate value of the starting coordinate of the original tensor in the first dimension, determining the initial state as a first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, determining the initial state as a second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as a third state.

[0118] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, using the state machine to determine each request after the first request among multiple requests includes: based on the initial state, using the state machine to determine a request initial coordinate corresponding to a second request after the first request and the number of pixels to be acquired by the second request; determining a data read address of the second request according to the shape size of the original tensor, the request initial coordinate corresponding to the second request, and the data storage format of the original tensor; determining a data write address of the second request according to the number of pixels to be acquired by the second request; updating the state machine according to information related to the second request, and sequentially determining subsequent requests in each request based on the updated state machine.

[0119] Figure 9A A state diagram of a state machine according to some embodiments of the present disclosure is shown.

[0120] For example, the states of the state machine include a first state s0, a second state s1, and a third state s2, and the initial state of entering the state machine is determined by the starting coordinates c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b of the tensor to be processed.

[0121] For example, in response to a first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in a first dimension being less than a second coordinate value of the starting coordinate of the original tensor in the first dimension, determining the initial state as a first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, determining the initial state as a second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as a third state.

[0122] Reference Figure 9A, the tensor_t direction represents the data in the first dimension. The data with the coordinate in the t dimension less than the second coordinate value and greater than the third coordinate value is not in memory or not within the range of the original tensor data ( Figure 9A The part shown by the diagonal shading, and this part of the data can be called invalid data), and the data with the size of the t dimension between the second coordinate value and the third coordinate value ( Figure 9A The white part) is stored in memory, which is the actual size of the original tensor data in the first dimension.

[0123] The state only switches to itself and adjacent states. For example Figure 9A in it, the first state s0 can jump to the first state s0 or the second state s1, the second state s1 can jump to the first state s0, the second state s1 and the third state s2, and the third state s2 can jump to the first state s0, the second state s1 and the third state s2.

[0124] Figure 9A The 6 cases (① to ⑥) in

[0125] show the state transition changes that the tensor to be processed experiences for different sizes in the first dimension. Figure 9A For example,

[0126] in case ①, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the second coordinate value. The state machine always loops and jumps in the first state s0 on the left. Figure 9A For example,

[0127] in case ②, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value and less than the third coordinate value. The state machine loops and jumps between the first state s0 and the second state s1. Figure 9A For example,

[0128] in case ③, the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value. The state machine always loops and jumps in the second state s1. Figure 9A For example,

[0129] in case ④, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1 and the third state s2. Figure 9AIn case ⑤, the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops between the second state s1 and the third state s2.

[0130] For example, Figure 9A In case ⑥, the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops in the third state s2.

[0131] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and combine with the trigger condition to determine the state entered by the second sent request. Then, based on the state entered by the second sent request, determine the request initial coordinates corresponding to the third sent request and the data length loaded by the second sent request, and combine with the trigger condition to determine the state entered by the third sent request, and so on.

[0132] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request in each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state. After determining the request initial coordinates corresponding to the current request, it can be judged whether the data loaded by the current request is in the memory according to the request initial coordinates.

[0133] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the data to be loaded by the current request is in the memory. It can be understood that for the data range that belongs to the original tensor, it will all be within the memory range (this part of the data is actually stored in the memory), and for the data range that does not belong to the original tensor, it will all not be within the memory range and is not the actually stored data. That is, the data exceeding the original tensor range can be understood as invalid data. For example, this part of the invalid data can be directly filled with zeros in the buffer. At this time, the data read address and data write address of the current request can be determined with reference to the above formula, and then the current request is sent to the memory to load the corresponding data into the buffer.

[0134] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the data to be loaded by the current request is not in the memory. At this time, in response to all the data to be loaded by the current request not being in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values into the buffer, where the number of the multiple predetermined values is determined by the data length specified by the current request. For example, the predetermined value is 0.

[0135] The process of using the state machine to determine the request initial coordinates and the data length loaded by each request will be specifically described below.

[0136] Figure 9B It shows the state transition of the data storage format for NDHWC in the PerW division manner. The condition for the PerW division manner is that c_coord is not equal to 0, or copy_c is not equal to tensor_c. At this time, the two adjacent pixels to be obtained are not continuous in memory, and split requests for each pixel are required to obtain the data.

[0137] In Figure 9B on the left side, only two states, s1 and s2, are shown for the C dimension. This is because in practical applications, generally the coordinate of the C dimension is not less than 0, that is, there is no first state. Specifically, s1 can correspond to the above-mentioned second state, and s2 can correspond to the above-mentioned third state. In addition, in Figure 9B the state transition schematic diagram on the right side, there is also an idle state, which is used to represent the idle state before data loading or after loading all the tensors to be processed. Table 1 below shows the meanings corresponding to each state: Table 1

[0138] That is to say, in the example of Figure 9B the blank box on the right corresponds to the size of the original tensor in the C dimension (tensor_c). Its left boundary can correspond to, for example, Figure 9A the second coordinate value in Figure 9B which is equal to 0, for example, and its right boundary can correspond to, for example,

[0139] Figure 9B the third coordinate value in Table 2

[0140] In Table 2 above, the parameter remain_copy_p represents the number of pixels that have not been fetched yet. For example, before the first request, remain_copy_p is equal to the number of pixels copy_pixel_num corresponding to the tensor to be processed. After the first request, remain_copy_p is equal to copy_pixel_num - req_size_1. remain_copy_c represents the size of the data that has not been fetched yet in the C dimension, which is used to characterize the data fetching information in the C dimension. For example, before the first request, remain_copy_c = copy_c.

[0141] The following describes the information that needs to be updated in each state and how to update it.

[0142] First, in the s1 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table: Table 3

[0143] It should be noted that in the example of Table 3 above, the situation of setting boundary values and data fetching step sizes for the original tensor is further considered. Among them, right_bound represents the right boundary value set for the W dimension, down_bound represents the boundary value in the H dimension, and hind_bound represents the boundary value in the D dimension. The setting of boundary values and step sizes can refer to the description in combination with Figure 8A - Figure 8C above. In addition, for the case where no boundary value is set, it can be understood that each boundary is equal to the size of the original tensor in that dimension. In Table 3, stride_x, stride_y, and stride_z respectively represent the step size values in the W dimension, H dimension, and D dimension. As an implementation, the step sizes can all be set to 1. In addition, it can be understood that in other embodiments according to the present disclosure, boundary values and step sizes can also be set for the C dimension, for example, and there is no limitation here.

[0144] The parameter update rules in Table 3 are described below. First, for the C dimension with the lowest priority, the parameter c_coord represents the coordinate of the current request in the C dimension. When c_coord + remain_copy_c <= tensor_c, that is, the next request for the C dimension does not exceed the boundary tensor_c of the original tensor in the C dimension, the parameter c_coord is updated to the starting coordinate c_coord_b of the C dimension; when c_coord + remain_copy_c > tensor_c, that is, the data acquisition for the next request in the C dimension exceeds the boundary tensor_c of the original tensor in the C dimension, the parameter c_coord is updated to tensor_c.

[0145] The coordinate update logic for the other dimensions in Table 3 is the same as that for the W dimension, which is described separately as follows. For the W dimension, the parameter w_coord represents the coordinate of the current request in the W dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be acquired by the next request is still valid data, when w_coord + stride_x >= right_bound, that is, when the W dimension exceeds the boundary, the coordinate w_coord of the W dimension is updated to w_coord + stride_x - copy_w, where copy_w represents the size of the tensor to be processed in the W dimension; otherwise (w_coord + stride_x < right_bound), that is, when it does not exceed the boundary, the coordinate w_coord of the W dimension is updated to w_coord + stride_x. When stride_x = 1, it means that w_coord is incremented by 1, that is, the data to be acquired by the next request is the next data immediately following the data to be acquired by the current request. Otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate w_coord of the W dimension remains unchanged.

[0146] For the H dimension, the parameter h_coord represents the coordinate of the current request in the H dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be fetched by the next request is still valid data, at this time w_coord + stride_x >= right_bound. When h_coord + stride_y >= down_bound, that is, when the H dimension exceeds the boundary, the coordinate h_coord of the H dimension is updated to h_coord + stride_y - copy_h, where copy_h represents the size of the tensor to be processed in the H dimension; otherwise (h_coord + stride_y < down_bound), that is, when it does not exceed the boundary, the coordinate h_coord of the H dimension is updated to h_coord + stride_y. Otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate h_coord of the H dimension remains unchanged.

[0147] For the D dimension, the parameter d_coord represents the coordinate of the current request in the D dimension. For the case of jumping from state s1 to s1, that is, the case where the data to be fetched by the next request is still valid data, at this time w_coord + stride_x >= right_bound and h_coord + stride_y >= down_bound. When d_coord + stride_z >= hind_bound, that is, when the D dimension exceeds the boundary, the coordinate d_coord of the D dimension is updated to d_coord + stride_z - copy_d, where copy_d represents the size of the tensor to be processed in the D dimension; otherwise (d_coord + stride_z < hind_bound), that is, when it does not exceed the boundary, the coordinate d_coord of the D dimension is updated to d_coord + stride_z. Otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate d_coord of the D dimension remains unchanged.

[0148] For the N dimension, the parameter n_coord represents the coordinates of the current request in the N dimension. For the case of jumping from state s1 to s1, that is, when the data to be retrieved by the next request is still valid data, at this time w_coord + stride_x >= right_bound and h_coord + stride_y >= down_bound and d_coord + stride_z >= hind_bound, the coordinate n_coord of the N dimension is incremented by 1, that is, updated to n_coord + 1; otherwise (for cases other than jumping from state s1 to s1, that is, jumping from state s1 to state s2 or idle), the coordinate n_coord of the N dimension remains unchanged. It can be seen that the update logic of the N dimension is slightly different from that of the W, H, and D dimensions because generally, no boundary values and step sizes are set for the N dimension.

[0149] In addition, in addition to the above dimension coordinate parameters, the state machine also needs to maintain two parameters remain_copy_p and remain_copy_c. Among them, remain_copy_p represents the number of remaining pixels not yet retrieved. For example, before the first request, remain_copy_p is equal to the number of pixels corresponding to the tensor to be processed, copy_pixel_num. After the first request, remain_copy_p is equal to copy_pixel_num - req_size_1; remain_copy_c represents the size of the remaining data not yet retrieved in the C dimension. For example, before the first request, remain_copy_c = copy_c. After the first request, remain_copy_c is equal to copy_c minus the size corresponding to the C dimension retrieved by the first request.

[0150] Specifically, as shown in Table 3, for remain_copy_c, when c_coord + remain_copy_c <= tensor_c, remain_copy_c is updated to copy_c; otherwise (c_coord + remain_copy_c > tensor_c), remain_copy_c is updated to c_coord + remain_copy_c - tensor_c.

[0151] For remain_copy_p, when c_coord + remain_copy_c <= tensor_c, remain_copy_p is updated to remain_copy_p - 1; otherwise (c_coord + remain_copy_c > tensor_c), remain_copy_p remains unchanged.

[0152] Second, in the s2 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table: Table 4

[0153] For Table 4, the state updates of each parameter can refer to the description for Table 3 and will not be elaborated here. Based on the above update logic, the update results of Table 4 can be derived similarly.

[0154] The state jumps of the data storage format of NDHWC according to the PerW division method are described above in combination with Tables 1 - 4, which involve states s1 and s2, as well as the update outputs of the state machine in these two states. It can be understood that the principles and update logics of the state machine for the data storage format of NDHWC according to other division methods (e.g., PerH) are similar to those described above and will not be described here.

[0155] Of course, as the coordinates continue to increase, when the number of acquired pixels is equal to the number of pixels included in the tensor to be processed, the state machine can enter the idle state, indicating that the data of the current pen instruction has been acquired. Taking the data storage format of NDHWC as an example, for the PerW request division method, when the remaining values of the N, D, H, and W dimensions are equal to 1 and the data of the C dimension has been acquired, the idle state is entered; for the PerH request division method, when the remaining values of the N, D, and H dimensions are equal to 1 and the data of the W dimension has been acquired, the idle state is entered; for the PerD request division method, when the remaining values of the N and D dimensions are equal to 1 and the data of the H dimension has been acquired, the idle state is entered; for the PerN request division method, when the remaining value of the N dimension is equal to 1 and the data of the D dimension has been acquired, the idle state is entered; for the Per1 request division method, when the data of the N dimension has been acquired, the idle state is entered.

[0156] The state transition conditions of the specific state machine and the updated coordinate content can be changed and set according to actual needs, and the logic is similar to the state machine logic described above, and no further examples will be given here.

[0157] It should be noted that the process of using the state machine to determine the request initial coordinates and the loaded data length corresponding to the request is basically the same for NDHWC and N(C / x)DHW(xC). The difference is that in response to the data storage format of the tensor to be processed being NDHWC, the coordinate value of the channel number dimension is incremented by 1 when updated, and in response to the data storage format of the tensor to be processed being N(C / x)DHW(xC), the coordinate value of the channel number dimension is incremented by x when updated.

[0158] In the above embodiments, requests are divided according to whether tensors can be continuously loaded in the first dimension. Each request is used to load sub-data that is either entirely in memory or entirely not in memory. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby enhancing the hardware utilization rate of the computing unit and improving the hardware performance.

[0159] According to some embodiments of the present disclosure, another data loading method is provided. Figure 10 The schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure is shown. As Figure 10 shown, the data loading method according to the embodiments of the present disclosure includes step S201 and step S202.

[0160] In step S201: Receive a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in memory to a buffer area.

[0161] According to the embodiments of the present disclosure, the data storage format of the original tensor in memory is the same as the data storage format of the tensor to be processed in the buffer area. The data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, depth dimension, height dimension, width dimension, and number of channels dimension. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers.

[0162] In step S202: After parsing the data loading instruction, use an execution unit to execute the data loading instruction.

[0163] As Figure 10 shown, wherein, step S202 uses an execution unit to execute the data loading instruction, including: S2021: Obtain the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; S2022: Combine the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed. Among them, the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and S2023: Send multiple requests in sequence, and write the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer.

[0164] For the related descriptions of the tensor to be processed and the original tensor, as well as the specific implementation processes of steps S2021 - S2023, reference can be made to the related descriptions of the foregoing data loading method, and repeated parts will not be elaborated.

[0165] As an example, the data loading instruction can be a machine instruction, or the data loading instruction can also be a micro-instruction. For example, the data loading instruction is implemented in the form of a Load instruction.

[0166] According to some embodiments of the present disclosure, a data storage method is also provided, which is used to obtain a second tensor based on the first tensor in the buffer and write the second tensor into the memory. It can be understood that the data storage method according to the embodiments of the present disclosure can be understood as the reverse process of the data loading method described above, that is, moving data from the buffer to the memory. The implementation principle thereof according to the embodiments of the present disclosure is similar to the above data loading method, and the repeated parts will not be described again, and only the different parts will be described in detail.

[0167] In the data storage method according to the embodiments of the present disclosure, the first tensor is a 5D tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. Among them, for the 5D tensor of the first tensor, the data determined by the height dimension and the width dimension represents a pixel, and it accumulates gradually to higher dimensions. Among them, the number of channels dimension does not calculate the number of pixels. As an example, the first tensor may refer to the tensor stored in the buffer, and it can be any one of the data storage formats of NDHWC or N(C / x)DHW(xC), and there is no limitation thereto. In terms of understanding the implementation principle, it can be correspondingly understood as the original tensor in the data loading method described above, that is, obtaining part or all of the data from the first tensor and transferring it to the memory. This part of the tensor obtained from the first tensor is represented as the second tensor, which can be correspondingly understood as the tensor to be processed in the data loading method described above.

[0168] Figure 11 Shows a schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure, as Figure 11 shown, the data storage method includes steps S301 - S303.

[0169] In step S301, obtain the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor. Then, in step S302, combine the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine multiple requests for storing the second tensor, where the multiple requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor. In step S303, sequentially send the multiple requests and sequentially write the data corresponding to each request into the memory to store the second tensor into the memory.

[0170] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, divide the requests in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, where the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0171] In the data storage method according to an embodiment of the present disclosure, the manner of sequentially obtaining data from the first tensor can be referred to and combined with Figure 6 the description manner shown. Compared with the block-based data storage method in the related art, the data storage method provided by the embodiment of the present disclosure can achieve pixel-by-pixel data storage, rather than Figure 4 the block-based acquisition in . This data storage method is more beneficial to calculation processes such as convolution operations and is conducive to improving the operation efficiency. It can be understood that, for example, the processor can reasonably use according to the type of operation to be performed or the data processing characteristics whether to use the continuous data loading method or the block-based data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0172] In actual applications, for example, for convolution operations in general computing operations, the operation process is multi-layered. For example, after the first layer is processed, the data result needs to be stored for use when performing the next layer of processing. As described above, for data that needs to perform convolution operations, the sequential data acquisition method provided according to the embodiment of the present disclosure is adopted (which can be referred to the combination above with Figure 5 、 Figure 6 、Figure 7A , Figure 7B , Figure 8A , Figure 8B and Figure 8C The description (as described) is more conducive to improving the computing efficiency. As an example, the weight data of 1x3 can only obtain data in the form of a sliding window, so it can only be obtained row by row, unless the shape of the graph to be obtained is very regular, exactly a cube and no operations such as padding are required for the next layer.

[0173] According to some embodiments of the present disclosure, a plurality of tensors are stored in the buffer area. The first tensor is one of the plurality of tensors. The data storage format of the plurality of tensors in the buffer area can be any one of NDHWC or N(C / x)DHW(xC). Among them, for the data storage format of NDHWC, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension; for the data storage format of N(C / x)DHW(xC), N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension. Among them, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions. Among them, the number of channels dimension does not calculate the number of pixels. Generally, x can be set to an integer multiple of 4.

[0174] According to some embodiments of the present disclosure, in the case where the first dimension is the number of channels dimension, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the starting coordinate of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the first dimension, it is determined that the second tensor cannot be continuously obtained in the first dimension.

[0175] According to some embodiments of the present disclosure, in the case where the first dimension is the number of channels dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the number of channels dimension being equal to the size of the original tensor in the number of channels dimension, and the starting coordinate of the tensor to be processed in the number of channels dimension being equal to the starting coordinate of the original tensor in the number of channels dimension, it is determined that the size relationship in the first dimension enables continuous loading in the first dimension when loading the tensor to be processed.

[0176] According to some embodiments of the present disclosure, determining a plurality of requests for storing a second tensor in combination with the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor includes: determining a first request among the plurality of requests and an initial state in which the first request enters a state machine based on the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor, where the first request indicates the number of pixels to be acquired by the request; and determining each request after the first request among the plurality of requests using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request, the number of pixels included in the second tensor, and the data storage format of the first tensor.

[0177] According to some embodiments of the present disclosure, each request includes a data read address for indicating a starting position for reading data from a buffer, a data write address for indicating a starting position for writing data to memory, and the number of pixels to be acquired by the request. Among them, determining a first request among the plurality of requests and an initial state in which the first request enters a state machine based on the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor includes: determining the initial state in which the first request enters the state machine based on the starting coordinates of the second tensor; using the starting coordinates of the second tensor as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the first tensor, the request initial coordinates corresponding to the first request, and the data storage format of the second tensor; determining the starting address for writing the second tensor in memory as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the second tensor, the data storage format of the first tensor, the starting coordinates of the second tensor, and the shape and size of the first tensor.

[0178] According to some embodiments of the present disclosure, determining the initial state in which the first request enters the state machine based on the starting coordinates of the second tensor includes: in response to the first coordinate value of the starting coordinates of the second tensor in the first dimension being less than the second coordinate value of the starting coordinates of the first tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the first tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0179] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired in the first request, the number of pixels included in the second tensor, and the data storage format of the first tensor, a state machine is used to determine each request after the first request among a plurality of requests, including: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired in the second request; determining the data read address of the second request according to the shape and size of the first tensor, the request initial coordinates corresponding to the second request, and the data storage format of the first tensor; determining the data write address of the second request according to the number of pixels to be acquired in the second request; and updating the state machine based on the information related to the second request, and sequentially determining subsequent requests among each request based on the updated state machine.

[0180] According to some embodiments of the present disclosure, when the data storage format of the first tensor is NDHWC, for a plurality of requests for storing the second tensor, determining the data read address of the nth request among the plurality of requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord), where u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the first tensor in five dimensions.

[0181] According to some embodiments of the present disclosure, when the data storage format of the first tensor is N(C / x)DHW(xC), for a plurality of requests for storing the second tensor, determining the data read address of the nth request among the plurality of requests includes: Calculating the data read address Addr2_n of the nth request according to the following formula: Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Wherein, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, and n_coord, d_coord, h_coord, w_coord, c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape dimensions of the first tensor in five dimensions.

[0182] According to some embodiments of the present disclosure, for multiple requests for storing a second tensor, determining the data write address Addr3_n of the nth request among the multiple requests includes: Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n - 1 Wherein, b_addr_base is the starting address for writing the second tensor in the memory, and req_size_1, req_size_2,..., req_size_n - 1 represent the number of pixels to be acquired by the first n - 1 requests respectively.

[0183] It can be understood that the data storage method according to the embodiments of the present disclosure can achieve similar technical effects to the data loading method according to the embodiments of the present disclosure. Figure 8A For example, the boundary value can be any value compared to the dimension size of the first tensor in this dimension. As an example, the boundary value set for the W dimension can be any of the

[0184] situations shown. Further, for the dimension with the boundary value set, the data loaded by each request either all belongs to the valid data range defined by the data boundary of the first tensor itself and the boundary value, or all does not belong to the valid data range defined by both the data boundary of the tensor itself and the set boundary value.

[0185] According to some embodiments of the present disclosure, another data storage method is provided. Figure 12 The schematic flowchart of the data storage method provided by at least one embodiment of the present disclosure is shown. As Figure 12As shown, the data storage method according to an embodiment of the present disclosure includes step S401 and step S402.

[0186] In step S401: Receive a data storage instruction indicating to execute obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into the memory.

[0187] According to an embodiment of the present disclosure, the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory. The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape size of the second tensor is represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, the depth dimension, the height dimension, the width dimension, and the number of channels dimension. The shape size of the first tensor is represented by b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers.

[0188] In step S402: After parsing the data storage instruction, use the execution unit to execute the data storage instruction.

[0189] As Figure 12 shown, wherein step S402 uses the execution unit to execute the data storage instruction, including: S4021: Obtain the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; S4022: Combine the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine multiple requests for storing the second tensor. Among them, the multiple requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and S4023: Sequentially send multiple requests and write the data corresponding to each request into the memory in sequence to store the second tensor into the memory.

[0190] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be achieved in the first dimension but can be achieved in dimensions lower than the first dimension, partitioning of requests is performed in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer area, wherein the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0191] Regarding the related descriptions of the first tensor and the second tensor and the specific implementation processes of steps S4021 - S4023, reference can be made to the related descriptions of the foregoing data storage method, and repeated parts will not be elaborated.

[0192] As an example, the data storage instruction can be a machine instruction, or the data storage instruction can also be a micro-instruction. For example, the data storage instruction is implemented in the form of a Store instruction.

[0193] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. Among them, the instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the data loading instruction instructs to load a tensor to be processed from an original tensor in a memory into a buffer area, wherein the data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the buffer area, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in a storage component. The shape and size of the tensor to be processed are represented by a1, a2, a3, a4, a5, and a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape and size of the original tensor are represented by b1, b2, b3, b4, b5, and b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers; and the execution unit is configured to: execute the data loading instruction.

[0194] According to an embodiment of the present disclosure, an execution unit executes a data loading instruction, including: obtaining the number of pixels included in a tensor to be processed, the data storage format of an original tensor, and the starting coordinates of the tensor to be processed in a coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to write the data corresponding to each request into a buffer in sequence to load the tensor to be processed into the buffer.

[0195] According to an embodiment of the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the data storage format of the original tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0196] For example, there are multiple tensors stored in memory, and the original tensor is one of the multiple tensors. The data storage format of the multiple tensors in memory can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. Among them, for the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension; for the N(C / x)DHW(xC) data storage format, N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0197] For example, the tensor to be processed is used for a convolution operation in a computing unit within a processor. It can be understood that the above-mentioned processor can be reasonably used according to the type of operation to be performed or the data processing characteristics, whether it is a continuous data loading method or a block data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0198] For example, in the case where the first dimension is the channel number dimension, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0199] For example, in the case where the first dimension is the channel number dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the channel number dimension being equal to the size of the channel number dimension of the original tensor, and the starting coordinate of the tensor to be processed in the channel number dimension being equal to the starting coordinate of the original tensor in the channel number dimension, it is determined that the size relationship in the first dimension enables continuous loading in the first dimension when loading the tensor to be processed.

[0200] For example, in the case where the data storage format of the original tensor is NDHWC, for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord), where u_addr_base represents the storage address in the memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in 5 dimensions.

[0201] For example, in the case where the data storage format of the original tensor is N(C / x)DHW(xC), for multiple requests for loading the tensor to be processed, determining the data read address of the nth request among the multiple requests includes: Calculate the data read address Addr2_n of the nth request according to the following formula: Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Wherein, u_addr_base represents the storage address in the memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in five dimensions.

[0202] For example, for multiple requests for loading tensors to be processed, determining the data write address Addr3_n of the nth request among the multiple requests includes: Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 Wherein, b_addr_base is the starting address for writing the tensor to be processed in the buffer, and req_size_1, req_size_2, ..., req_size_n-1 represent the number of pixels to be obtained by the first n-1 requests respectively.

[0203] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, where the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in memory, and the data storage format is used to indicate the storage order and dimension arrangement of the tensor in the storage component. The shape and size of the second tensor are represented by a1, a2, a3, a4, a5, where a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in 5 dimensions and are all positive integers. The 5 dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a number of channels dimension. The shape and size of the first tensor are represented by b1, b2, b3, b4, b5, where b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in 5 dimensions and are all positive integers; and the execution unit is configured to: execute the data storage instruction.

[0204] According to an embodiment of the present disclosure, the execution unit executes the data storage instruction, including: obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for storing the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into memory to store the second tensor into memory.

[0205] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when storing the second tensor, continuous acquisition cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data obtained by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data obtained by the request all belonging to the data range of the first tensor, the data obtained by the request comes from the first tensor and is continuously stored in the buffer, where the data storage format of the first tensor indicates that the first dimension has priority over the second dimension during storage or loading, and the first dimension and the second dimension are adjacent.

[0206] For example, there are multiple tensors stored in a buffer. The first tensor is one of the multiple tensors. The data storage format of the multiple tensors in the buffer can be any one of NDHWC or N(C / x)DHW(xC). The data storage format is used to indicate the storage order and dimension arrangement of the tensors in the storage component. Among them, for the NDHWC data storage format, N corresponds to b1, indicating the batch dimension; D corresponds to b2, indicating the depth dimension; H corresponds to b3, indicating the height dimension; W corresponds to b4, indicating the width dimension; C corresponds to b5, indicating the number of channels dimension. For the N(C / x)DHW(xC) data storage format, N corresponds to b1, indicating the batch dimension; (C / x) corresponds to b2, indicating the number of channels dimension; D corresponds to b3, indicating the depth dimension; H corresponds to b4, indicating the height dimension; W(xC) corresponds to b5, indicating the width dimension. In the N(C / x)DHW(xC) data storage format, x is a positive integer, and the number of channels of xC is bound to the width dimension.

[0207] For example, in the case where the first dimension is the number of channels dimension, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the first dimension, it is determined that the second tensor cannot be continuously fetched in the first dimension.

[0208] For example, in the case where the first dimension is the number of channels dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the number of channels dimension being equal to the size of the original tensor in the number of channels dimension, and the starting coordinate of the tensor to be processed in the number of channels dimension being equal to the starting coordinate of the original tensor in the number of channels dimension, it is determined that the size relationship in the first dimension enables continuous loading in the first dimension when loading the tensor to be processed.

[0209] For example, in the case where the data storage format of the first tensor is NDHWC, for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + ((((n_coord * tensor_d + d_coord) * tensor_h + h_coord) * tensor_w + w_coord) * tensor_c + c_coord), Among them, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, and c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, and tensor_c represent the shape dimensions of the first tensor in five dimensions.

[0210] For example, in the case where the data storage format of the first tensor is N(C / x)DHW(xC), for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr2_n of the nth request according to the following formula: Addr2_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Among them, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, and c_coord represent the initial coordinates of the request corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, and tensor_c represent the shape dimensions of the first tensor in five dimensions.

[0211] For example, for multiple requests for storing the second tensor, determining the data write address Addr3_n of the nth request among the multiple requests includes: Addr3_n = b_addr_base + req_size_1 + req_size_2 + … + req_size_n-1 Among them, b_addr_base is the starting address for writing the second tensor in the memory, and req_size_1, req_size_2, ..., req_size_n-1 represent the number of pixels to be obtained by the first n-1 requests respectively.

[0212] As an example, Figure 13A schematic block diagram of a processor according to some embodiments of the present disclosure is shown. As Figure 13 shown, the processor 1000 may include an instruction parsing unit 1010 and an execution unit 1020. It can be understood that the processor 1000 may be implemented to execute the data loading method according to the embodiments of the present disclosure to load the tensor to be processed from the original tensor in the memory into the buffer area, or to implement the data storage method according to the embodiments of the present disclosure to obtain the second tensor based on the first tensor in the buffer area and write the second tensor into the memory.

[0213] Regarding the specific implementation processes of the data storage method and the data loading method, reference may be made to the above description and will not be repeated here. The processor provided by at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0214] According to some embodiments of the present disclosure, an electronic device is further provided. Figure 14 A schematic block diagram of an electronic device according to some embodiments of the present disclosure is shown. As Figure 14 shown, the electronic device 2000 may include a processor 2010 and a memory 2020 connected to the processor 2010. In addition, the processor 2010 may further include a buffer area. According to the embodiments of the present disclosure, the memory 2020 may be implemented in the form of a high-bandwidth memory HBM, which is not limited thereto. Specifically, according to the embodiments of the present disclosure, the processor 2010 is configured to run computer-executable instructions, and when the computer-executable instructions are run by the processor 2010, the data loading method according to the embodiments of the present disclosure is implemented to load the tensor to be processed from the original tensor in the memory into the buffer area, or the data storage method according to the embodiments of the present disclosure is implemented to obtain the second tensor based on the first tensor in the buffer area and write the second tensor into the memory.

[0215] The processor 2010 can perform various actions and processes according to a program stored in a non-transitory memory such as a non-volatile memory. Specifically, the processor 2010 can refer to a processor chip capable of performing parallel computing. For example, it can be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In addition, the processor 2010 can also be implemented as other conventional types of processors, which are not limited herein.

[0216] Regarding the specific implementation processes of the data storage method and the data loading method, reference can be made to the above description, which will not be repeated here. The processor provided in at least one embodiment of the present disclosure can achieve technical effects similar to those of the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0217] Figure 15 A block diagram showing an example computing device implementing some embodiments of the present disclosure is shown. As Figure 15 shown, the computing device 3000 is, for example, suitable for implementing the data loading method or the data storage method provided by the embodiments of the present disclosure. It should be noted that Figure 15 the components of the computing device 3000 shown are exemplary and not restrictive. According to actual application needs, the computing device 3000 may also have other components.

[0218] As Figure 15 shown, the computing device 3000 may include a processing device 3010 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.

[0219] For example, when the computer-readable instructions are run by the processing device 3010, one or more steps of the data loading method according to any of the above embodiments, or one or more steps of the data storage method according to any of the above embodiments may be executed. It should be noted that for the detailed description of the processing process of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for the detailed description of the processing process of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0220] For example, the processing device 3010, the read-only memory (ROM) 3020, and the random access memory (RAM) 3030 are connected to each other via the bus 3040. The input / output (I / O) interface 3050 is also connected to the bus 3040.

[0221] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 3030 and / or cache memory, etc. For example, the computer-readable instructions may be loaded from the storage device 3080 into the random access memory (RAM) 3030 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 3020, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage medium, as well as various data used and / or generated by the application programs, etc.

[0222] Generally, the following devices may be connected to the I / O interface 3050: an input device 3060 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 3070 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 3080 including, for example, magnetic tape, a hard disk, flash memory, etc.; and a communication device 3090. The communication device 3090 may allow the computing device 3000 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 15A computing device 3000 with various devices is shown. However, it should be understood that it is not required to implement or have all the shown devices, and the computing device 3000 may alternatively implement or have more or fewer devices. For example, the processing device 3010 may control other components in the computing device 3000 to perform desired functions. The processing device 3010 may be a Central Processing Unit (CPU), a Tensor Processing Unit (TPU), or a Graphics Processing Unit (GPU), etc., which has data processing capabilities and / or program execution capabilities. The GPU may be directly integrated into a System on Chip (SOC), directly integrated onto the motherboard, or built into the North Bridge chip of the motherboard.

[0223] According to some embodiments of the present disclosure, a non-transitory computer-readable storage medium is further provided. The non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they implement the data loading method according to the embodiments of the present disclosure, or implement the data storage method according to the embodiments of the present disclosure.

[0224] Figure 16 It is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 16 shown, the computer-readable storage medium 4000 may be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 4010 may be non-temporarily stored on the storage medium 4000. For example, when the computer-readable instructions 4010 are executed by a processor, one or more steps of the data loading method described in any of the above embodiments, or one or more steps of the data storage method described in any of the above embodiments, may be executed. It should be noted that for a detailed description of the processing procedure of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing procedure of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0225] As an example, the storage medium 4000 may be applied to the electronic device 2000 and / or the computing device 3000. For example, the storage medium 4000 may be implemented as the storage device 3080 in the computing device 3000.

[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0227] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0228] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on. The above description is only for the preferred embodiments of the present disclosure and the explanation of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0229] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0230] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

[0231] The following points also need to be explained regarding the present disclosure: (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.

[0232] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.

[0233] The above description is only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the appended claims.

Claims

1. A data loading method for loading a to-be-processed tensor from an original tensor in memory into a buffer, wherein: The data storage format of the original tensor in the memory is the same as the data storage format of the tensor to be processed in the cache area, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component. The shape size of the tensor to be processed is represented by a1, a2, a3, a4, and a5, and a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers. The 5 dimensions include batch dimension, depth dimension, height dimension, width dimension, and channel number dimension. The shape size of the original tensor is represented by b1, b2, b3, b4, and b5, and b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the 5 dimensions and are all positive integers. The data loading method comprises: Acquire the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Determine a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, wherein the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area, In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, and the request is divided in the second dimension, each data requested to be loaded belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and the data loaded in response to the request belongs to the data range of the original tensor, the data requested to be loaded comes from the original tensor and is stored continuously in the memory, wherein the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.

2. The method according to claim 1, wherein: The memory stores a plurality of tensors, the original tensor is one of the plurality of tensors, and the data storage format of the plurality of tensors in the memory is any one of NDHWC or N(C / x)DHW(xC), wherein for the data storage format of NDHWC, N corresponds to b1, indicating a batch dimension, D corresponds to b2, indicating a depth dimension, H corresponds to b3, indicating a height dimension, W corresponds to b4, indicating a width dimension, and C corresponds to b5, indicating a channel number dimension; for the data storage format of N(C / x)DHW(xC), N corresponds to b1, indicating a batch dimension, (C / x) corresponds to b2, indicating a channel number dimension, D corresponds to b3, indicating a depth dimension, H corresponds to b4, indicating a height dimension, and W(xC) corresponds to b5, indicating a width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension, wherein the data determined by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels.

3. The method according to claim 1, wherein: The tensor to be processed is used for convolution operation in a computing unit within the processor.

4. The method according to claim 1, wherein: In the case where the first dimension is a channel number dimension, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

5. The method according to claim 4, wherein: When the first dimension is the channel number dimension and the data storage format of the original tensor is NDHWC, in response to the size of the tensor to be processed in the channel number dimension being equal to the size of the original tensor in the channel number dimension, and the starting coordinates of the tensor to be processed in the channel number dimension being equal to the starting coordinates of the original tensor in the channel number dimension, the size relationship of the first dimension is determined so that when loading the tensor to be processed, the first dimension is loaded continuously.

6. The method according to claim 1, wherein: Determining multiple requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, includes: Determine a first request among the multiple requests and an initial state in a state machine in which the first request enters, based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, wherein the first request indicates the number of pixels to be obtained by the request; and Based on the initial state, in combination with the number of pixels to be obtained by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, the state machine is used to determine each request after the first request in the multiple requests.

7. The method according to claim 6, wherein: Each request includes a data read address for indicating a starting position for reading data from the memory, a data write address for indicating a starting position for writing data to the buffer area, and the number of pixels to be acquired by the request. Wherein, based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, determining the first request among the multiple requests and the initial state of the first request entering the state machine includes: Determining, based on the starting coordinates of the tensor to be processed, an initial state in which the first request enters the state machine; Using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; Determine a data reading address for the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; Determine a starting address in the buffer area for writing the tensor to be processed as a data writing address for the first request; and The number of pixels to be obtained by the first request is determined according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

8. The method according to claim 7, wherein: Determining the initial state of the first request entering the state machine based on the starting coordinates of the tensor to be processed includes: In response to a first coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension being less than a second coordinate value of the starting coordinate of the original tensor in the first dimension, determining that the initial state is the first state; In response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, determining that the initial state is a second state, wherein a difference between the second coordinate value and the third coordinate value is equal to a size of the original tensor in the first dimension; and In response to the first coordinate value being greater than or equal to the third coordinate value, the initial state is determined to be the third state.

9. The method according to claim 6, wherein: Based on the initial state, in combination with the number of pixels to be obtained by the first request, the number of pixels included in the tensor to be processed, and the data storage format of the original tensor, the state machine is used to determine each request after the first request in the multiple requests, including: Based on the initial state, using the state machine to determine the request initial coordinates corresponding to a second request after the first request and the number of pixels to be obtained by the second request; Determine a data reading address for the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; Determining a data writing address for the second request according to the number of pixels to be acquired by the second request; The state machine is updated according to information related to the second request, and subsequent requests among the requests are sequentially determined based on the updated state machine.

10. The method according to claim 7, wherein: In a case where the data storage format of the original tensor is NDHWC, for multiple requests for loading the tensor to be processed, determining a data read address of an nth request among the multiple requests includes: The data read address Addr1_n of the nth request is calculated according to the following formula: Addr1_n=u_addr_base+((((n_coord*tensor_d+d_coord)*tensor_h+h_coord)*tensor_w+w_coord) *tensor_c+c_coord), Among them, u_addr_base represents the storage address of the pixel at the starting coordinate position of the original tensor in the memory, n_coord, d_coord, h_coord, w_coord, c_coord represent the requested initial coordinates corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape size of the original tensor in the five dimensions.

11. The method according to claim 7, wherein: In a case where the data storage format of the original tensor is N(C / x)DHW(xC), for a plurality of requests for loading the tensor to be processed, determining a data read address of an nth request among the plurality of requests comprises: The data read address Addr2_n of the nth request is calculated according to the following formula: Addr2_n=u_addr_base+ n_coord*tensor_c*tensor_d*tensor_h*tensor_w+ c_coord *tensor_d*tensor_h*tensor_w+ d_coord *tensor_h*tensor_w*xC+ h_coord *tensor_w*xC+ w_coord *xC Among them, u_addr_base represents the storage address of the pixel at the starting coordinate position of the original tensor in the memory, n_coord, d_coord, h_coord, w_coord, c_coord represent the requested initial coordinates corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape size of the original tensor in the five dimensions.

12. The method according to claim 10 or 11, wherein: For a plurality of requests for loading the tensor to be processed, determining a data write address Addr3_n of an nth request among the plurality of requests includes: Addr3_n= b_addr_base+req_size_1+ req_size_2+…+ req_size_n-1 Among them, b_addr_base is the starting address of the tensor to be processed written into the buffer area, and req_size_1, req_size_2, ..., req_size_n-1 represent the number of pixels to be obtained by the first n-1 requests respectively.

13. The method according to claim 1, further comprising: Obtain boundary values ​​for the original tensor in at least a portion of the five dimensions, respectively, where the boundary values ​​are used to limit a boundary range for obtaining data from the original tensor. Among them, for the dimension with a boundary value set, each data requested to be loaded belongs to the original tensor and the valid data range specified by the boundary value or does not belong to the valid data range.

14. A data loading method, comprising: Receive a data loading instruction instructing execution of loading a to-be-processed tensor from an original tensor in a memory into a cache area, wherein the data storage format of the original tensor in the memory is the same as the data storage format of the to-be-processed tensor in the cache area, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensors in the storage component, and the shape size of the to-be-processed tensor is represented by a1, a2, a3, a4, and a5, and a1, a2, a3, a4, and a5 respectively indicate the sizes of the to-be-processed tensor in five dimensions and are all positive integers, and the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension, and the shape size of the original tensor is represented by b1, b2, b3, b4, and b5, and b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers; and After parsing the data loading instruction, using the execution unit to execute the data loading instruction, Wherein, using the execution unit to execute the data loading instruction includes: Acquire the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Determine a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, wherein the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area, In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, and the request is divided in the second dimension, each data requested to be loaded belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and the data loaded in response to the request belongs to the data range of the original tensor, the data requested to be loaded comes from the original tensor and is stored continuously in the memory, wherein the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.

15. A data storage method for acquiring a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, wherein: The data storage format of the first tensor in the cache area is the same as the data storage format of the second tensor in the memory, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensors in the storage component. The shape size of the second tensor is represented by a1, a2, a3, a4, and a5, and a1, a2, a3, a4, and a5 respectively indicate the sizes of the second tensor in five dimensions and are all positive integers. The five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension. The shape size of the first tensor is represented by b1, b2, b3, b4, and b5, and b1, b2, b3, b4, and b5 respectively indicate the sizes of the first tensor in the five dimensions and are all positive integers. The data storage method comprises: Acquire the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Determine, in combination with the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data of the first tensor from the first tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request into the memory in sequence to store the second tensor into the memory, In which, in response to the size relationship between the second tensor and the first tensor in the first dimension, when storing the second tensor, the first dimension cannot be continuously acquired but the dimension lower than the first dimension can be continuously acquired, and the request is divided in the second dimension, each requested data belongs to the data range of the first tensor or does not belong to the data range of the first tensor, and the data acquired in response to the request belongs to the data range of the first tensor, the requested data comes from the first tensor and is stored continuously in the cache area, wherein the data storage format of the first tensor indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.

16. A data storage method, comprising: Receive a data storage instruction instructing execution of acquiring a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, wherein the data storage format of the first tensor in the buffer is the same as the data storage format of the second tensor in the memory, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensors in the storage component, the shape size of the second tensor is represented by a1, a2, a3, a4, a5, a1, a2, a3, a4, a5 respectively indicate the sizes of the second tensor in five dimensions and are all positive integers, the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension and a channel number dimension, the shape size of the first tensor is represented by b1, b2, b3, b4, b5, b1, b2, b3, b4, b5 respectively indicate the sizes of the first tensor in the five dimensions and are all positive integers; and After parsing the data storage instruction, use the execution unit to execute the data storage instruction, The step of using the execution unit to execute the data storage instruction includes: Obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Determine, in combination with the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data of the first tensor from the first tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request into the memory in sequence to store the second tensor into the memory, In which, in response to the size relationship between the second tensor and the first tensor in the first dimension, when storing the second tensor, the first dimension cannot be continuously acquired but the dimension lower than the first dimension can be continuously acquired, and the request is divided in the second dimension, each requested data belongs to the data range of the first tensor or does not belong to the data range of the first tensor, and the data acquired in response to the request belongs to the data range of the first tensor, the requested data comes from the first tensor and is stored continuously in the cache area, wherein the data storage format of the first tensor indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.

17. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the data loading instruction instructs execution to load a to-be-processed tensor from an original tensor in a memory into a cache area, wherein the data storage format of the original tensor in the memory is the same as the data storage format of the to-be-processed tensor in the cache area, and the data storage format is used to indicate the storage order and dimensional arrangement of the tensors in the storage component, and the shape size of the to-be-processed tensor is represented by a1, a2, a3, a4, and a5, and a1, a2, a3, a4, and a5 respectively indicate the sizes of the to-be-processed tensor in five dimensions and are all positive integers, and the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension, and the shape size of the original tensor is represented by b1, b2, b3, b4, and b5, and b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in the five dimensions and are all positive integers; and The execution unit is configured to: execute the data loading instruction, The execution unit executes the data loading instruction, including: Acquire the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Determine a plurality of requests for loading the tensor to be processed in combination with the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, wherein the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area, In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, and the request is divided in the second dimension, each data requested to be loaded belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and the data loaded in response to the request belongs to the data range of the original tensor, the data requested to be loaded comes from the original tensor and is stored continuously in the memory, wherein the data storage format of the original tensor indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.

18. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is configured to: receive and parse a data storage instruction, wherein the data storage instruction instructs execution to obtain a second tensor based on a first tensor in a cache area and write the second tensor into a memory, wherein a data storage format of the first tensor in the cache area is the same as a data storage format of the second tensor in the memory, and the data storage format is used to indicate a storage order and dimensional arrangement of tensors in a storage component, and a shape size of the second tensor is represented by a1, a2, a3, a4, and a5, and a1, a2, a3, a4, and a5 respectively indicate the sizes of the second tensor in five dimensions and are all positive integers, and the five dimensions include a batch dimension, a depth dimension, a height dimension, a width dimension, and a channel number dimension, and a shape size of the first tensor is represented by b1, b2, b3, b4, and b5, and b1, b2, b3, b4, and b5 respectively indicate the sizes of the first tensor in the five dimensions and are all positive integers; and The execution unit is configured to: execute the data storage instruction, The execution unit executes the data storage instruction, including: Obtaining the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Determine, in combination with the number of pixels included in the second tensor, the data storage format of the first tensor, and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, a plurality of requests for storing the second tensor, wherein the plurality of requests are used to sequentially obtain data of the first tensor from the first tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request into the memory in sequence to store the second tensor into the memory, In which, in response to the size relationship between the second tensor and the first tensor in the first dimension, when storing the second tensor, the first dimension cannot be continuously acquired but the dimension lower than the first dimension can be continuously acquired, and the request is divided in the second dimension, each requested data belongs to the data range of the first tensor or does not belong to the data range of the first tensor, and the data acquired in response to the request belongs to the data range of the first tensor, the requested data comes from the first tensor and is stored continuously in the cache area, wherein the data storage format of the first tensor indicates that the first dimension takes precedence over the second dimension when storing or loading, and the first dimension and the second dimension are adjacent.

19. An electronic device comprising a processor and a memory connected to the processor, wherein: The processor includes a cache area, wherein The processor is configured to run computer-executable instructions, which, when executed by the processor, implement the data loading method according to any one of claims 1-14 to load the tensor to be processed from the original tensor in the memory to the cache area, or implement the data storage method according to claim 15 or 16 to obtain the second tensor based on the first tensor in the cache area and write the second tensor to the memory.

20. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the data loading method according to any one of claims 1 to 14 is implemented, or the data storage method according to claim 15 or 16 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN116822612A

  • Image processing method and device, electronic equipment and computer storage medium

    CN117765244A

  • Data processing method and device, electronic equipment and computer readable storage medium

    CN118152713A

  • Processor, chip product, computer equipment and tensor processing method

    CN119917166A

Cited By

  • Data loading method, processor, electronic equipment and storage medium

    CN120743196A

  • Data loading method and device, processor, electronic equipment and medium

    CN122364112A

  • Data loading method, processor, electronic equipment and medium

    CN122489460A