Data loading method, data storage method, processor, electronic equipment and medium

By using multiple request order to obtain and write data in a parallel processor, the pending tensor is loaded from the original tensor of memory to the cache area, solving the problem of low data loading efficiency in the prior art and improving computing performance.

CN120123264AActive Publication Date: 2025-06-10SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510607240.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to provide fast and efficient data loading/storage methods, which affects the computing performance of parallel processors.

Method used

Through a data loading method, the pending tensor is loaded from the original tensor in memory to the cache area, and the data of the original tensor is obtained and written to the cache area in sequence using multiple requests to adapt to the loading situation in different dimensions.

Benefits of technology

It improves the efficiency of data loading, improves memory access bandwidth, enhances the hardware utilization of computing devices, and improves computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123264A_ABST
    Figure CN120123264A_ABST
Patent Text Reader

Abstract

The invention provides a data loading method, a data storage method, a processor, electronic equipment and a medium. The data loading method is used for loading a to-be-processed tensor from an original tensor of a memory to a cache region, and comprises the following steps: determining a plurality of requests for loading the to-be-processed tensor in combination with the number of pixels included in the to-be-processed tensor and an initial coordinate of the to-be-processed tensor in a coordinate system determined by the original tensor, the plurality of requests are used for sequentially acquiring data of the original tensor from the original tensor by taking the starting coordinate as a starting point until the number of the acquired data is equal to the number of pixels included in the tensor to be processed, and the data loaded by each request belong to the data range of the original tensor or do not belong to the data range of the original tensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data loading method, a data storage method, a processor, an electronic device, and a medium. Background Art

[0002] A tensor is a data structure of a multi-dimensional array. Tensor operations are widely used in processors such as parallel processors. For example, in the field of deep learning, the dimensions of the input data, the intermediate data processed during the deep learning process, and the output data are elastic and not exact. Therefore, an elastic data form is needed to describe various types of data, and thus the concept of tensors is generated. In the field of deep learning, all data to be operated on can be stored and exist in the form of tensors. If the data is not in the form of tensors, it needs to be first converted into the data structure form of tensors. As an example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor itself can have any number of dimensions. For example, it can be represented as a 5-dimensional array.

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processing devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of parallel processors need to process a large amount of data, and this data is generally stored in the storage components of the parallel processors. For example, the storage components can be high-speed memories. Through data loading instructions, this data can be loaded from the storage components to the buffer for calculation, and through data storage instructions, the data in the buffer can be stored in the memory.

[0004] How to provide a fast and efficient data loading / storing method is crucial for the computing performance of the device. Summary of the Invention

[0005] Embodiments of the present disclosure provide a data loading method, a data storage method, a processor, an electronic device, and a medium, which are used to provide a fast and efficient data loading / storing method, improve the memory access bandwidth, and increase the hardware utilization rate of computing devices.

[0006] According to a first aspect of the present disclosure, there is provided a data loading method for loading a tensor to be processed from an original tensor in memory into a buffer. The data storage format of the original tensor in memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the number of channels dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is a pixel, and is accumulated to higher dimensions level by level. The number of pixels is not calculated for the number of channels dimension. The data loading method includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests and writing the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer. When the size relationship between the tensor to be processed and the original tensor in the first dimension causes the tensor to be processed to not be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, the requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And when the data loaded by the request all belongs to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as or adjacent to the second dimension.

[0007] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the H dimension or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. When the first dimension is the C dimension or the first dimension is the N dimension and the size of the tensor to be processed in the C dimension is greater than x, the second dimension is the C dimension. When the first dimension is the N dimension and the size of the tensor to be processed in the C dimension is equal to x, the second dimension is the overall dimension higher than the N dimension.

[0008] According to some embodiments of the present disclosure, when the size of the tensor to be processed in the first dimension is not equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension is not equal to the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0009] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension, and it is determined to adopt a partitioning for requests in the H dimension.

[0010] According to some embodiments of the present disclosure, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be continuously loaded in the H dimension, and it is determined to adopt a partitioning for requests in the D dimension.

[0011] According to some embodiments of the present disclosure, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a partitioning for requests in the C dimension.

[0012] According to some embodiments of the present disclosure, the tensor to be processed is used for a convolution operation in a computing unit within a processor.

[0013] According to some embodiments of the present disclosure, each request includes a data read address for indicating the starting position to read data from the memory, a data write address for indicating the starting position to write data to the buffer, and the number of pixels to be acquired by this request. The request initial coordinate corresponding to the next request is determined based on the request initial coordinate corresponding to the previous request. The request initial coordinate corresponding to each request is used to determine the data read address, data write address, and the number of pixels to be acquired by the request. Among them, when sequentially determining the request initial coordinate corresponding to the request, the coordinate value of the request initial coordinate is incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0014] According to some embodiments of the present disclosure, multiple requests for loading a tensor to be processed are determined in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, including: determining a first request among the multiple requests and an initial state when the first request enters a state machine based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, where the first request indicates the number of pixels to be acquired by the request; and determining each request after the first request among the multiple requests using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed.

[0015] According to some embodiments of the present disclosure, each request includes a data read address for indicating a starting position for reading data from memory, a data write address for indicating a starting position for writing data to a buffer, and the number of pixels to be acquired by the request. Among them, determining a first request among the multiple requests and an initial state when the first request enters a state machine based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0016] According to some embodiments of the present disclosure, determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0017] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired for the first request and the number of pixels included in the tensor to be processed, a state machine is used to determine each request among multiple requests after the first request, including: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired for the second request; determining the data read address of the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; determining the data write address of the second request according to the number of pixels to be acquired for the second request; updating the state machine according to the information related to the second request, and sequentially determining subsequent requests among each request based on the updated state machine.

[0018] According to some embodiments of the present disclosure, for multiple requests for loading a tensor to be processed, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Wherein, u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the original tensor in five dimensions respectively.

[0019] According to some embodiments of the present disclosure, for multiple requests for loading a tensor to be processed, determining the data write address of the nth request among the multiple requests includes: for m requests with the same coordinate value range of the request data in the C dimension, calculating the data write address Addr_2_i of the ith request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i-1 * copy_c Among them, b_addr represents the data write address of the first request sent among m requests, req_size_1, req_size_2, …, req_size_i-1 represent the number of pixels to be obtained by each of the first i-1 requests respectively, and copy_c is the size of the tensor to be processed in the C dimension; among them, in response to the data of each of the m requests including the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer area, and in response to the data of each of the m requests including the t-th data element to the (t + x)-th data element of the original tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1) * x t, m, and i are positive integers.

[0020] According to some embodiments of the present disclosure, a plurality of requests are sequentially sent, and the data corresponding to each request is sequentially written into the buffer area to load the tensor to be processed into the buffer area, including: for any request, in response to all the data loaded by any request being located in the memory, sending any request to the memory; in response to all the data loaded by any request not being located in the memory, converting any request into writing a plurality of predetermined values into the buffer area, where the number of the plurality of predetermined values is determined by the number of pixels to be obtained by any request.

[0021] According to some embodiments of the present disclosure, the data loading method further includes: obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, where the boundary values are used to define the boundary range for obtaining data from the original tensor, and for the dimension for which the boundary value is set, the data loaded by each request belongs to the valid data range defined by the original tensor and the boundary value or does not belong to the valid data range.

[0022] According to a second aspect of the present disclosure, a data loading method is provided, including: receiving a data loading instruction indicating to execute loading a tensor to be processed from an original tensor in memory into a buffer, wherein the data storage format of the original tensor in memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the number of channels dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension, and wherein, in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, and the number of pixels is not calculated for the number of channels dimension; and after parsing the data loading instruction, executing the data loading instruction using an execution unit, wherein executing the data loading instruction using the execution unit includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data from the original tensor with the starting coordinates as the starting point until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the buffer to load the tensor to be processed into the buffer, wherein in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, partitioning requests are made in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor, and in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, wherein the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

[0023] According to a third aspect of the present disclosure, a data storage method is provided for obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, wherein the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the number of channels dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the number of channels dimension does not calculate the number of pixels. The data storage method includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests, and sequentially writing the data corresponding to each request into the memory to store the second tensor into the memory. Wherein, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

[0024] According to a fourth aspect of the present disclosure, a data storage method is provided, including: receiving a data storage instruction indicating to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory, where the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the channel number of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not calculate the number of pixels; and after parsing the data storage instruction, using an execution unit to execute the data storage instruction, where using the execution unit to execute the data storage instruction includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests and writing the data corresponding to each request into memory in sequence to store the second tensor into memory, where in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, request division is performed in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor, and in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0025] According to a fifth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction. In this case, the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is regarded as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not count the number of pixels; and the execution unit is configured to: execute the data loading instruction. When the execution unit executes the data loading instruction, it includes: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests, and writing the data corresponding to each request into the buffer in sequence to load the tensor to be processed into the buffer. In this case, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, division of requests is performed in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in the memory, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0026] According to a sixth aspect of the present disclosure, a processor is provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory. The data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the number of channels dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the number of channels dimension does not count the number of pixels; and the execution unit is configured to: execute the data storage instruction, where the execution unit executing the data storage instruction includes: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into memory to store the second tensor into memory. Wherein, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, partitioning of requests is performed in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0027] According to a seventh aspect of the present disclosure, an electronic device is provided, including a processor and a memory connected to the processor. The processor includes a buffer, and the processor is configured to run computer-executable instructions, which when run by the processor, implement the data loading method according to the embodiments of the present disclosure to load a tensor to be processed from an original tensor in the memory into the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain a second tensor based on a first tensor in the buffer and write the second tensor into the memory.

[0028] According to an eighth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, a data loading method according to an embodiment of the present disclosure is implemented, or a data storage method according to an embodiment of the present disclosure is implemented. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 Shows a schematic structural diagram of a General-Purpose computing on Graphics Processing Unit (GPGPU); Figure 2 Shows a schematic structure of a tensor; Figure 3A Shows a schematic diagram of the data storage format of NDHWC; Figure 3B Shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32; Figure 4 Shows a schematic diagram of obtaining a tensor in related technologies; Figure 5 Shows a schematic flowchart of a data loading method provided by at least one embodiment of the present disclosure; Figure 6 Shows a schematic diagram of continuously obtaining a tensor to be processed according to an embodiment of the present disclosure; Figure 7A Shows a schematic diagram of continuously obtaining a tensor to be processed for the data storage format of NDHWC according to an embodiment of the present disclosure; Figure 7B Shows a schematic diagram of continuously obtaining a tensor to be processed for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure; Figure 8A Shows a schematic diagram of setting boundary values for an original tensor according to an embodiment of the present disclosure; Figure 8B Shows a schematic diagram of continuously obtaining data in the case of setting boundaries for an original tensor according to an embodiment of the present disclosure; Figure 8CShows a schematic diagram of continuously obtaining a tensor to be processed according to a set step size according to an embodiment of the present disclosure; Figure 9A Shows a schematic diagram of continuously loading an original tensor of N(C / x)DHW(xC) as a tensor to be processed stored in NDHWC according to an embodiment of the present disclosure; Figure 9B Shows a state schematic diagram of a state machine provided according to some embodiments of the present disclosure; Figure 9C Shows, according to an embodiment of the present disclosure, for Figure 9A An example of a state transition schematic diagram according to the PerH partitioning method; Figure 10 Shows a schematic flowchart of a data loading method provided by at least one embodiment of the present disclosure; Figure 11 Shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure; Figure 12 Shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure; Figure 13 Shows a schematic block diagram of a processor according to some embodiments of the present disclosure; Figure 14 Shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure; Figure 15 Shows a block diagram of an example computing device implementing some embodiments of the present disclosure; and Figure 16 Shows a schematic block diagram of a computer-readable storage medium according to some embodiments of the present disclosure. Detailed implementation manners

[0031] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0032] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Words such as "upper", "lower", "left", "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of this disclosure clear and concise, some detailed descriptions of known functions and known components are omitted in this disclosure.

[0033] Figure 1 A schematic structural diagram of a GPGPU is shown. As Figure 1 shown, a GPGPU is actually an array of programmable multi-processors. For example, the programmable multi-processors can be Streaming Processor Clusters (SPCs), for example including Figure 1 the streaming processor cluster 1 shown, ..., the streaming processor cluster M, where M is a positive integer. In a general-purpose graphics processor, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing between multiple streaming processor clusters is performed through a global cache or High Bandwidth Memory (HBM).

[0034] As Figure 1 shown, taking the streaming processor cluster 1 as an example, 1 streaming processor cluster can include multiple Compute Units (CUs), for example Figure 1 the compute unit 1, compute unit 2, ..., compute unit K in it, where K is a positive integer. Each compute unit is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A compute unit can include multiple cores (also called computing cores or compute cores), and each compute core includes an Arithmetic Logic Unit (ALU), a floating-point computing unit, etc., and the compute core is used to perform specific computing tasks. In addition, the compute unit also includes registers (such as Figure 1The register bank) and shared memory are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit. In addition, the buffer can be understood as being used to share data among the computing units within the streaming processor cluster.

[0035] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processing unit (or parallel computing processor), and then multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundle (or simply called thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0036] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates the thread bundles so that multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-shared. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the intermediate-level cache or global cache or high-bandwidth memory for read / write operations, etc.

[0037] As Figure 1 shown, general computing operations, such as matrix computing operations in the field of artificial intelligence, usually require a large amount of data. These data are usually stored in a memory, such as in a high-bandwidth memory HBM. When performing general computing operations, data needs to be loaded from the memory (Load operation), and when obtaining the computing result, data needs to be stored to the memory (Store operation). The storage method of data in the memory will affect the memory access bandwidth, and thus affect the hardware utilization rate of the computing unit.

[0038] For example, general computing operations include General Matrix Multiplication (abbreviated as GEMM). As an example, the data required for general matrix multiplication is represented as two 5D arrays, such as the matrix multiplication calculation of two tensors A and tensor B. In addition, general computing operations also include convolution operations, which are manifested as data dot products. It can be understood that in the field of artificial intelligence, other computing operations are also involved, which will not be listed one by one here. The data involved in these calculations is usually embodied in the form of tensors.

[0039] For example, for a certain tensor A in the buffer, its shape and size can be represented by a1, a2, a3, a4, and a5. a1, a2, a3, a4, and a5 respectively indicate the sizes of the tensor data in 5 dimensions, and a1, a2, a3, a4, and a5 are positive integers. For example, the 5 dimensions include [N, D, H, W, C]. The N dimension represents the batch size, that is, N represents the batch dimension, which is the number of data samples grabbed in one training. The D dimension represents the depth dimension, the H dimension represents the height dimension of the input data, the W dimension represents the width dimension of the input data, and the C dimension represents the number of channels dimension. For example, taking tensor A as an example, a1 can be the size of the N dimension, a2 can be the size of the D dimension, a3 can be the size of the H dimension, a4 can be the size of the W dimension, and a5 can be the size of the C dimension. Of course, the present disclosure does not make specific limitations on this.

[0040] As an example, Figure 2 shows a schematic structure of a tensor. In Figure 2 the shown tensor, a1 is the size of the N dimension and is equal to 1, a2 is the size of the D dimension and is equal to 1, a3 is the size of the H dimension and is equal to 5, a4 is the size of the W dimension and is equal to 4, and a5 is the size of the C dimension and is equal to 64. For example, Figure 2 the pixel elements of the tensor in

[0041] are represented as 0, 1, 2, 3,... and so on. Figure 2 The placement of tensors in the memory (such as memory or buffer) can have various formats, called data storage formats (layout). The data storage format is used to indicate the storage order and dimension arrangement of tensors in the storage component. The following uses the

[0042] shown tensor to describe different data storage formats. Figure 3A shows a schematic diagram of the NDHWC data storage format.

[0043] For example, for the NDHWC linear mode, as Figure 3A shown, starting from the first element (element 0 in Figure 3A ) of the first channel (a5 = 0, c0 in Figure 3A ), then storing the first element (element 20 in Figure 3A ) of the second channel (a5 = 1, c1 in Figure 3A ), and so on, until all the first elements of all channels are laid out, for example, until the first element (element in Figure 3A ) of the 64th channel (a5 = 63, c63 in Figure 3AAfter the element 1260) in, select the first channel (a5 = 0, Figure 3A the second element of c0) in Figure 3A element 1) in, and then store the second element of the second channel (a5 = 1, Figure 3A c1) in Figure 3A element 21) in, and so on until the second elements of all channels are laid out, and so on.

[0044] In the related art, the data storage format may also include N(C / x)DHW(xC), also known as the Interleave mode, where x can be set to 8, 16, 32, etc. as needed.

[0045] The N(C / x)DHW(xC) data storage format is similar to the NDHWC data storage format, but a key difference is that in the layout of N(C / x)DHW(xC), a5 channels are divided into a5 / x groups, with each group having x channels: the first group consists of channels a5 = 0 to a5 = x - 1, the second group consists of channels a5 = x to a5 = 2x - 1, and each group is arranged in the NDHWC format.

[0046] Figure 3B Shows a schematic diagram of the data storage format of N(C / x)DHW(xC), where x = 32.

[0047] As Figure 3B shown, 64 channels are divided into two groups, with each group having 32 channels. The first group consists of channels a5 = 0 ( Figure 3B c0) in to a5 = 31 ( Figure 3B c31) in, and the second group consists of channels a5 = 32 to a5 = 63. Then each group is arranged in the NDHWC format.

[0048] In memory, for example, a certain tensor B, similar to a certain tensor A in the buffer, this tensor B can be stored in memory according to one of the two data storage formats described above, and the shape dimensions of the tensor can be similarly expressed as b1×b2×b3×b4×b5, where b1, b2, b3, b4, b5 respectively indicate the dimensions of the tensor B in these 5 dimensions and are all positive integers.

[0049] It can be understood that in the related art and possible future developments, the data storage format of tensors is not limited to the above-described N(C / x)DHW(xC) data storage format and NDHWC data storage format. Further, for multiple tensors that can be stored in memory and the buffer during the calculation process, usually the storage space of memory is much larger than that of the buffer, but it is farther from the calculation unit, and the data transfer efficiency is lower than that of the buffer.

[0050] In the related art, during the computing process of a processing device, a large amount of computing data will be generated, for example, in the form of tensors, which can be temporarily stored in a buffer. For example, the buffer here can refer to Figure 1 the buffer in the streaming processor cluster shown in Figure 1 Furthermore, these data can also be transferred from the buffer or directly stored in the memory. For example, the memory can be

[0051] It can be understood that in this article, the data storage process of storing the tensors in the buffer into the memory and the data loading process of loading the tensors in the memory into the buffer can be implemented in a similar manner. Therefore, for the sake of convenience of description, in some embodiments or examples, only the data loading process is described as an example, and those skilled in the art can apply it similarly to the data storage process. For the differences between the two, additional descriptions will be made.

[0052] As an example, Figure 4 shows a schematic diagram of obtaining tensors in the related art. As Figure 4 shown, there is a tensor A stored in the memory, which can be placed in the memory in the above-mentioned N(C / x)DHW(xC) or NDHWC data storage format. That is, tensor A is a 5D array. In Figure 4 only three dimensions W, H, and C are schematically shown. Schematically, on the left side of Figure 4 a dimension coordinate system of tensor A is shown, where they are the W dimension, the H dimension, and the C dimension respectively. During the data loading process, all or part of the data in tensor A can be loaded into the buffer through a loading instruction. For example, Figure 4 tensor B in Figure 4In the 3D schematic diagram shown, the tensor B is a cuboid determined by the above-mentioned second starting point and the respective dimensional sizes. Based on the information about the second starting point and the respective dimensional sizes of the tensor to be loaded, the memory can load the tensor B into the buffer area.

[0053] In the above related technologies, the block-based data loading method can be applied to processes such as the general matrix multiplication GEMM in general computing operations. However, the above block-based data loading method is not applicable to convolution operations, whose computational feature is data dot multiplication. The above block-based data loading method will limit the computational efficiency of this type of per-pixel convolution operation, increase data access time, and reduce the overall performance of the processor, which limits the further development space of efficient and general-purpose processors.

[0054] Furthermore, in the memory, the original tensor is continuously stored in the memory in the data storage format as described above. The shape of the original tensor is represented as b1×b2×b3×b4×b5, where b1, b2, b3, b4, and b5 respectively indicate the sizes of the original tensor in these 5 dimensions and are all positive integers. When extracting or storing a partial tensor in the original tensor, since it is a partial tensor inside the original tensor, it may not be possible to continuously extract in each dimension during extraction, and it is impossible to accurately obtain the amount of data and the data position required when loading or storing data to optimally achieve data loading or storing, which greatly reduces the data bandwidth during data loading and reduces the hardware computational efficiency.

[0055] In view of the above technical problems in the related technologies, the present disclosure provides a data loading method, a data storage method, a processor, an electronic device, and a non-transitory computer-readable storage medium. In the present disclosure, first, the tensor is no longer obtained and transported in a block form, but is sequentially obtained and transported in units of pixels to be applicable to computational operations such as convolution operations. Further, for the operation of loading the tensor to be processed from the original tensor in the memory into the buffer area, multiple requests for loading the tensor to be processed are determined according to the number of pixels included in the tensor to be loaded, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, so as to sequentially obtain data from the original tensor starting from the starting coordinates using the determined multiple requests until the number of obtained data is equal to the number of pixels included in the tensor to be processed, thereby improving the data transportation efficiency for the data loading operation, reducing the memory access time, and improving the overall performance of the processor.

[0056] In the data loading method provided by at least one embodiment of the present disclosure, requests are divided according to whether the tensors can be continuously loaded in the first dimension. The data to be loaded by each request either belongs to the data range of the original tensor or does not belong to the data range of the original tensor. The division of data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, thereby improving the hardware utilization rate of the computing unit and enhancing the hardware performance.

[0057] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0058] The data loading method provided by at least one embodiment of the present disclosure is used to load a tensor to be processed from an original tensor in memory to a buffer area. The memory according to the embodiments of the present disclosure may be, for example, a high-bandwidth memory (HBM), and the buffer area may be, for example, a buffer (buffer) in a streaming processor cluster, which is not limited herein.

[0059] In the method according to the embodiments of the present disclosure, the data storage format of the original tensor in memory is different from the data storage format of the tensor to be processed in the buffer area. The data storage format is used to indicate the storage order and dimensional arrangement of the tensor in the storage component. For example, the storage component refers to the above-mentioned HBM or buffer area. That is to say, in the data loading method according to the embodiments of the present disclosure, the placement manner of the original tensor in memory is different from the placement manner of the tensor to be processed to be loaded in the buffer area next. Specifically, the original tensor from which data is to be obtained is placed in memory in an N(C / x)DHW(xC) interleaved manner, and after being loaded into the buffer area, it is changed to be placed in an NDHWC linear manner. The data loading / storage method proposed by the present disclosure is applicable to the case of changing the data placement manner during data transfer.

[0060] The original tensor in memory is a 5D tensor, and its shape size is represented by 5 parameters b1, b2, b3, b4, b5. b1, b2, b3, b4, b5 respectively indicate the sizes of the original tensor in 5 dimensions and are all positive integers. The 5 dimensions include the batch dimension, depth dimension, height dimension, width dimension, and number of channels dimension. Similarly, the shape size of the tensor to be processed is represented by a1, a2, a3, a4, a5. a1, a2, a3, a4, a5 respectively indicate the sizes of the tensor to be processed in 5 dimensions and are all positive integers.

[0061] For the NDHWC data storage format, N corresponds to b1, representing the batch dimension, D corresponds to b2, representing the depth dimension, H corresponds to b3, representing the height dimension, W corresponds to b4, representing the width dimension, and C corresponds to b5, representing the number of channels dimension.

[0062] For the data storage format of N(C / x)DHW(xC), N corresponds to b1, representing the batch dimension, (C / x) corresponds to b2, representing the number of channels dimension, D corresponds to b3, representing the depth dimension, H corresponds to b4, representing the height dimension, and W(xC) corresponds to b5, representing the width dimension. In the data storage format of N(C / x)DHW(xC), x is a positive integer, and the number of channels of xC is bound to the width dimension. Generally, x can be set to an integer multiple of 4.

[0063] Regarding the characteristics of the above two data storage formats, reference can be made to the description above in combination with Figure 3A - Figure 3B and will not be repeated here.

[0064] Furthermore, for the 5D data of the original tensor, the data determined by the height dimension and the width dimension represents a pixel (or, it can also be called an element), and accumulates step by step towards higher dimensions. Among them, the number of pixels is not calculated for the number of channels dimension. As an example, as Figure 4 shown in, the values of each W dimension and H dimension can determine a pixel. For example, W = 0 and H = 0 correspond to the first pixel (or element) in tensor A, and W = 1 and H = 0 correspond to the second pixel in tensor A. In the tensor, the C dimension does not affect the number of pixels.

[0065] Figure 5 is a schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure. As Figure 5 shown, the data loading method provided by at least one embodiment of the present disclosure at least includes steps S101 - S103.

[0066] In step S101, obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. Then, in step S102, combine the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine multiple requests for loading the tensor to be processed, where the multiple requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed. In step S103, sequentially send the multiple requests and sequentially write the data corresponding to each request into the buffer area to load the tensor to be processed into the buffer area. According to the embodiment of the present disclosure, the tensor to be processed can be used for convolution operations in the computing units within the processor. Herein, the number of data can also be expressed as the number of pixels.

[0067] In the data loading method according to the embodiment of the present disclosure, in order to obtain the tensor to be processed from the original tensor in the memory, it is necessary to indicate the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor. For example, the starting coordinates can be combined withFigure 4 The described second starting point (C = 0, W = 3, H = 0). Further, in the data loading method according to an embodiment of the present disclosure, it is also necessary to indicate the number of pixels included in the tensor to be processed (for example, denoted as copy_pixel_num), that is, the total number of pixels in the tensor to be loaded into the buffer. This continuous data copying method is different from the block-based data loading method adopted in the related art above (where it is necessary to indicate the sizes of the tensor to be processed in each dimension). According to the data loading method of an embodiment of the present disclosure, the range of the tensor to be processed is determined by the number of pixels to be obtained. Further, instead of obtaining the data in a block from the original tensor, the data is sequentially obtained from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed.

[0068] Next, first, the implementation process of realizing the above continuous data loading based on copy_pixel_num and the starting coordinates will be described, and then, the implementation process of determining multiple requests for loading the tensor to be processed will be described later.

[0069] As an example, Figure 6 shows a schematic diagram of obtaining the tensor to be processed in a continuous manner according to an embodiment of the present disclosure. In Figure 6 , tensor A represents the original tensor in memory. In Figure 6 , each pixel point is shown as a square. Tensor A covers all the squares, that is, both the white squares and the gray-shaded squares. The starting point of this original tensor is represented as (C = 0, W = 0, H = 0), for example, used to indicate the position of this original tensor in memory. Then, in Figure 6 , tensor B represents the tensor to be obtained, and its starting point is represented as (C = 0, W = 3, H = 0), that is, start obtaining data from the 3rd pixel point in the first row of tensor A. Then, in Figure 6 , the process of obtaining the tensor to be processed in a sequential (or also called continuous) manner according to an embodiment of the present disclosure is schematically shown, that is, starting from the starting point (C = 0, W = 3, H = 0) as shown by the dashed arrow, obtaining data from tensor A pixel by pixel, for example, until the number of obtained pixels reaches copy_pixel_num. In the example of Figure 6 , only the three dimensions of W, H, and C are shown. It can be understood that if the data in these 3 dimensions still does not reach copy_pixel_num, the data of higher dimensions can be further obtained, which is not limited here. In addition, as shown in Figure 6 , the C dimension itself does not affect the number of pixels. Assume that Figure 6 the data storage format of tensor A in memory is NDHWC. In Figure 6In the example shown, the total number of pixels of the tensor to be processed is copy_pixel_num = 27.

[0070] In an embodiment according to the present disclosure, obtaining data from the original tensor sequentially with the starting coordinate as the starting point for the data of the original tensor includes: excluding the channel number dimension from the five dimensions of the original tensor, and in the order of dimensions from low to high, using the starting coordinate as the starting point for data acquisition, and sequentially obtaining data until the number of pixels obtained reaches copy_pixel_num. The reason for excluding the channel number dimension from the five dimensions is that for a tensor, the channel number dimension does not count the number of pixels. That is to say, S102 may include obtaining pixel data from the original tensor with the starting coordinate as the starting point for data acquisition in the order of the width dimension, height dimension, depth dimension, and batch dimension until the number of pixels obtained reaches copy_pixel_num.

[0071] Compare Figure 4 and Figure 6 For the two ways of obtaining the tensors to be processed shown, the data loading method provided by the embodiments of the present disclosure can achieve pixel-by-pixel data loading, rather than Figure 4 the block-by-block acquisition in [reference], this data loading method is more beneficial to the calculation process such as convolution operation and is conducive to improving the operation efficiency. Thus, based on the method provided by the embodiments of the present disclosure, for the data to be subjected to convolution operation next, the processor can, for example, instruct in the form of an instruction to fetch this part of the data from the memory to the buffer according to the above continuous data loading method for use in the convolution operation. It can be understood that the above processor can reasonably use, according to the type of operation to be performed or the data processing characteristics, whether to use the continuous data loading method or the block-by-block data loading method, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer may also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0072] According to some embodiments of the present disclosure, there are multiple tensors stored in the memory, and the data loading method may further include: obtaining indication information of the storage location of the original tensor in the memory. As an example, the indication information may include the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. It can be understood that the storage space of the memory is generally much larger than the buffer area, and various tensor data generated during the calculation process can be stored therein. These tensor data can be arranged in the memory in any one of the above-mentioned N(C / x)DHW(xC) or NDHWC data storage formats. In the embodiments according to the present disclosure, in order to enable the memory to know the specific location of the tensor to be obtained in the memory, indication information of the storage location of the original tensor in the memory may also be obtained. The indication information includes the starting coordinates of the original tensor in the memory and the size values in 5 dimensions, respectively represented as tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5. As an example, in the case where the original tensor is in the N(C / x)DHW(xC) data storage format, tensor_b1, tensor_b2, tensor_b3, tensor_b4, tensor_b5 respectively represent the numerical values of the original tensor in the 5 dimensions of N, C / x, D, H, and W.

[0073] Regarding the influence of the data storage format on the loading process of the tensor to be processed ("tensor B"), it will be described below in conjunction with Figure 7A - Figure 7B for description. In conjunction with Figure 7A and Figure 7B In the example shown, the data storage method of the tensor to be processed in the buffer area is the same as the data arrangement method of the original tensor in the memory.

[0074] Figure 7A FIG. shows a schematic diagram of continuously obtaining the tensor to be processed for the NDHWC data storage format according to an embodiment of the present disclosure. For Figure 7A the tensor shown, its data storage format in the memory is in the form of NDHWC, or is arranged in the NDHWC manner.

[0075] In Figure 7A In the example shown, the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 8 (the C dimension is not shown in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Taking the N dimension as an example, N_dim = 2 corresponds toFigure 7A For N0 and N1 in [description], based on the above parameters, the total number of pixel points included in the original tensor is 64.

[0076] Next, in Figure 7A the example of [description], the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixel points included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7A the pixel points where W = 2 and H = 2 under the N = 0 and D = 0 dimensions shown in [description]. Starting from this pixel point, in the data arrangement manner, in the order from low dimension to high dimension (i.e., in the order of W, H, D, N, as shown by the data acquisition schematic arrow in Figure 7A [description]), data is sequentially acquired until the number of acquired data is equal to the number of pixel points included in the tensor to be processed. In Figure 7A the example shown in [description], starting from the determined starting point, data is acquired pixel by pixel in the arrow order to load this part of the data in the original tensor in memory into the buffer as the tensor to be processed for subsequent operations such as convolution calculations.

[0077] Figure 7B [description] shows a schematic diagram of continuously acquiring the tensor to be processed for the data storage format of N(C / x)DHW(xC) according to an embodiment of the present disclosure. For Figure 7B the tensor shown in [description], its data storage format in memory is in the form of N(C / x)DHW(xC), or is arranged in the manner of N(C / x)DHW(xC). In Figure 7B the example of [description], x = 8, that is, every 8 C channels are grouped and bound to the W dimension. Compared with Figure 7A the NDHWC data storage format shown in [description], for the data storage format of N(C / x)DHW(xC), the rank of the C dimension in the tensor is higher and is located between the N dimension and the D dimension.

[0078] In Figure 7B the example shown in [description], the original tensor corresponds to the data shown in the squares in the figure. Among them, the size of the original tensor in the C dimension is 16 (shown as two 8*C dimensions in the figure), the size in the W dimension is 4 (W_dim = 4), the size in the H dimension is 4 (H_dim = 4), the size in the D dimension is 2 (D_dim = 2), and the size in the N dimension is 2 (N_dim = 2). Thus, the total number of pixel points included in the original tensor can be obtained as 64.

[0079] Next, in the example of Figure 7B , the starting point coordinates of the tensor to be processed in the original tensor are represented as w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0, and the number of pixels included in the tensor to be processed is copy_pixel_num = 40. Based on the above information, it can be obtained that the starting position of the tensor to be processed in the original tensor is Figure 7B the pixel point where W = 2 and H = 2 under N = 0 and D = 0 and the first 8*C dimension shown in Figure 7B . Starting from this pixel point, according to the data arrangement method, in the order from low dimension to high dimension (i.e., in the order of W, H, D, C, N, as shown by the data acquisition schematic arrow in Figure 7A ), data is sequentially acquired until the number of acquired data is equal to the number of pixels included in the tensor to be processed. Compared with the acquisition order shown in Figure 7B , since the rank of the C dimension is higher, in the example of Figure 7A and Figure 7B , first, data of W, H, D, and the first 8*C dimension is acquired according to the starting point, then, data of the second 8*C dimension is acquired in the order of W, H, D, and finally, data of the N dimension is acquired. It can be understood that in the examples of Figure 7A and Figure 7B , the number of pixels of the tensor to be processed acquired is 40, and the difference is only caused by the different arrangement methods of the C dimension. In addition, it should be noted that for the second 8*C dimension shown in Figure 7B , its starting point at the W and H dimensions should be aligned with the starting point of the first 8*C dimension.

[0080] In the example shown in Figure 7B , starting from the determined starting coordinates, data is acquired pixel by pixel in the arrow order to load this part of the data in the original tensor in the memory into the buffer area as the tensor to be processed for subsequent operations such as convolution calculations.

[0081] In some embodiments according to the present disclosure, boundary values can also be defined for at least a part of the five dimensions of the original tensor respectively. As an example, left and right boundary values can be set for, for example, the C dimension, the W dimension, the H dimension, and the D dimension, which are represented as (L-tensor_C, R-tensor_C), (L-tensor_W, R-tensor_W), (L-tensor_H, R-tensor_H), and (L-tensor_D, R-tensor_D) respectively. The above boundary values are used to define the boundary range for obtaining data from the original tensor, where the boundary values are arbitrary values compared to the size values of the original tensor in this dimension.

[0082] In the method described in combination with Figure 7A and Figure 7B no limitation is imposed on the range of the original tensor (i.e., the boundary values), that is, the tensor to be processed is obtained from the complete data of the original tensor. In the method according to the embodiments of the present disclosure, it is also proposed that the range for obtaining the tensor to be processed from the original tensor can be delimited by setting boundary values. Further, in the implementation process, the boundary values can be set to arbitrary values compared to the size values of the original tensor in this dimension. That is to say, the boundary values can exceed the range of the original tensor itself.

[0083] As an example, Figure 8A shows a schematic diagram of setting boundary values for the original tensor according to the embodiments of the present disclosure. As Figure 8A shown, the rectangular box represents the range covered by the original tensor, which can be any dimension in the original tensor, such as the C dimension, the W dimension, the H dimension, or the D dimension. Generally, no boundary values are set for the N dimension. In the six sub-pictures of Figure 8A the relationship between the left boundary value and the right boundary value (shown as "L" and "R" in Figure 8A ) and the size value of the original tensor in this dimension is shown respectively.

[0084] According to the embodiments of the present disclosure, when setting boundary values, sequentially obtaining data from the original tensor includes: for the data part of the original tensor covered by the range defined by the boundary values, obtaining data sequentially from the range of the original tensor defined by the boundary values, for example, referring to the order described in Figure 7A and Figure 7B . In comparison, for the data part of the original tensor not covered by the range defined by the boundary values, it is represented as invalid data. For the invalid data, the buffer directly fills the tensor to be processed with a predetermined value, where the predetermined value is equal to 0.

[0085] As an example, in Figure 8AIn the first sub - picture, both the left boundary value L and the right boundary value R are on the left side of the original tensor. That is, the data to be obtained in this dimension are all invalid data. In this case, for the invalid data, for example, this part of the data can be automatically filled by sending a zero - filling instruction to the buffer area. Another example is, in Figure 8A In the second sub - picture, the left boundary value L is on the left side of the left boundary of the original tensor data range, and the right boundary value R is on the left side of the right boundary of the original tensor. That is, for the tensor to be processed to be obtained, part of it is invalid data and part of it is valid data in the original tensor. Schematically, in Figure 8A , the data corresponding to the slanted shaded part is represented as valid data, and the rest are all invalid data. In this case, for the valid data, for example, refer to Figure 7A and Figure 7B for the described order, while for the invalid data, for example, this part of the data can be automatically filled by sending a zero - filling instruction to the buffer area.

[0086] In the method according to the embodiments of the present disclosure, by setting boundary values for the original tensor, the range from which the tensor to be processed is to be taken can be further delimited. In practical applications, this implementation method can adapt to the characteristics of operations such as convolution operations. For example, it is beneficial to reduce the amount of calculation and greatly improve the flexibility of data. As an example, assume that the original tensor corresponds to an intermediate tensor for feature extraction of an entire input picture, and the input picture includes a specific target, such as an object to be recognized, and the object does not cover the entire picture. That is, the picture includes a background part. In this case, by setting boundary values, the range of the tensor to be processed to be obtained can be limited to the part of the original tensor corresponding to the specific target, so as to reduce the amount of calculation of subsequent operations such as convolution operations and improve the processing efficiency.

[0087] As an example, Figure 8B shows a schematic diagram of continuous data acquisition in the case where boundaries are set for the original tensor. Among them, compared with the situation shown in Figure 6 , Figure 8B can be understood as setting boundary values (bound) for the C - dimension, W - dimension, and H - dimension of the tensor A in Figure 6 , and the boundaries are all within the size ranges of the tensor A in the C - dimension, W - dimension, and H - dimension, which is equivalent to the boundary situation shown in the fourth sub - picture in Figure 8A . Specifically, Figure 8B the outer square box in Figure 6 corresponds to the tensor A (corresponding to the tensor A in Figure 6 ). After setting the boundaries therein, according to the data loading method of the embodiments of the present disclosure, data will be sequentially acquired starting from the starting - point coordinates within the set boundary range until the number of acquired data is equal to the number of pixels included in the tensor to be processed. InFigure 8B Among them, the data part composed of blocks corresponds to the data range framed by the boundary values set for the C dimension, W dimension, and H dimension, and data is sequentially acquired from the data range framed by the boundary. Data located outside the data range framed by the boundary can be regarded as invalid data. Regarding Figure 8B The process of sequentially acquiring data in Figure 6 can be referred to in combination with the description of

[0088] In some embodiments according to the present disclosure, a step value for acquiring data can also be set. As an example, the set data step value is used to specify the stride for sequentially acquiring data from the original tensor. Among them, for the data of the original tensor, starting from the starting coordinate, sequentially acquiring data from the original tensor includes: for the data of the original tensor, starting from the starting coordinate, sequentially acquiring data from the original tensor according to the data step value.

[0089] Figure 8C FIG. shows a schematic diagram of acquiring a tensor to be processed in a continuous manner according to the set step in an embodiment of the present disclosure. Among them, the step is equal to 2, that is, one data is taken every other pixel point, that is, only the pixels shown in the shaded part are sequentially acquired. The step being equal to 2 means skipping one pixel. Similarly, when the step is equal to 3, it means skipping two pixels, and so on. It can be understood that when the set step is equal to 1, it corresponds to Figure 7A - Figure 7B the data loading method of acquiring each pixel point shown. In practical applications, by setting the step, the computational amount of subsequent operations such as convolution operations on the data can be further reduced, and the processing efficiency can be improved. In addition, the flexibility of data loading can be further enhanced.

[0090] In the data loading method provided according to an embodiment of the present disclosure, a new data acquisition mode different from the block-based data acquisition method is provided, that is, pixel-by-pixel data loading can be achieved according to the starting point and the number of pixels to be acquired (as shown in Figure 6 ), rather than Figure 4 the block-based acquisition in

[0091] This data loading method is more conducive to the calculation process of operations such as convolution operations and is beneficial to improving the operation efficiency. Specifically, the tensor is no longer acquired and transported in a block form, but is sequentially acquired and transported in units of pixel points to be applicable to calculation operations such as convolution operations, improving the data transportation efficiency for such operations, reducing the memory access time, and improving the overall performance of the processor. Figure 6 、 Figure 7A - Figure 7B 、 Figure 8A - Figure 8C, which details the implementation process in the data loading method according to an embodiment of the present disclosure. In this process, data of the original tensor is sequentially fetched from the original tensor starting from the starting coordinates until the number of fetched data is equal to the number of pixels (copy_pixel_num) included in the tensor to be processed.

[0092] Next, the implementation process regarding Figure 5 in step S102 shown in determining multiple requests for loading the tensor to be processed by combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor will be described.

[0093] In an embodiment according to the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but can be performed in dimensions lower than the first dimension, requests are divided in the second dimension. Data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as or adjacent to the second dimension. As an example, when continuous loading cannot be performed in the first dimension but can be performed in each dimension lower than the first dimension, requests are divided in the second dimension.

[0094] For example, a coordinate system is determined with a certain element in the original tensor as the origin of the coordinate system. For example, referring to Figure 6 the embodiment of, the top - left vertex in the original tensor is used as the origin of the coordinate system. Of course, the present disclosure is not limited to this.

[0095] The starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor are, for example, Figure 6 the coordinates of the top - left vertex of the tensor to be processed in the embodiment of. The coordinate values of the starting coordinates of the tensor to be processed are the minimum coordinate values among the coordinate values of all elements in the tensor to be processed.

[0096] For example, the first dimension is the W dimension, H dimension, or D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. Specifically, if the first dimension is the W dimension, the second dimension is the H dimension; if the first dimension is the H dimension, the second dimension is the D dimension; if the first dimension is the D dimension, the second dimension is the C dimension.

[0097] In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension not being equal to x, the second dimension is the C dimension.

[0098] In response to the first dimension being an N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is an overall dimension higher than the N dimension. This can be understood as a special case of the requested partitioning. The overall dimension higher than the N dimension refers to the entire tensor. For example, it is expressed as requesting partitioning in the Per1 manner, that is, the entire tensor to be processed can be loaded into the buffer through one request. In this case, there is only one request for loading the tensor to be processed.

[0099] In the present disclosure, the original tensor is stored in the memory in the data storage format of N(C / x)DHW(xC). The data storage format of the original tensor in the memory is different from the data storage format (NDHWC) of the tensor to be processed in the buffer. In the memory, the C dimension of the original tensor is the lowest dimension, followed by the W dimension, then the H dimension, then the D dimension, and the highest dimension is the N dimension. When the tensor to be processed is stored in the buffer, the W dimension is the lowest dimension, followed by the H dimension, then the D dimension, then the C dimension, and the highest dimension is the N dimension. And for the W dimension, continuous xC is substantially considered.

[0100] For example, in response to the size copy_t of the tensor to be processed in the first dimension being not equal to the size tensor_t of the original tensor in the first dimension, and / or the first coordinate value t_coord_b of the starting coordinate of the tensor to be processed in the first dimension being not equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0101] According to some embodiments of the present disclosure, when the first dimension is the W dimension and the second dimension is the H dimension, in response to the size of the tensor to be processed in the W dimension being not equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension being not equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension, and it is determined to adopt partitioning of the request in the H dimension. For ease of description, this way of partitioning the request can be referred to as the request partitioning rule according to the H dimension (PerH).

[0102] Taking the second coordinate value as 0 as an example, when the first dimension is the W dimension and the second dimension is the H dimension, it is determined that continuous loading cannot be performed in the W dimension when any of the following conditions is satisfied: (1) The starting coordinate of the tensor to be processed in the W dimension is not equal to 0 (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension At this time, it can be understood that the data loaded by each request belongs to the same row (the same H dimension), and the data loaded by different requests is located in different rows, that is, the requests are divided in the H dimension, and the requests are split by row (H dimension).

[0103] According to some embodiments of the present disclosure, when the first dimension is the H dimension, the second dimension is the D dimension, in response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be continuously loaded in the H dimension, and it is determined to use the request division in the D dimension. For ease of description, this request division method can be referred to as the request division rule according to the D dimension (PerD).

[0104] Assume that the first dimension is the H dimension and the second dimension is the D dimension. When any of the following conditions is satisfied, it is determined that continuous loading cannot be performed in the H dimension: (1) The starting coordinate of the tensor to be processed in the H dimension is not equal to 0 (2) The size copy_h of the tensor to be processed in the H dimension is greater than the size tensor_h of the original tensor in the H dimension (3) The size copy_h of the tensor to be processed in the H dimension is less than the size tensor_h of the original tensor in the H dimension At this time, it can be understood that the coordinates of the data to be loaded by each request in the W dimension and the H dimension are different, but the coordinates in the D dimension and the N dimension are the same. Therefore, the requests are divided in the D dimension.

[0105] According to some embodiments of the present disclosure, when the first dimension is the D dimension, the second dimension is the C dimension, in response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension, and when the size of the tensor to be processed in the C dimension is greater than x, it is determined to use the request division in the C dimension. For ease of description, this request division method can be referred to as the request division rule according to the C dimension (PerC).

[0106] Assume that the first dimension is the D dimension and the second dimension is the C dimension. When any of the following conditions is satisfied, it is determined that continuous loading cannot be performed in the D dimension: (1) The starting coordinate of the tensor to be processed in the D dimension is not equal to 0 (2) The size copy_d of the tensor to be processed in the D dimension is greater than the size tensor_d of the original tensor in the D dimension (3) The size copy_d of the tensor to be processed in dimension D is less than the size tensor_d of the original tensor in dimension D And when the size of the tensor to be processed in dimension C is greater than x (denoted as copy_c greater than xC), then the requested partitioning is performed in dimension C. Specifically, considering the particularity of the N(C / x)DHW(xC) interleaving pattern, the partitioning is performed in dimension C with every x data elements as a group. This is because for the N(C / x)DHW(xC) data storage format, every x data elements are continuously stored in the channel number dimension in the storage component.

[0107] Assume the first dimension is dimension C, and the second dimension is also dimension C at this time. Determine that continuous loading cannot be performed in dimension C when any of the following conditions holds: (1) The starting coordinate of the tensor to be processed in dimension C is not equal to 0 (2) The size copy_c of the tensor to be processed in dimension C is greater than the size tensor_c of the original tensor in dimension C (3) The size copy_c of the tensor to be processed in dimension C is less than the size tensor_c of the original tensor in dimension C At this time, the requested partitioning is still performed in dimension C. Specifically, considering the particularity of the N(C / x)DHW(xC) interleaving pattern, the partitioning is performed in dimension C with every x data elements as a group. This is because for the N(C / x)DHW(xC) format, every x data elements are continuously stored in the channel number dimension in the storage component. For ease of description, this way of requested partitioning can be called the request partitioning rule according to dimension C (PerC).

[0108] Assume the first dimension is dimension N. Determine that continuous loading cannot be performed in dimension N when any of the following conditions holds: (1) The starting coordinate of the tensor to be processed in dimension N is not equal to 0 (2) The size copy_n of the tensor to be processed in dimension N is greater than the size tensor_n of the original tensor in dimension N (3) The size copy_n of the tensor to be processed in dimension N is less than the size tensor_n of the original tensor in dimension N If the size of the tensor to be processed in dimension C is greater than x, and the second dimension is dimension C at this time, the requested partitioning is still performed in dimension C. Specifically, considering the particularity of the N(C / x)DHW(xC) interleaving pattern, the partitioning is performed in dimension C with every x data elements as a group. This is because for the N(C / x)DHW(xC) format, every x data elements are continuously stored in the channel number dimension in the storage component.

[0109] If the size of the tensor to be processed in the C dimension is equal to x, and at this time the second dimension is the overall dimension higher than the N dimension, for the tensor data located in the storage component, in essence, a request can be sent to the memory to load the data in the memory at this time, and other requests can be used to fill the data not located in the memory, such as the above-mentioned invalid data. For ease of description, this way of dividing requests can be called request division according to the whole of the tensor to be processed, which can be expressed as Per1.

[0110] In addition, since the data in the N(C / x)DHW(xC) format needs to be written into the buffer in the NDHWC format, which is different from the data storage format of the original tensor itself, and the dimension arrangement orders of the two data storage formats are also different, the update logic of the initial coordinates of the requests during request splitting is different from the dimension arrangement order of the data storage format itself.

[0111] Each request includes a data read address for indicating the starting position of reading data from the memory, a data write address for indicating the starting position of writing data into the buffer, and the number of pixels to be obtained by the request. The initial coordinates of the next request are determined according to the initial coordinates of the previous request, and the initial coordinates of each request are used to determine the data read address, data write address, and the number of pixels to be obtained by the request.

[0112] As described later, except for the first dimension, the coordinate values of the other dimensions higher than the first dimension are updated according to the coordinate values lower than that other dimension. Since the data of the original tensor needs to be written into the buffer in the NDHWC storage format, for ease of hardware implementation, the coordinate values of the initial coordinates of the requests are incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and the coordinate value of the C dimension is incremented by x when updated. That is to say, the N dimension is updated prior to the C dimension during the update, but since in fact the N dimension of the tensor to be processed is the highest dimension originally, and the N dimension is given priority when writing into the buffer, it is necessary to reserve in advance the data positions that should be written into the buffer but have not been written yet when writing into the buffer, and write the data into the reserved data positions when writing this data subsequently, thereby ensuring that the tensor to be processed in the buffer can be arranged in the NDHWC dimension order.

[0113] For example, taking the W dimension as the first dimension, the order of coordinate update dimensions is illustrated. The coordinate value h_coord of the request initial coordinate corresponding to the next request in the H dimension is incremented by 1, and when it reaches the boundary of the H dimension of the tensor to be processed, the initial coordinate value h_coord_b is returned; the coordinate value d_coord of the request initial coordinate corresponding to the next request in the D dimension is incremented by 1 when h_coord returns h_coord_b. If it does not reach the boundary of the D dimension of the tensor to be processed, it remains unchanged, and when it reaches the boundary of the D dimension of the tensor to be processed, the initial coordinate value d_coord_b is returned; the coordinate value n_coord of the request initial coordinate corresponding to the next request in the N dimension is incremented by 1 when d_coord returns d_coord_b. If it does not reach the boundary of the N dimension of the tensor to be processed, it remains unchanged, and when it reaches the boundary of the N dimension of the tensor to be processed, the initial coordinate value n_coord_b is returned; the coordinate value c_coord of the request initial coordinate corresponding to the next request in the C dimension is incremented by x when n_coord returns the initial coordinate value n_coord_b, and remains unchanged in other cases.

[0114] The calculation method for the data writing address also varies considering the need to reserve data positions in advance. For specific details, please refer to the relevant descriptions in the following text.

[0115] The object to be loaded in this disclosure, that is, the tensor to be processed, is split into multiple different requests according to a certain rule and sent sequentially, and the tensor data is written into the buffer area in order, thereby efficiently loading the tensor to be processed into the buffer area. One principle during splitting is that the data loaded by each request is either all in memory or all not in memory. This is because when the requested data is in memory, a request is sent to memory, and when the requested data is not in memory, the request can be converted into an operation such as writing a predetermined value to the buffer area by hardware. This splitting method can send corresponding loading requests to different hardware more reasonably.

[0116] In addition, during splitting, if the size relationship between the tensor to be processed and the original tensor in the first dimension causes the tensor to be processed to not be continuously loadable in the first dimension but continuously loadable in dimensions lower than the first dimension, then the requests are further divided with the first dimension as the boundary, so that the requests can be split more reasonably, the continuously stored data in memory can be retained as much as possible, the number of requests can be reduced, the tensor to be processed can be loaded efficiently, the bandwidth when loading or storing data from memory can be greatly increased, the performance can be improved, and the efficiency of the hardware computing unit can be improved.

[0117] In an implementation solution where boundary values are set for the original tensor, the data loading method according to an embodiment of the present disclosure may further include: obtaining boundary values for at least a part of the five dimensions of the original tensor respectively, and the boundary values are used to define the boundary range for obtaining data from the original tensor. In some implementation manners, the boundary value may be any value compared with the size value of the original tensor in this dimension. As an example, the boundary value set for the W dimension may be any situation as shown in Figure 8A . Further, for the dimension with the boundary value set, either all the data requested to be loaded belongs to the data boundary of the original tensor itself and within the valid data range specified by the boundary value, or all does not belong to the valid data range specified by both the data boundary of the tensor itself and the set boundary value.

[0118] As shown in Figure 8A , for the first sub - figure, both the left and right boundary values are located on the left side of the data range of the original tensor, that is, both are outside the data range of the original tensor, indicating that all the data to be obtained for this dimension is invalid (for example, represented as out of bound, oob). For the Figure 8A second sub - figure, the left boundary is located on the left side of the data range of the original tensor, and the right boundary is located within the data range of the original tensor. Thus, only the part of the data that is both within the data range of the original tensor and to the left of the right boundary R is valid data (in Figure 8A , the data corresponding to the diagonal shaded part is represented as valid data), and the rest of the data to be obtained is invalid data. According to an embodiment of the present disclosure, for the dimension with the boundary value set, the rule for dividing requests is that either all the data loaded by each request is valid data, or all is invalid data (i.e., oob), that is, valid data and invalid data cannot be loaded by the same request.

[0119] The following specifically describes the determination method for multiple requests for loading the tensor to be processed.

[0120] According to an embodiment of the present disclosure, for the multiple requests determined according to step S102, in the actual implementation process of the processor, it can be implemented by setting a state machine. For the state machine, multiple parameters required to determine the above - mentioned requests can be set, and according to the specific values of the parameters and the initial parameters of the original tensor and the tensor to be processed, the specific tensor to be processed in the original tensor is loaded into the buffer area. The following will describe the implementation solution related to determining multiple requests for loading the tensor to be processed in combination with the state machine.

[0121] According to some embodiments of the present disclosure, in step S102, multiple requests for loading the tensor to be processed are determined in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, including: determining the first request among the multiple requests and the initial state of the first request entering the state machine based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, where the first request indicates the number of pixels to be acquired by this request; and determining each request after the first request among the multiple requests by using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed.

[0122] As an example, for the task of loading the tensor to be processed in the original tensor in memory into the buffer, the first request can be determined first based on the initial parameters. For example, based on the number of pixels copy_pixel_num included in the tensor to be processed, the data storage format of the original tensor (N(C / x)DHW(xC)), and the starting coordinates of the tensor to be processed (the starting coordinates indicate the position of the first pixel of the tensor to be processed in the coordinate system of the original tensor, for example, represented as its positions in 5 dimensions, c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b), the first request among the multiple requests and the initial state of the first request entering the state machine are determined, where the first request indicates the number of pixels to be acquired by this request (req_size_1).

[0123] Next, in combination with Figure 9A Describe how to determine the partitioning method of the requests and the first request. Specifically, Figure 9A FIG. shows a schematic diagram of continuously loading the original tensor of N(C / x)DHW(xC) as the tensor to be processed stored in NDHWC according to an embodiment of the present disclosure. That is, in Figure 9A In the example of, the original tensor is placed in memory in the N(C / x)DHW(xC) data storage format, and the tensor to be processed is continuously loaded from it into the buffer and the tensor to be processed is placed in the buffer in the NDHWC data storage format, so that during the data loading process, the change in the data placement method is realized. It is precisely due to this change that Figure 9A The data acquisition order shown in is different from Figure 7B The order shown in, in Figure 7B In the example of, it is a schematic diagram of continuously loading the original tensor of N(C / x)DHW(xC) as the tensor to be processed stored in N(C / x)DHW(xC), that is, in Figure 7B In the example of, there is no change in the data storage format.

[0124] In Figure 9A the example, copy_pixel_num = 40, and the starting coordinates are w_coord_b = 2, h_coord_b = 2, d_coord_b = 0, n_coord_b = 0, c_coord_b = 0 respectively. As Figure 9A shown, for the above starting coordinates, the starting pixel corresponding to the starting coordinates in the original tensor, that is, starting from this pixel to obtain data. Then, for the data storage format of N(C / x)DHW(xC), the W dimension is the lowest dimension. Thus, taking the W dimension as the first dimension, it is determined whether continuous loading can be performed in the W dimension. In response to the size relationship between the tensor to be processed and the original tensor in the W dimension such that continuous loading cannot be performed in the W dimension when loading the tensor to be processed, the above-mentioned PerH request partitioning method is adopted, that is, partitioning the requests one by one for each H. That is to say, each row of pixels is loaded into the buffer by one request. According to the judgment on whether continuous loading can be performed in the W dimension described above: In response to the size of the tensor to be processed in the W dimension (copy_w) not being equal to the size of the original tensor in the W dimension (tensor_w), and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension (w_coord_b) not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension (for example, 0), it is determined that continuous loading cannot be performed in the W dimension for the tensor to be processed. Summarized as follows: It is determined that the data to be obtained cannot be continuously loaded in the W dimension when any of the following conditions is satisfied: (1) The starting coordinate w_coord_b of the tensor to be processed in the W dimension is not equal to 0; (2) The size copy_w of the tensor to be processed in the W dimension is greater than the size tensor_w of the original tensor in the W dimension; (3) The size copy_w of the tensor to be processed in the W dimension is less than the size tensor_w of the original tensor in the W dimension; After determining the partitioning method, for example, in the PerH manner, the first request can be correspondingly determined. Refer to Figure 9A , the first request is used to load the row of pixels where the starting pixel corresponding to the starting coordinates is located. Its initial state is equal to the starting coordinates of the tensor to be processed, and the number of pixels to be obtained by this first request is req_size_1 = 2 because the subsequent requests are partitioned in the row-by-row (PerH) manner. The specific acquisition order is as shown by the arrows in Figure 9A until the number of pixels to be obtained is equal to copy_pixel_num = 40 of the tensor to be processed.

[0125] Specifically, in Figure 9A the example, compared with Figure 7BAn example, the difference lies in the loading priority for the C dimension. Specifically, in Figure 9A the example is that, in order to change the data placement method during the data loading process, during the actual loading process, the C dimension is taken as the highest priority. Refer to Figure 9A the loading order indicated by the arrow in. First, load the H dimension of the first 8*C dimension, then the D dimension, and then the N dimension. After the N dimension (N1) corresponding to the first 8*C dimension is loaded, jump back to the second 8*C dimension and continue to load data according to the H dimension, D dimension, and N dimension until the number of pixels obtained is equal to the number of pixels included in the tensor to be processed, that is, copy_pixel_num = 40.

[0126] According to an embodiment of the present disclosure, after determining the request, it is further necessary to determine the data reading address of the starting position of reading data from the memory for the request, which is used to indicate the data writing address of the starting position of writing data to the buffer area. Among them, based on the number of pixels included in the tensor to be processed, the data storage format of the original tensor, and the starting coordinates of the tensor to be processed, determining the first request among multiple requests and the initial state when the first request enters the state machine includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the tensor to be processed; taking the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; determining the data reading address of the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; determining the starting address of writing the tensor to be processed in the buffer area as the data writing address of the first request; and determining the number of pixels to be obtained by the first request according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

[0127] According to some embodiments of the present disclosure, the data storage format of the original tensor is N(C / x)DHW(xC). For multiple requests for loading the tensor to be processed, determining the data reading address of the nth request among multiple requests includes: Calculating the data reading address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Among them, u_addr_base represents the storage address in memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in five dimensions respectively.

[0128] According to some embodiments of the present disclosure, for multiple requests for loading tensors to be processed, determining the data write address of the nth request among the multiple requests includes: For m requests with the same coordinate value range of the data in the C dimension of the request, for example, the data with the same coordinate value range is split into m requests for sending, for example, the coordinate values of the m requests are all from 0 to x - 1, or from x to 2*x - 1, etc., calculate the data write address Addr_2_i of the ith request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i - 1 * copy_c Among them, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2, …, req_size_i - 1 represent the number of pixels to be obtained by the previous i - 1 requests respectively, and copy_c is the size of the tensor to be processed in the C dimension.

[0129] In response to that the data of each request among the m requests includes the first x data elements of the original tensor in the C dimension, for example, the coordinate values are from 0 to x - 1, b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer.

[0130] In response to that the data of each request among the m requests includes the tth data element to the (t + x)th data element of the original tensor in the C dimension, where t is greater than x, for example, t is equal to 2*x or 3*x, etc., b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1) * x Among them, t, m, and i are positive integers.

[0131] For the first sent request, the corresponding initial request coordinates are the starting coordinates of the tensor to be processed. Thus, referring to the above formula, the data read address of the first request can be calculated. The starting address of the tensor to be processed written into the buffer is used as the data write address of the first request.

[0132] The number of pixels to be acquired by the first request can be determined based on parameters such as the determined request partitioning method, the starting coordinates of the tensor to be processed, and the size of the original tensor. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is less than the second coordinate value (i.e., the coordinate value of the starting point of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the shape size copy_t of the tensor to be processed in the first dimension. For example, if the sum of the first coordinate value t_coord_b of the tensor to be processed in the first dimension and the shape size copy_t of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value (i.e., the coordinate value of the starting point of the original tensor in the first dimension, for example, 0), the data length loaded by the first request is the absolute value of the first coordinate value t_coord_b, and this data length represents the number of pixels to be acquired by this request.

[0133] Considering that the state machine has the advantages of a clear logical structure, being easy to maintain and expand, and being particularly suitable for processing logical scenarios with multiple conditions and multiple branches, in this disclosure, the state machine is adopted to automatically update the initial request coordinates and the loaded data length corresponding to each request, avoiding complex conditional nesting, with clear logic, being easy to maintain, and having strong scalability.

[0134] According to some embodiments of the present disclosure, based on the starting coordinates of the tensor to be processed, determining the initial state of the first request entering the state machine includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0135] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired for the first request and the number of pixels included in the tensor to be processed, a state machine is used to determine each request after the first request among multiple requests, including: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired for the second request; determining the data read address of the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; determining the data write address of the second request according to the number of pixels to be acquired for the second request; updating the state machine according to the information related to the second request, and sequentially determining subsequent requests in each request based on the updated state machine.

[0136] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the data to be loaded for the current request is located in the memory. At this time, the data read address and data write address of the current request can be determined with reference to the above calculation formula, and then the current request is sent to the memory to load the corresponding data into the buffer.

[0137] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the data to be loaded for the current request is not located in the memory. At this time, in response to all the data to be loaded for the current request not being located in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values into the buffer, where the number of multiple predetermined values is determined by the length of the data to be loaded specified by the current request. For example, the predetermined value is 0.

[0138] Figure 9B Shows a state diagram of the state machine provided according to some embodiments of the present disclosure.

[0139] For example, the states of the state machine include a first state s0, a second state s1, and a third state s2. The initial state of entering the state machine is determined by the starting coordinates (c_coord_b, w_coord_b, h_coord_b, d_coord_b, n_coord_b) of the tensor to be processed.

[0140] For example, in response to the first coordinate value t_coord_b of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, it is determined that the initial state is the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that the initial state is the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; in response to the first coordinate value being greater than or equal to the third coordinate value, it is determined that the initial state is the third state.

[0141] Reference Figure 9B , the tensor_t direction represents the data of the first dimension. The data with the coordinate of the t dimension less than the second coordinate value and greater than the third coordinate value is not in memory or not within the range of the original tensor data ( Figure 9B The part shown by the diagonal shading, and this part of the data can be called invalid data), and the data with the size of the t dimension between the second coordinate value and the third coordinate value ( Figure 9B The white part in) is stored in memory, which is the actual size of the original tensor data in the first dimension.

[0142] The state only switches to itself and adjacent states. For example Figure 9B In, the first state s0 can jump to the first state s0 or the second state s1, the second state s1 can jump to the first state s0, the second state s1 and the third state s2, and the third state s2 can jump to the first state s0, the second state s1 and the third state s2.

[0143] Figure 9B The 6 cases (① to ⑥) in show the state transition changes that the tensor to be processed experiences for different sizes in the first dimension.

[0144] For example, Figure 9B In case ①, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the second coordinate value. The state machine always loops and jumps in the first state s0 on the left.

[0145] For example, Figure 9B In case ②, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the second coordinate value and less than the third coordinate value. The state machine loops and jumps between the first state s0 and the second state s1.

[0146] For example, Figure 9B In case ③, the first coordinate value is greater than or equal to the second coordinate value but less than the third coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is less than the third coordinate value. The state machine always loops and jumps in the second state s1.

[0147] For example, Figure 9B In case ④, the first coordinate value is less than the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps among the first state s0, the second state s1 and the third state s2.

[0148] For example, Figure 9BIn case ⑤, the first coordinate value is greater than or equal to the second coordinate value, and the sum of the first coordinate value and the size of the tensor to be processed in the first dimension is greater than or equal to the third coordinate value. The state machine loops and jumps between the second state s1 and the third state s2.

[0149] For example, Figure 9B In case ⑥, the first coordinate value is greater than or equal to the third coordinate value. The state machine always loops and jumps in the third state s2.

[0150] After determining the initial state, based on the initial state, determine the request initial coordinates corresponding to the second sent request, and determine the state entered by the second sent request in combination with the trigger condition. Then, based on the state entered by the second sent request, determine the request initial coordinates corresponding to the third sent request and the data length loaded by the second sent request, and determine the state entered by the third sent request in combination with the trigger condition, and so on.

[0151] For example, the state machine outputs the request initial coordinates corresponding to the current request and the data length loaded by the current request in each state. In addition, the state machine also prepares the request initial coordinates corresponding to the next request for the next state. After determining the request initial coordinates corresponding to the current request, it can be judged whether the data loaded by the current request is in the memory according to the request initial coordinates.

[0152] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being greater than or equal to the second coordinate value and less than the third coordinate value, it is determined that all the data to be loaded by the current request is in the memory. It can be understood that for the data range that all belongs to the original tensor, it will all be within the memory range (this part of the data is actually stored in the memory), and for the data range that all does not belong to the original tensor, it will all not be within the memory range and is not the actually stored data. That is, the data exceeding the original tensor range can be understood as invalid data. For example, this part of the invalid data can be directly filled with zeros in the buffer. At this time, the data read address and data write address of the current request can be determined with reference to the above formula, and then the current request is sent to the memory to load the corresponding data into the buffer.

[0153] For example, in response to the coordinate value of the request initial coordinates corresponding to the current request in the first dimension being less than the second coordinate value or greater than or equal to the third coordinate value, it is determined that all the data to be loaded by the current request is not in the memory. At this time, in response to all the data to be loaded by the current request not being in the memory, the current request is converted to, for example, using hardware to write multiple predetermined values into the buffer, where the number of the multiple predetermined values is determined by the data length loaded by the current request specified. For example, the predetermined value is 0.

[0154] The process of using the state machine to determine the request initial coordinates and the data length loaded by each request will be specifically described below.

[0155] Figure 9C shows a state transition schematic diagram according to the PerH partitioning method for an example of Figure 9A For the PerH partitioning method, the condition is that continuous loading cannot be performed in the W dimension. At this time, the adjacent two rows of pixels to be acquired are discontinuous in memory, and a split request for each row of pixels is required to obtain data.

[0156] In Figure 9C on the left side of, only three states s0, s1, and s2 are shown for the W and H dimensions. Specifically, s0 can correspond to the first state, s1 can correspond to the second state, and s2 can correspond to the third state. In addition, in Figure 9C the state transition schematic diagram on the right side of, there is also an idle state, which is used to represent the idle state before data loading or after all the tensors to be processed have been loaded. Table 1 below shows the meanings corresponding to each state: Table 1

[0157] That is to say, in the example of Figure 9C the middle blank box can correspond to the size of the original tensor in the W dimension (tensor_w). Its left boundary can correspond to, for example, Figure 9B the second coordinate value in, for example, equal to 0, and its right boundary can correspond to, for example, Figure 9C the third coordinate value in, for example, equal to tensor_w. According to Table 1, when w_coord < 0, it corresponds to state s0; when tensor_w > w_coord >= 0, it corresponds to state s1; when w_coord >= tensor_w & tensor_w < copy_w, it corresponds to state s2. It can be understood that the description of the meanings of the corresponding states in Table 1 does not set boundary values, and it can also be understood that the left boundary value for the W dimension is equal to 0, and the right boundary value is equal to tenser_w, that is, the boundary is the same as the size of the W dimension of the original tensor.

[0158] Figure 9C The various states shown in can be jumped between. Table 2 below shows the jumps between the various states and the corresponding conditions for making the corresponding jumps: Table 2

[0159] It should be noted that in the examples of Table 2 and subsequent tables above, the case of setting boundary values for the original tensor is further considered. Among them, right_bound represents the right boundary value set for the W dimension, down_bound represents the boundary value in the H dimension, and hind_bound represents the boundary value in the D dimension. The setting of boundary values and strides can refer to the description in combination with Figure 8A - Figure 8C above. In addition, for the case where no boundary value is set, it can be understood that each boundary is equal to the size of the original tensor in that dimension. In addition, it can be understood that in other embodiments according to the present disclosure, boundary values and strides can also be set for the C dimension, for example, which is not limited herein. Regarding the left boundary left_bound and right boundary right_bound for the W dimension, the upper boundary up_bound and lower boundary down_bound for the H dimension, and the front boundary front_bound and rear boundary hind_bound for the D dimension, reference can be made to Figure 6 the cube shown in for easy understanding.

[0160] In each table in this article, the parameter remain_copy_p represents the number of remaining pixels not yet acquired. For example, before the first request, remain_copy_p is equal to the number of pixels copy_pixel_num corresponding to the tensor to be processed. After the first request, remain_copy_p is equal to copy_pixel_num - req_size_1, where req_size_1 is the number of pixels to be acquired in the first request. The parameters Is_last_c and Is_not_last_c are respectively used to indicate whether the currently acquired data is for the last 8*C dimension. For example, referring to Figure 9A the schematic diagram, for the entire C dimension (C_Dim), for the data acquisition process of the first 8*C dimension, it can correspond to Is_not_last_c, and for the data acquisition process of the second 8*C dimension, it can correspond to Is_last_c. That is to say, Is_last_c indicates that the data in the C dimension has been completely acquired, while Is_not_last_c indicates that the data acquisition process in the C dimension has not been completed and there is still data in the C dimension that needs to be acquired.

[0161] Next, the information that needs to be updated in each state and how to update it will be introduced.

[0162] First, in the s0 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table: Table 3

[0163] The parameter update rules in Table 3 are described below. Here, stride_x, stride_y, and stride_z respectively represent the stride values in the W dimension, H dimension, and D dimension. As an implementation, the strides can all be set to 1. In addition, in the examples of Tables 3-5, x = 8 is used for illustration. For example, it is represented as 8C, that is, the size of the C dimension bound to the W dimension is 8*C.

[0164] Specifically, in the solutions of Table 3 and Tables 4-5, it is first necessary to determine whether to jump within the same 8C or between different 8Cs. Figure 9A For example, jumping within the same 8C can be, for example, to obtain data within D0-D1 of the first 8*C in the N0 dimension. Jumping between different 8Cs can be, for example, to jump from the N0 dimension to the first 8*C in the N1 dimension to obtain data. That is, for some implementation manners of the present disclosure, it is necessary to distinguish whether to jump the 8*C dimension. This is because in the data loading method of the present application, the tensor to be processed is loaded from the original tensor with the data storage format of N(C / x)DHW(xC), and the tensor to be processed is placed in the buffer area in the NDHWC data storage format. Due to the change in the data storage format, the priority of the C dimension has changed, which also makes it necessary to reserve in advance the request addresses for this part of the C dimension during the process of loading data for multiple requests based on partitioning. The specific implementation process is reflected in the state update and jump conditions of the state machine.

[0165] According to Table 3, first, for the W dimension, (1) for the case of jumping within the same 8C, if jumping from state s0 to state s0, the coordinate w_coord of the W dimension is updated to left_bound, where left_bound represents the left boundary value set for the W dimension. If jumping from state s0 to state s1, the coordinate w_coord of the W dimension is updated to 0; (2) for the case of jumping between different 8Cs, the coordinate w_coord of the W dimension is directly updated to the starting coordinate w_coord_b of the W dimension.

[0166] The coordinate update logic for the other dimensions in Table 3 is the same as that of the W dimension and is described separately as follows.

[0167] For the H dimension, (1) For the case of jumping within the same 8C, if jumping from state s0 to state s0, when h_coord + stride_y >= down_bound, the coordinate h_coord of the H dimension is updated to h_coord + stride_y - copy_h; otherwise (h_coord + stride_y < down_bound), the coordinate h_coord of the H dimension is updated to h_coord + stride_y. Otherwise (jumping from state s0 to a state other than state s0), the coordinate h_coord of the H dimension remains unchanged. (2) For the case of jumping between different 8Cs, the coordinate h_coord of the H dimension is directly updated to the starting coordinate h_coord_b of the H dimension.

[0168] For the D dimension, (1) For the case of jumping within the same 8C, if jumping from state s0 to state s0 and h_coord + stride_y >= down_bound, when d_coord + stride_z >= hind_bound, the coordinate d_coord of the D dimension is updated to d_coord + stride_z - copy_d, where hind_bound represents the rear boundary value set for the D dimension; otherwise (d_coord + stride_z < hind_bound), the coordinate d_coord of the D dimension is updated to d_coord + stride_z. Otherwise (jumping from state s0 to a state other than state s0), the coordinate d_coord of the D dimension remains unchanged. (2) For the case of jumping between different 8Cs, the coordinate d_coord of the D dimension is directly updated to the starting coordinate d_coord_b of the D dimension.

[0169] For the C dimension, if it is for jumping between different 8Cs, the coordinate c_coord of the C dimension is updated to c_coord_b + 8; otherwise, the coordinate c_coord of the C dimension remains unchanged. The coordinate update logic of the C dimension is slightly different from that of the W dimension, H dimension, and D dimension because, as described above, the placement method after data loading has changed from the original N(C / x)DHW(xC) to NDHWC, which makes the data in the C dimension have the lowest priority after loading.

[0170] For the N dimension, (1) For the case of jumping within the same 8C, if jumping from state s0 to state s0, at this time h_coord + stride_y >= down_bound and d_coord + stride_z >= hind_bound, the coordinate of the N dimension is updated to n_coord + 1; otherwise, the coordinate n_coord of the N dimension remains unchanged. (2) For the case of jumping between different 8Cs, the coordinate n_coord of the N dimension is directly updated to the starting coordinate n_coord_b of the N dimension.

[0171] In addition, in addition to the above dimension coordinate parameters, the state machine can also maintain some other parameters required for the state update process and the request address calculation. For example, the parameter remain_copy_p in Table 3, which represents the remaining number of pixels to be acquired. (1) For the case of jumping within the same 8C, if jumping from state s0 to state s0, then remain_copy_p is updated to remain_copy_p - (right_bound - w_coord), if jumping from state s0 to state s1, then remain_copy_p is updated to remain_copy_p + w_coord. (2) For the case of jumping between different 8Cs, remain_copy_p is updated to copy_pixel_num.

[0172] Second, in the s1 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table: Table 4

[0173] Third, in the s2 state, to prepare for the next state, that is, as the initial state of the next state, the state machine can update the parameters according to the following table: Table 5

[0174] For Table 4 and Table 5, the state update process of each parameter can refer to the description for Table 3, which will not be elaborated here. Based on the above update logic, the update results of Table 4 and Table 5 can be derived similarly.

[0175] The state transitions in the PerH partitioning method during the process of loading the original tensor stored in N(C / x)DHW(xC) as a tensor to be processed in the NDHWC data storage format are described above in conjunction with Tables 1-5. The states involved are s0, s1, and s2, as well as the updated outputs of the state machine in these three states. It can be understood that the principles and update logics of the state machine for other partitioning methods (e.g., PerD, PerC, etc.) are similar to those described above and will not be described here.

[0176] Of course, as the coordinates continue to increase, when the number of acquired data is equal to the number of pixels included in the tensor to be processed, it is determined that the state machine can end the coordinate determination for each request, and the state machine can enter the idle state, which indicates that the data for the current pen instruction has been acquired.

[0177] The specific state transition conditions of the state machine and the updated coordinate content can be changed and set according to actual needs. The logic is similar to the state machine logic described above, and no further examples will be given here.

[0178] In the above embodiment, the requests are partitioned according to whether the tensor can be continuously loaded in the first dimension. The data used for each request is either all located in the memory or all not located in the memory. The partitioning of the data requests is more reasonable, more suitable for the loading and storage of tensor data, greatly improving the bandwidth and efficiency during data access, and thus improving the hardware utilization rate of the computing unit and enhancing the hardware performance.

[0179] According to some embodiments of the present disclosure, another data loading method is provided. Figure 10 The schematic flowchart of the data loading method provided by at least one embodiment of the present disclosure is shown. As Figure 10 shown, the data loading method according to the embodiment of the present disclosure includes step S201 and step S202.

[0180] In step S201: Receive a data loading instruction indicating to execute the loading of the tensor to be processed from the original tensor in the memory to the buffer area.

[0181] According to the embodiment of the present disclosure, the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer area is NDHWC. Here, N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is regarded as a pixel, and is accumulated step by step to the higher dimensions. The channel number dimension does not calculate the number of pixels.

[0182] After parsing the data loading instruction in step S202, the execution unit executes the data loading instruction.

[0183] As Figure 10 shown, wherein, in step S202, the execution unit executes the data loading instruction, including: S2021: Obtain the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; S2022: Combine the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data from the original tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and S2023: Sequentially send a plurality of requests, and sequentially write the data corresponding to each request into the buffer area to load the tensor to be processed into the buffer area.

[0184] For the related descriptions of the tensor to be processed and the original tensor and the specific implementation processes of steps S2021 - S2023, reference may be made to the related descriptions of the foregoing data loading method, and the repeated parts will not be elaborated.

[0185] As an example, the data loading instruction may be a machine instruction, or the data loading instruction may also be a micro-instruction. For example, the data loading instruction is implemented in the form of a Load instruction.

[0186] According to some embodiments of the present disclosure, a data storage method is further provided, which is used to obtain a second tensor based on the first tensor in the buffer area and write the second tensor into the memory. It can be understood that the data storage method according to the embodiments of the present disclosure can be understood as the reverse process of the data loading method described above, that is, moving data from the buffer area to the memory. The implementation principle according to the embodiments of the present disclosure is similar to the above data loading method, and the repeated parts will not be described again. Only the different parts will be described in detail.

[0187] In the data storage method according to the embodiments of the present disclosure, the data storage format of the first tensor in the buffer area is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension represents a pixel, and accumulates gradually to the higher dimensions. The channel number dimension does not calculate the number of pixels.

[0188] As an example, the first tensor may refer to a tensor stored in a buffer in the data storage format of N(C / x)DHW(xC). For the second tensor obtained therefrom, it is placed in the memory in the NDHWC data storage format. In terms of understanding the implementation principle, the first tensor can be correspondingly understood as the original tensor in the data loading method described above, that is, part or all of the data is obtained from the first tensor and transferred to the memory. This part of the tensor obtained from the first tensor is denoted as the second tensor, which can be correspondingly understood as the tensor to be processed in the data loading method described above.

[0189] Figure 11 FIG. shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure, as Figure 11 shown, the data storage method includes steps S301-S303.

[0190] In step S301, obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor. Then, in step S302, combine the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor. In step S303, sequentially send the plurality of requests and sequentially write the data corresponding to each request into the memory to store the second tensor in the memory.

[0191] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, divide the requests in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as or adjacent to the second dimension.

[0192] In the data storage method according to an embodiment of the present disclosure, the manner of sequentially obtaining data from the first tensor can be referred to and described in combination with Figure 6 the description shown. Compared with the block-based data storage method in the related art, the data storage method provided by the embodiment of the present disclosure can achieve per-pixel data storage, rather than Figure 4The acquisition of the whole block in this way is more conducive to the calculation process such as convolution operation and helps improve the operation efficiency. It can be understood that, for example, the processor can reasonably use the continuous data loading method or the whole-block data loading method according to the type of operation to be performed or the data processing characteristics, that is, it can support the adaptive switching between these two loading methods, which will not be further elaborated here. In addition, the memory or buffer can also include corresponding identifiers to indicate the specific acquisition method of this tensor.

[0193] In practical applications, for example, in the convolution operation of general computing operations, the operation process is multi-layered. For example, after the first layer of processing is completed, the data result needs to be stored for use in the next layer of processing. As described above, for the data that needs to perform convolution operations, adopting the sequential data acquisition method provided by the embodiments of the present disclosure (which can refer to the description made above in combination with Figure 5 , Figure 6 , Figure 7A , Figure 7B , Figure 8A , Figure 8B and Figure 8C ) is more conducive to improving the calculation efficiency. As an example, the 1x3 weight data can only obtain data in the form of a sliding window, so it can only be obtained row by row, unless the shape of the graph to be obtained is very regular, exactly a cube and no operations such as padding are required in the next layer.

[0194] According to some embodiments of the present disclosure, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the second tensor in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the second tensor in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

[0195] According to some embodiments of the present disclosure, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the starting coordinate of the second tensor in the first dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the first dimension, it is determined that the second tensor cannot be continuously loaded in the first dimension.

[0196] According to some embodiments of the present disclosure, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the second tensor in the W dimension not being equal to the size of the first tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the W dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the W dimension, it is determined that the second tensor cannot be continuously loaded in the W dimension, and it is determined to adopt a division of requests in the H dimension.

[0197] According to some embodiments of the present disclosure, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the second tensor in the H dimension not being equal to the size of the first tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the H dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the H dimension, it is determined that the second tensor cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

[0198] According to some embodiments of the present disclosure, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the second tensor in the D dimension not being equal to the size of the first tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the second tensor in the D dimension not being equal to the second coordinate value of the starting coordinate of the first tensor in the D dimension, it is determined that the second tensor cannot be continuously loaded in the D dimension. When the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

[0199] According to some embodiments of the present disclosure, each request includes a data read address for indicating the starting position of reading data from the buffer, a data write address for indicating the starting position of writing data to the memory, and the number of pixels to be acquired by the request. The request initial coordinate corresponding to the next request is determined according to the request initial coordinate corresponding to the previous request. The request initial coordinate corresponding to each request is used to determine the data read address, data write address, and the number of pixels to be acquired by the request. Among them, when sequentially determining the request initial coordinates corresponding to each request, the coordinate values of the request initial coordinates are incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0200] According to some embodiments of the present disclosure, determining a plurality of requests for storing a second tensor in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor includes: determining a first request among the plurality of requests and an initial state when the first request enters a state machine based on the number of pixels included in the second tensor and the starting coordinates of the second tensor, where the first request indicates the number of pixels to be acquired by the request; and determining each request after the first request among the plurality of requests by using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the second tensor.

[0201] According to some embodiments of the present disclosure, each request includes a data read address for indicating a starting position for reading data from a buffer, a data write address for indicating a starting position for writing data to a memory, and the number of pixels to be acquired by the request. Among them, determining a first request among the plurality of requests and an initial state when the first request enters a state machine based on the number of pixels included in the second tensor and the starting coordinates of the second tensor includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the second tensor; using the starting coordinates of the second tensor as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the first tensor, the request initial coordinates corresponding to the first request, and the data storage format of the second tensor; determining the starting address for writing the second tensor in the memory as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the second tensor, the data storage format of the first tensor, the starting coordinates of the second tensor, and the shape and size of the first tensor.

[0202] According to some embodiments of the present disclosure, determining the initial state when the first request enters the state machine based on the starting coordinates of the second tensor includes: in response to the first coordinate value of the starting coordinates of the second tensor in the first dimension being less than the second coordinate value of the starting coordinates of the first tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the first tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0203] According to some embodiments of the present disclosure, based on the initial state, in combination with the number of pixels to be acquired in the first request and the number of pixels included in the second tensor, a state machine is used to determine each request after the first request among multiple requests, including: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired in the second request; determining the data read address of the second request according to the shape and size of the first tensor, the request initial coordinates corresponding to the second request, and the data storage format of the first tensor; determining the data write address of the second request according to the number of pixels to be acquired in the second request; updating the state machine based on the information related to the second request, and sequentially determining subsequent requests in each request based on the updated state machine.

[0204] According to some embodiments of the present disclosure, for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Wherein, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the first tensor in five dimensions respectively.

[0205] According to some embodiments of the present disclosure, for multiple requests for storing the second tensor, determining the data write address of the nth request among the multiple requests includes: For m requests with the same coordinate value range of the request data in the C dimension, calculating the data write address Addr_2_i of the ith request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i-1 * copy_c Wherein, b_addr represents the data write address of the first request sent among m requests, req_size_1, req_size_2, …, req_size_i-1 represent the number of pixels to be acquired for each of the first i-1 requests respectively, and copy_c is the size of the second tensor in the C dimension; Wherein, in response to the data of each of the m requests including the first x data elements of the first tensor in the C dimension, b_addr is the starting address b_addr_base for writing the second tensor into the memory, in response to the data of each of the m requests including the t-th data element to the (t + x)-th data element of the first tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1) * x t, m, and i are positive integers.

[0206] According to some embodiments of the present disclosure, a plurality of requests are sequentially sent, and the data corresponding to each request is sequentially written into the memory to store the second tensor into the memory, including: for any request, in response to all the data corresponding to any request being located in the buffer area, sending any request to the buffer area; in response to all the data corresponding to any request not being located in the buffer area, abandoning the request.

[0207] The data storage method according to the embodiments of the present disclosure may further include: obtaining boundary values for at least a part of the five dimensions of the first tensor respectively, and the boundary values are used to define the boundary range for obtaining data from the first tensor. For example, the boundary value may be any value compared to the size value of the first tensor in this dimension. As an example, the boundary value set for the W dimension may be any of the Figure 8A circumstances shown. Further, for the dimension with the boundary value set, the data loaded by each request either all belongs to the data boundary of the first tensor itself and the valid data range defined by the boundary value, or all does not belong to the valid data range defined by both the data boundary of the tensor itself and the set boundary value.

[0208] It can be understood that the data storage method according to the embodiments of the present disclosure can achieve similar technical effects to the data loading method according to the embodiments of the present disclosure.

[0209] According to some embodiments of the present disclosure, another data storage method is provided. Figure 12The figure shows a schematic flowchart of a data storage method provided by at least one embodiment of the present disclosure. As Figure 12 shown, the data storage method according to an embodiment of the present disclosure includes step S401 and step S402.

[0210] In step S401: Receive a data storage instruction indicating to execute obtaining a second tensor based on the first tensor in the buffer and writing the second tensor into memory.

[0211] According to an embodiment of the present disclosure, the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. Among the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not calculate the number of pixels.

[0212] In step S402: After parsing the data storage instruction, use an execution unit to execute the data storage instruction.

[0213] As Figure 12 shown, where step S402 uses an execution unit to execute the data storage instruction, including: S4021: Obtain the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; S4022: Combine the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine multiple requests for loading the second tensor, where the multiple requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and S4023: Sequentially send multiple requests and sequentially write the data corresponding to each request into memory to store the second tensor in memory.

[0214] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, divide the requests in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0215] For the related descriptions of the first tensor and the second tensor and the specific implementation processes of steps S4021 - S4023, reference may be made to the related descriptions of the foregoing data storage method, and repeated parts will not be elaborated.

[0216] As an example, the data storage instruction may be a machine instruction, or the data storage instruction may also be a micro-instruction. For example, the data storage instruction is implemented in the form of a Store instruction.

[0217] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data loading instruction, where the original tensor has a data storage format of N(C / x)DHW(xC) in memory, and the tensor to be processed is stored in the buffer area in a data storage format of NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is represented as a pixel, and is accumulated step by step to higher dimensions, where the channel number dimension does not calculate the number of pixels; and the execution unit is configured to: execute the data loading instruction.

[0218] According to an embodiment of the present disclosure, the execution unit executes the data loading instruction, including: obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; combining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor to determine a plurality of requests for loading the tensor to be processed, where the plurality of requests are used to sequentially obtain data from the original tensor with the starting coordinates as the starting point until the number of obtained data is equal to the number of pixels included in the tensor to be processed; and sequentially sending the plurality of requests to write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area.

[0219] According to an embodiment of the present disclosure, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension such that when loading the tensor to be processed, continuous loading cannot be performed in the first dimension but continuous loading can be performed in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the original tensor or all does not belong to the data range of the original tensor. And in response to the data loaded by the request all belonging to the data range of the original tensor, the data loaded by the request comes from the original tensor and is continuously stored in memory, where the first dimension is the same as or adjacent to the second dimension.

[0220] For example, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the tensor to be processed in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

[0221] For example, in response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be continuously loaded in the first dimension.

[0222] For example, when the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be continuously loaded in the W dimension, and it is determined to adopt a division of requests in the H dimension.

[0223] For example, when the first dimension is the H dimension, the second dimension is the D dimension. In response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

[0224] For example, when the first dimension is the D dimension, the second dimension is the C dimension. In response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be continuously loaded in the D dimension, and when the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

[0225] For example, each request includes a data read address for indicating the starting position of reading data from memory, a data write address for indicating the starting position of writing data to the buffer, and the number of pixels to be acquired by the request. The initial request coordinates corresponding to the next request are determined based on the initial request coordinates corresponding to the previous request. The initial request coordinates corresponding to each request are used to determine the data read address, data write address, and the number of pixels to be acquired by the request. Among them, when sequentially determining the initial request coordinates corresponding to each request, the coordinate values of the initial request coordinates are incrementally updated in the order of the W dimension, H dimension, D dimension, N dimension, and C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0226] For example, in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, multiple requests for loading the tensor to be processed are determined, including: based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determining the first request among the multiple requests and the initial state of the first request entering the state machine, where the first request indicates the number of pixels to be acquired by the request; and based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the tensor to be processed, using the state machine to determine each request among the multiple requests after the first request.

[0227] For example, each request includes a data read address for indicating the starting position of reading data from memory, a data write address for indicating the starting position of writing data to the buffer, and the number of pixels to be acquired by the request. Among them, based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determining the first request among the multiple requests and the initial state of the first request entering the state machine includes: based on the starting coordinates of the tensor to be processed, determining the initial state of the first request entering the state machine; using the starting coordinates of the tensor to be processed as the initial request coordinates corresponding to the first request; according to the shape and size of the original tensor, the initial request coordinates corresponding to the first request, and the data storage format of the tensor to be processed, determining the data read address of the first request; determining the starting address for writing the tensor to be processed in the buffer as the data write address of the first request; and according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor, determining the number of pixels to be acquired by the first request.

[0228] For example, based on the starting coordinates of the tensor to be processed, determining the initial state when the first request enters the state machine, includes: in response to the first coordinate value of the starting coordinates of the tensor to be processed in the first dimension being less than the second coordinate value of the starting coordinates of the original tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the original tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0229] For example, based on the initial state, in combination with the number of pixels to be obtained by the first request and the number of pixels included in the tensor to be processed, using the state machine to determine each request after the first request among multiple requests, includes: based on the initial state, using the state machine to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be obtained by the second request; according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor, determining the data read address of the second request; according to the number of pixels to be obtained by the second request, determining the data write address of the second request; updating the state machine based on the information related to the second request, and sequentially determining subsequent requests in each request based on the updated state machine.

[0230] For example, for multiple requests used to load the tensor to be processed, determining the data read address of the nth request among multiple requests, includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC Wherein, u_addr_base represents the storage address in the memory of the pixel at the starting coordinate position of the original tensor, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the original tensor in 5 dimensions respectively.

[0231] For example, for multiple requests for loading tensors to be processed, determining the data write address of the nth request among the multiple requests includes: for m requests with the same coordinate value range of the data of the requests in the C dimension, calculating the data write address Addr_2_i of the ith request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i-1 * copy_c where b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2, …, req_size_i-1 represent the number of pixels to be acquired by the respective previous i-1 requests, and copy_c is the size of the tensor to be processed in the C dimension; wherein, in response to the data of each of the m requests including the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base for writing the tensor to be processed in the buffer, and in response to the data of each of the m requests including the tth data element to the (t + x)th data element of the original tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1) * x t, m, and i are positive integers.

[0232] For example, sequentially sending multiple requests and sequentially writing the data corresponding to each request to the buffer to load the tensor to be processed into the buffer includes: for any request, in response to all the data loaded by any request being located in the memory, sending any request to the memory; in response to all the data loaded by any request not being located in the memory, converting any request into writing multiple predetermined values to the buffer, where the number of the multiple predetermined values is determined by the number of pixels to be acquired by any request.

[0233] The data storage method according to an embodiment of the present disclosure may further include: obtaining boundary values for at least a part of the five dimensions of the first tensor respectively, the boundary values being used to define the boundary range for acquiring data from the first tensor, and for the dimension for which the boundary values are set, the data loaded by each request belongs to the valid data range defined by the original tensor and the boundary values or does not belong to the valid data range.

[0234] According to some embodiments of the present disclosure, a processor is further provided, including an instruction parsing unit and an execution unit. The instruction parsing unit is configured to: receive and parse a data storage instruction, where the data storage instruction instructs to execute obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into memory. The data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in memory is NDHWC, where N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data represented by the height dimension and the width dimension is a pixel, and it accumulates to higher dimensions step by step, where the channel number dimension does not calculate the number of pixels; and the execution unit is configured to: execute the data storage instruction.

[0235] According to an embodiment of the present disclosure, the execution unit executes the data storage instruction, including: obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; combining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor to determine a plurality of requests for loading the second tensor, where the plurality of requests are used to sequentially obtain data from the first tensor starting from the starting coordinates until the number of obtained data is equal to the number of pixels included in the second tensor; and sequentially sending the plurality of requests to sequentially write the data corresponding to each request into memory to store the second tensor in memory.

[0236] According to an embodiment of the present disclosure, in response to the size relationship between the second tensor and the first tensor in the first dimension such that when loading the second tensor, it cannot be continuously loaded in the first dimension but can be continuously loaded in dimensions lower than the first dimension, requests are divided in the second dimension, and the data loaded by each request either all belongs to the data range of the first tensor or all does not belong to the data range of the first tensor. And in response to the data loaded by the request all belonging to the data range of the first tensor, the data loaded by the request comes from the first tensor and is continuously stored in the buffer, where the first dimension is the same as the second dimension or the first dimension is adjacent to the second dimension.

[0237] For example, in response to the first dimension being the W dimension, the H dimension, or the D dimension, the second dimension is adjacent to the first dimension and the first dimension has priority over the second dimension during loading. In response to the first dimension being the C dimension or the first dimension being the N dimension and the size of the second tensor in the C dimension being greater than x, the second dimension is the C dimension. In response to the first dimension being the N dimension and the size of the second tensor in the C dimension being equal to x, the second dimension is the overall dimension higher than the N dimension.

[0238] For example, in response to the size of the second tensor in the first dimension not being equal to the size of the first tensor in the first dimension, and / or the starting coordinate of the second tensor in the first dimension having a first coordinate value that is not equal to the starting coordinate of the first tensor in the first dimension having a second coordinate value in the first dimension, it is determined that the second tensor cannot be continuously loaded in the first dimension.

[0239] For example, when the first dimension is the W dimension and the second dimension is the H dimension, in response to the size of the second tensor in the W dimension not being equal to the size of the first tensor in the W dimension, and / or the starting coordinate of the second tensor in the W dimension having a first coordinate value that is not equal to the starting coordinate of the first tensor in the W dimension having a second coordinate value in the W dimension, it is determined that the second tensor cannot be continuously loaded in the W dimension, and it is determined to adopt a division of requests in the H dimension.

[0240] For example, when the first dimension is the H dimension and the second dimension is the D dimension, in response to the size of the second tensor in the H dimension not being equal to the size of the first tensor in the H dimension, and / or the starting coordinate of the second tensor in the H dimension having a first coordinate value that is not equal to the starting coordinate of the first tensor in the H dimension having a second coordinate value in the H dimension, it is determined that the second tensor cannot be continuously loaded in the H dimension, and it is determined to adopt a division of requests in the D dimension.

[0241] For example, when the first dimension is the D dimension and the second dimension is the C dimension, in response to the size of the second tensor in the D dimension not being equal to the size of the first tensor in the D dimension, and / or the starting coordinate of the second tensor in the D dimension having a first coordinate value that is not equal to the starting coordinate of the first tensor in the D dimension having a second coordinate value in the D dimension, it is determined that the second tensor cannot be continuously loaded in the D dimension, and when the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt a division of requests in the C dimension.

[0242] For example, each request includes a data read address for indicating the starting position of reading data from the buffer, a data write address for indicating the starting position of writing data to the memory, and the number of pixels to be acquired by the request. The request initial coordinate corresponding to the next request is determined based on the request initial coordinate corresponding to the previous request. The request initial coordinate corresponding to each request is used to determine the data read address, the data write address, and the number of pixels to be acquired by the request. Among them, when sequentially determining the request initial coordinates corresponding to each request, the coordinate values of the request initial coordinates are incrementally updated in the order of the W dimension, the H dimension, the D dimension, the N dimension, and the C dimension, and when the coordinate value of the C dimension is updated, it is incremented by x.

[0243] For example, determining a plurality of requests for storing a second tensor in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor includes: determining a first request among the plurality of requests and an initial state when the first request enters a state machine based on the number of pixels included in the second tensor and the starting coordinates of the second tensor, where the first request indicates the number of pixels to be acquired by the request; and determining each request after the first request among the plurality of requests using the state machine based on the initial state, in combination with the number of pixels to be acquired by the first request and the number of pixels included in the second tensor.

[0244] For example, each request includes a data read address for indicating a starting position for reading data from a buffer, a data write address for indicating a starting position for writing data to a memory, and the number of pixels to be acquired by the request. Among them, determining a first request among the plurality of requests and an initial state when the first request enters a state machine based on the number of pixels included in the second tensor and the starting coordinates of the second tensor includes: determining the initial state when the first request enters the state machine based on the starting coordinates of the second tensor; taking the starting coordinates of the second tensor as the request initial coordinates corresponding to the first request; determining the data read address of the first request according to the shape and size of the first tensor, the request initial coordinates corresponding to the first request, and the data storage format of the second tensor; determining the starting address for writing the second tensor in the memory as the data write address of the first request; and determining the number of pixels to be acquired by the first request according to the number of pixels included in the second tensor, the data storage format of the first tensor, the starting coordinates of the second tensor, and the shape and size of the first tensor.

[0245] For example, determining the initial state when the first request enters the state machine based on the starting coordinates of the second tensor includes: in response to the first coordinate value of the starting coordinates of the second tensor in the first dimension being less than the second coordinate value of the starting coordinates of the first tensor in the first dimension, determining the initial state as the first state; in response to the first coordinate value being greater than or equal to the second coordinate value and less than the third coordinate value, determining the initial state as the second state, where the difference between the second coordinate value and the third coordinate value is equal to the size of the first tensor in the first dimension; and in response to the first coordinate value being greater than or equal to the third coordinate value, determining the initial state as the third state.

[0246] For example, based on the initial state, in combination with the number of pixels to be acquired for the first request and the number of pixels included in the second tensor, a state machine is used to determine each request among multiple requests that comes after the first request, including: based on the initial state, the state machine is used to determine the request initial coordinates corresponding to the second request after the first request and the number of pixels to be acquired for the second request; according to the shape and size of the first tensor, the request initial coordinates corresponding to the second request, and the data storage format of the first tensor, the data read address of the second request is determined; according to the number of pixels to be acquired for the second request, the data write address of the second request is determined; the state machine is updated based on the information related to the second request, and subsequent requests among each request are sequentially determined based on the updated state machine.

[0247] For example, for multiple requests for storing the second tensor, determining the data read address of the nth request among the multiple requests includes: Calculating the data read address Addr1_n of the nth request according to the following formula: Addr1_n = u_addr_base + n_coord * tensor_c * tensor_d * tensor_h * tensor_w + c_coord * tensor_d * tensor_h * tensor_w + d_coord * tensor_h * tensor_w * xC + h_coord * tensor_w * xC + w_coord * xC wherein, u_addr_base represents the storage address of the pixel at the starting coordinate position of the first tensor in the buffer, n_coord, d_coord, h_coord, w_coord, c_coord represent the request initial coordinates corresponding to the nth request, and tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape and size of the first tensor in 5 dimensions respectively.

[0248] For example, for multiple requests for storing the second tensor, determining the data write address of the nth request among the multiple requests includes: For m requests with the same coordinate value range of the request data in the C dimension, calculating the data write address Addr_2_i of the ith request sent among the m requests according to the following formula: Addr_2_i = b_addr + req_size_1 * copy_c + req_size_2 * copy_c + … + req_size_i-1 * copy_c Wherein, b_addr represents the data write address of the first request sent among m requests, req_size_1, req_size_2, ..., req_size_i-1 represent the number of pixels to be acquired for each of the first i-1 requests respectively, and copy_c is the size of the second tensor in the C dimension; Wherein, in response to the data of each of the m requests including the first x data elements of the first tensor in the C dimension, b_addr is the starting address b_addr_base for writing the second tensor into the memory, In response to the data of each of the m requests including the t-th data element to the (t + x)-th data element of the first tensor in the C dimension, where t is greater than x, b_addr is calculated according to the following formula: b_addr = b_addr_base + (t / x - 1) * x t, m, and i are positive integers.

[0249] For example, when sending multiple requests in sequence and writing the data corresponding to each request into the memory in order to store the second tensor into the memory, it includes: for any request, in response to all the data corresponding to any request being located in the buffer, sending any request to the buffer; in response to all the data corresponding to any request not being located in the buffer, then abandoning the request.

[0250] As an example, Figure 13 shows a schematic block diagram of a processor according to some embodiments of the present disclosure. As Figure 13 shown, the processor 1000 may include an instruction parsing unit 1010 and an execution unit 1020. It can be understood that the processor 1000 may be implemented to execute the data loading method according to the embodiments of the present disclosure to load the tensor to be processed from the original tensor in the memory into the buffer, or implement the data storage method according to the embodiments of the present disclosure to obtain the second tensor based on the first tensor in the buffer and write the second tensor into the memory.

[0251] Regarding the specific implementation processes of the data storage method and the data loading method, reference can be made to the above description and will not be repeated here. The processor provided by at least one embodiment of the present disclosure can achieve technical effects similar to those of the aforementioned data loading method / data storage method, and the repeated parts will not be elaborated.

[0252] According to some embodiments of the present disclosure, an electronic device is further provided. Figure 14A schematic block diagram of an electronic device according to some embodiments of the present disclosure is shown. As Figure 14 shown, the electronic device 2000 may include a processor 2010 and a memory 2020 connected to the processor 2010. In addition, the processor 2010 may further include a buffer. According to an embodiment of the present disclosure, the memory 2020 may be implemented in the form of a high-bandwidth memory HBM, which is not limited thereto. Specifically, according to an embodiment of the present disclosure, the processor 2010 is configured to run computer-executable instructions, which, when run by the processor 2010, implement a data loading method according to an embodiment of the present disclosure to load a tensor to be processed from an original tensor in the memory into the buffer, or implement a data storage method according to an embodiment of the present disclosure to obtain a second tensor based on a first tensor in the buffer and write the second tensor into the memory.

[0253] The processor 2010 may perform various actions and processes according to a program stored in a non-transitory memory. Specifically, the processor 2010 may refer to a processor chip capable of performing parallel computing. For example, it may be any one of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural network Processing Unit), a DPU (Deep learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). In addition, the processor 2010 may also be implemented as other conventional types of processors, which are not limited herein.

[0254] Regarding the specific implementation processes of the data storage method and the data loading method, reference may be made to the above description and will not be repeated herein. The processor provided by at least one embodiment of the present disclosure may achieve technical effects similar to those of the foregoing data loading method / data storage method, and the repeated parts will not be elaborated.

[0255] Figure 15 A block diagram of an example computing device implementing some embodiments of the present disclosure is shown. As Figure 15 shown, the computing device 3000 is, for example, suitable for implementing the data loading method or the data storage method provided by the embodiments of the present disclosure. It should be noted that Figure 15The components of the computing device 3000 shown are merely exemplary and not restrictive. According to actual application requirements, the computing device 3000 may also have other components.

[0256] As Figure 15 shown, the computing device 3000 may include a processing device 3010 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.

[0257] For example, when the computer-readable instructions are run by the processing device 3010, one or more steps of the data loading method described in any of the above embodiments, or one or more steps of the data storage method described in any of the above embodiments, may be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference may be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference may be made to the relevant descriptions in the embodiments of the data storage method above.

[0258] For example, the processing device 3010, the read-only memory (ROM) 3020, and the random access memory (RAM) 3030 are connected to each other through a bus 3040. The input / output (I / O) interface 3050 is also connected to the bus 3040.

[0259] For example, the memory may include any combination of one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 3030 and / or cache memory, etc. For example, the computer-readable instructions may be loaded from the storage device 3080 into the random access memory (RAM) 3030 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 3020, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage medium, as well as various data used and / or generated by the application programs, etc.

[0260] Typically, the following devices can be connected to the I / O interface 3050: input devices 3060 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 3070 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 3080 including, for example, magnetic tapes, hard disks, flash memories, etc.; and communication devices 3090. The communication device 3090 can allow the computing device 3000 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 15 the computing device 3000 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the computing device 3000 can alternatively implement or have more or fewer devices. For example, the processing device 3010 can control other components in the computing device 3000 to perform desired functions. The processing device 3010 can be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit (GPU) and other devices with data processing capabilities and / or program execution capabilities. The GPU can be directly integrated into a system on chip (SOC), directly integrated onto the motherboard, or built into the north bridge chip of the motherboard.

[0261] According to some embodiments of the present disclosure, a non-transitory computer-readable storage medium is also provided, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they implement the data loading method according to the embodiments of the present disclosure, or implement the data storage method according to the embodiments of the present disclosure.

[0262] Figure 16 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 16 shown, the computer-readable storage medium 4000 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 4010 can be non-temporarily stored on the storage medium 4000. For example, when the computer-readable instructions 4010 are executed by a processor, one or more steps of the data loading method described in any of the above embodiments, or one or more steps of the data storage method described in any of the above embodiments can be executed. It should be noted that for a detailed description of the processing process of the data loading method, reference can be made to the relevant descriptions in the embodiments of the data loading method above, and for a detailed description of the processing process of the data storage method, reference can be made to the relevant descriptions in the embodiments of the data storage method above.

[0263] As an example, the storage medium 4000 can be applied to the electronic device 2000 and / or the computing device 3000. For example, the storage medium 4000 can be implemented as the storage device 3080 in the computing device 3000.

[0264] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0265] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.

[0266] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, the exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on. The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0267] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0268] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

[0269] Regarding the present disclosure, the following points also need to be noted: (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0270] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0271] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the appended claims.

Claims

1. A data loading method for loading a to-be-processed tensor from an original tensor in memory into a buffer, wherein: The data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the tensor to be processed stored in the buffer area is NDHWC, wherein N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to higher dimensions, wherein the channel number dimension does not calculate the number of pixels, The data loading method comprises: Obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Determine, in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data of the original tensor from the original tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area, In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each data requested to be loaded belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the data loaded by the request all belong to the data range of the original tensor, the data requested to be loaded comes from the original tensor and is stored continuously in the memory, The first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

2. The method according to claim 1, wherein: In response to the first dimension being a W dimension, an H dimension, or a D dimension, the second dimension is adjacent to the first dimension and the first dimension takes precedence over the second dimension when loading, In response to the first dimension being a C dimension or the first dimension being an N dimension and the size of the tensor to be processed in the C dimension being greater than x, the second dimension being the C dimension, and In response to the first dimension being N-dimensional and the size of the tensor to be processed in the C-dimensionality being equal to x, the second dimension is an integral dimension higher than the N-dimensionality.

3. The method according to claim 2, wherein: In response to the size of the tensor to be processed in the first dimension not being equal to the size of the original tensor in the first dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the first dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the first dimension, it is determined that the tensor to be processed cannot be loaded continuously in the first dimension.

4. The method according to claim 1, wherein: In the case where the first dimension is the W dimension, the second dimension is the H dimension. In response to the size of the tensor to be processed in the W dimension not being equal to the size of the original tensor in the W dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the W dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the W dimension, it is determined that the tensor to be processed cannot be loaded continuously in the W dimension, and it is determined to adopt the requested division in the H dimension.

5. The method according to claim 1, wherein: In the case where the first dimension is H dimension, the second dimension is D dimension. In response to the size of the tensor to be processed in the H dimension not being equal to the size of the original tensor in the H dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the H dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the H dimension, it is determined that the tensor to be processed cannot be loaded continuously in the H dimension, and it is determined to adopt the requested division in the D dimension.

6. The method according to claim 1, wherein: When the first dimension is D dimension and the second dimension is C dimension, in response to the size of the tensor to be processed in the D dimension not being equal to the size of the original tensor in the D dimension, and / or the first coordinate value of the starting coordinate of the tensor to be processed in the D dimension not being equal to the second coordinate value of the starting coordinate of the original tensor in the D dimension, it is determined that the tensor to be processed cannot be loaded continuously in the D dimension, and when the size of the tensor to be processed in the C dimension is greater than x, it is determined to adopt the requested division in the C dimension.

7. The method according to claim 1, wherein: The tensor to be processed is used for convolution operation in a computing unit within the processor.

8. The method according to claim 1, wherein: Each request includes a data read address for indicating a starting position for reading data from the memory, a data write address for indicating a starting position for writing data to the buffer area, and the number of pixels to be acquired by the request. Determine the request initial coordinates corresponding to the next request according to the request initial coordinates corresponding to the previous request, and the request initial coordinates corresponding to each request are used to determine the data reading address, data writing address and number of pixels to be obtained by the request. Among them, when determining the requested initial coordinates corresponding to each request in turn, the coordinate values ​​of the requested initial coordinates are updated incrementally in the order of W dimension, H dimension, D dimension, N dimension, and C dimension, and the coordinate value of the C dimension is automatically incremented by x when updated.

9. The method according to claim 1, wherein: Determining multiple requests for loading the tensor to be processed based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, including: Determining a first request among the multiple requests and an initial state in a state machine in which the first request enters, based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, wherein the first request indicates the number of pixels to be obtained by the request; and Based on the initial state, in combination with the number of pixels to be obtained by the first request and the number of pixels included in the tensor to be processed, the state machine is used to determine each request after the first request in the multiple requests.

10. The method according to claim 9, wherein: Each request includes a data read address for indicating a starting position for reading data from the memory, a data write address for indicating a starting position for writing data to the buffer area, and the number of pixels to be acquired by the request. Wherein, based on the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed, determining the first request among the multiple requests and the initial state of the first request entering the state machine includes: Determining, based on the starting coordinates of the tensor to be processed, an initial state in which the first request enters the state machine; Using the starting coordinates of the tensor to be processed as the request initial coordinates corresponding to the first request; Determine a data reading address for the first request according to the shape and size of the original tensor, the request initial coordinates corresponding to the first request, and the data storage format of the tensor to be processed; Determine a starting address in the buffer area for writing the tensor to be processed as a data writing address for the first request; and The number of pixels to be obtained by the first request is determined according to the number of pixels included in the tensor to be processed, the data storage format of the original tensor, the starting coordinates of the tensor to be processed, and the shape and size of the original tensor.

11. The method according to claim 10, wherein: Determining the initial state of the first request entering the state machine based on the starting coordinates of the tensor to be processed includes: In response to a first coordinate value of the starting coordinate of the to-be-processed tensor in the first dimension being less than a second coordinate value of the starting coordinate of the original tensor in the first dimension, determining that the initial state is the first state; In response to the first coordinate value being greater than or equal to the second coordinate value and less than a third coordinate value, determining that the initial state is a second state, wherein a difference between the second coordinate value and the third coordinate value is equal to a size of the original tensor in the first dimension; and In response to the first coordinate value being greater than or equal to the third coordinate value, the initial state is determined to be the third state.

12. The method according to claim 9, wherein: Based on the initial state, in combination with the number of pixels to be obtained by the first request and the number of pixels included in the tensor to be processed, the state machine is used to determine each request after the first request in the multiple requests, including: Based on the initial state, using the state machine to determine the request initial coordinates corresponding to a second request after the first request and the number of pixels to be obtained by the second request; Determine a data reading address for the second request according to the shape and size of the original tensor, the request initial coordinates corresponding to the second request, and the data storage format of the original tensor; Determining a data writing address for the second request according to the number of pixels to be acquired by the second request; The state machine is updated according to information related to the second request, and subsequent requests among the requests are sequentially determined based on the updated state machine.

13. The method according to claim 10, wherein: For a plurality of requests for loading the tensor to be processed, determining a data read address of an nth request among the plurality of requests comprises: The data read address Addr1_n of the nth request is calculated according to the following formula: Addr1_n=u_addr_base+ n_coord*tensor_c*tensor_d*tensor_h*tensor_w+ c_coord *tensor_d*tensor_h*tensor_w+ d_coord *tensor_h*tensor_w*xC+ h_coord *tensor_w*xC+ w_coord *xC Among them, u_addr_base represents the storage address of the pixel at the starting coordinate position of the original tensor in the memory, n_coord, d_coord, h_coord, w_coord, c_coord represent the requested initial coordinates corresponding to the nth request, tensor_n, tensor_d, tensor_h, tensor_w, tensor_c represent the shape sizes of the original tensor in five dimensions respectively.

14. The method according to claim 10, wherein: For a plurality of requests for loading the tensor to be processed, determining a data write address of an nth request among the plurality of requests includes: For m requests whose requested data have the same coordinate value range in the C dimension, the data write address Addr_2_i of the i-th request sent among the m requests is calculated according to the following formula: Addr_2_i=b_addr + req_size_1*copy_c + req_size_2*copy_c +…+ req_size_i-1*copy_c Wherein, b_addr represents the data write address of the first request sent among the m requests, req_size_1, req_size_2, ..., req_size_i-1 represent the number of pixels to be obtained by each of the first i-1 requests, and copy_c is the size of the tensor to be processed in the C dimension; The data in response to each of the m requests includes the first x data elements of the original tensor in the C dimension, b_addr is the starting address b_addr_base of the tensor to be processed written into the buffer area, The data in response to each of the m requests includes the t-th data element to the t+x-th data element of the original tensor in the C dimension, where t is greater than x, and b_addr is calculated according to the following formula: b_addr = b_addr_base+ (t / x-1)*x t, m and i are positive integers.

15. The method according to claim 1, wherein: Sending the multiple requests in sequence, and writing the data corresponding to each request into the buffer area in sequence, so as to load the tensor to be processed into the buffer area, includes: For any request, all the data loaded in response to the any request are located in the memory, and the any request is sent to the memory; In response to the data loaded by any one of the requests not being located in the memory, the any one of the requests is converted into writing a plurality of predetermined values ​​into the buffer area, wherein the number of the plurality of predetermined values ​​is determined by the number of pixels to be acquired by the any one of the requests.

16. The method according to claim 1, further comprising: Obtain boundary values ​​for the original tensor in at least a portion of the five dimensions, respectively, where the boundary values ​​are used to limit a boundary range for obtaining data from the original tensor. Among them, for the dimension with a boundary value set, each data requested to be loaded belongs to the original tensor and the valid data range specified by the boundary value or does not belong to the valid data range.

17. A data loading method, comprising: Receive a data loading instruction instructing execution of loading a to-be-processed tensor from an original tensor in a memory into a cache area, wherein the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the to-be-processed tensor stored in the cache area is NDHWC, wherein N represents a batch dimension, D represents a depth dimension, W represents a width dimension, H represents a height dimension, C represents a channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension, wherein in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step toward higher dimensions, wherein the channel number dimension does not calculate the number of pixels; and After parsing the data loading instruction, using the execution unit to execute the data loading instruction, Wherein, using the execution unit to execute the data loading instruction includes: Obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Determine, in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data of the original tensor from the original tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area, In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each data requested to be loaded belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the data loaded by the request all belong to the data range of the original tensor, the data requested to be loaded comes from the original tensor and is stored continuously in the memory, The first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

18. A data storage method for acquiring a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, wherein: The data storage format of the first tensor in the cache area is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, wherein N represents the batch dimension, D represents the depth dimension, W represents the width dimension, H represents the height dimension, C represents the channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension. In the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension are represented as a pixel and accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels. The data storage method comprises: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Determine, in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, a plurality of requests for loading the second tensor, wherein the plurality of requests are used to sequentially obtain data of the first tensor from the first tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request in sequence into the memory to store the second tensor into the memory, In which, in response to the size relationship between the second tensor and the first tensor in the first dimension, when loading the second tensor, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each requested loaded data belongs to the data range of the first tensor or does not belong to the data range of the first tensor, and the data loaded in response to the request belongs to the data range of the first tensor, the data requested to be loaded comes from the first tensor and is stored continuously in the cache area, The first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

19. A data storage method, comprising: Receive a data storage instruction instructing execution of obtaining a second tensor based on a first tensor in a buffer and writing the second tensor into a memory, wherein the data storage format of the first tensor in the buffer is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, wherein N represents a batch dimension, D represents a depth dimension, W represents a width dimension, H represents a height dimension, C represents a channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension, wherein, in the data storage formats of N(C / x)DHW(xC) and NDHWC, data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and After parsing the data storage instruction, use the execution unit to execute the data storage instruction, The step of using the execution unit to execute the data storage instruction includes: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Determine, in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, a plurality of requests for loading the second tensor, wherein the plurality of requests are used to sequentially obtain data of the first tensor from the first tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request in sequence into the memory to store the second tensor into the memory, In which, in response to the size relationship between the second tensor and the first tensor in the first dimension, when loading the second tensor, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each requested loaded data belongs to the data range of the first tensor or does not belong to the data range of the first tensor, and the data loaded in response to the request belongs to the data range of the first tensor, the data requested to be loaded comes from the first tensor and is stored continuously in the cache area, The first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

20. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is configured to: receive and parse a data loading instruction, wherein the data loading instruction instructs execution to load a to-be-processed tensor from an original tensor in a memory into a cache area, wherein the data storage format of the original tensor in the memory is N(C / x)DHW(xC), and the data storage format of the to-be-processed tensor stored in the cache area is NDHWC, wherein N represents a batch dimension, D represents a depth dimension, W represents a width dimension, H represents a height dimension, C represents a channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension, wherein in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and The execution unit is configured to: execute the data loading instruction, The execution unit executes the data loading instruction, including: Obtaining the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor; Determine, in combination with the number of pixels included in the tensor to be processed and the starting coordinates of the tensor to be processed in the coordinate system determined by the original tensor, a plurality of requests for loading the tensor to be processed, wherein the plurality of requests are used to sequentially obtain data of the original tensor from the original tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the tensor to be processed; and Send the multiple requests in sequence, and write the data corresponding to each request into the buffer area in sequence to load the tensor to be processed into the buffer area, In which, in response to the size relationship between the tensor to be processed and the original tensor in the first dimension, when loading the tensor to be processed, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each data requested to be loaded belongs to the data range of the original tensor or does not belong to the data range of the original tensor, and in response to the data loaded by the request all belong to the data range of the original tensor, the data requested to be loaded comes from the original tensor and is stored continuously in the memory, The first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

21. A processor comprising an instruction parsing unit and an execution unit, wherein: The instruction parsing unit is configured to: receive and parse a data storage instruction, wherein the data storage instruction instructs execution to obtain a second tensor based on a first tensor in a cache area and write the second tensor into a memory, wherein the data storage format of the first tensor in the cache area is N(C / x)DHW(xC), and the data storage format of the second tensor stored in the memory is NDHWC, wherein N represents a batch dimension, D represents a depth dimension, W represents a width dimension, H represents a height dimension, C represents a channel number dimension, x is a positive integer, and the number of channels of xC is bound to the width dimension, wherein in the data storage formats of N(C / x)DHW(xC) and NDHWC, the data determined by the height dimension and the width dimension are represented as a pixel, and are accumulated step by step to a higher dimension, wherein the channel number dimension does not calculate the number of pixels; and The execution unit is configured to: execute the data storage instruction, The execution unit executes the data storage instruction, including: Obtaining the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor; Determine, in combination with the number of pixels included in the second tensor and the starting coordinates of the second tensor in the coordinate system determined by the first tensor, a plurality of requests for loading the second tensor, wherein the plurality of requests are used to sequentially obtain data of the first tensor from the first tensor starting from the starting coordinates until the number of data obtained is equal to the number of pixels included in the second tensor; and Send the multiple requests in sequence, and write the data corresponding to each request in sequence into the memory to store the second tensor into the memory, In which, in response to the size relationship between the second tensor and the first tensor in the first dimension, when loading the second tensor, the first dimension cannot be loaded continuously but the dimension lower than the first dimension can be loaded continuously, the request is divided in the second dimension, each requested loaded data belongs to the data range of the first tensor or does not belong to the data range of the first tensor, and the data loaded in response to the request belongs to the data range of the first tensor, the data requested to be loaded comes from the first tensor and is stored continuously in the cache area, The first dimension is the same as the second dimension, or the first dimension is adjacent to the second dimension.

22. An electronic device comprising a processor and a memory connected to the processor, wherein: The processor includes a cache area, wherein The processor is configured to run computer-executable instructions, which, when executed by the processor, implement the data loading method according to any one of claims 1-17 to load the tensor to be processed from the original tensor in the memory to the cache area, or implement the data storage method according to claim 18 or 19 to obtain the second tensor based on the first tensor in the cache area and write the second tensor to the memory.

23. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the data loading method according to any one of claims 1 to 17 is implemented, or the data storage method according to claim 18 or 19 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN116822612A

  • Data processing method and device, processor, electronic equipment and storage medium

    CN119312003A

  • Tensor-filled storage space compression layout method, device and equipment

    CN119902993A

  • Processor, chip product, computer equipment and tensor processing method

    CN119917166A

Cited By

  • Data loading method, processor, electronic equipment and storage medium

    CN120743196A

  • Data loading methods, processors, electronic devices, and storage media

    CN120743196B