Data processing method and device, electronic equipment and storage medium
By determining the address information of the second tensor and performing asynchronous processing, the problem of insufficient storage performance in the artificial intelligence chip is solved, and efficient and flexible data access and processing is achieved, which is suitable for feature data in deep learning tasks.
Patent Information
- Application Number
- CN202311865153.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-12-29
AI Technical Summary
In the prior art, artificial intelligence chips have low storage performance in data flow design, resulting in insufficient data access efficiency and flexibility, making it difficult to meet high performance needs.
By determining the address information of the second tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed, the address information of the second tensor is determined and read from the global memory to the local memory, and asynchronous processing is realized in combination with the interrupt mechanism, improving the efficiency and flexibility of data processing.
It realizes fast data reading from global memory to local memory, improves the efficiency and flexibility of tensor storage, supports blocked data and convolution kernel operations, and is suitable for feature data processing in deep learning tasks.
Smart Images

Figure CN120277005A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, an apparatus, an electronic device, and a storage medium. Background Art
[0002] With the continuous expansion of the application fields of artificial intelligence technologies, the requirements for the performance of artificial intelligence chips are also getting higher and higher. As one of the bottlenecks of the performance of artificial intelligence chips, the storage performance is expected to further improve the data stream design of artificial intelligence chips, so that artificial intelligence chips can access data more efficiently and flexibly. Summary of the Invention
[0003] The present disclosure proposes a data processing technical solution.
[0004] According to one aspect of the present disclosure, there is provided a data processing method, including: determining address information of a second tensor according to tensor information of a first tensor and access information of a second tensor to be accessed; reading the second tensor from a global memory storing the first tensor to a local memory according to the address information of the second tensor; and triggering an interrupt in response to the local memory receiving the second tensor.
[0005] In a possible implementation manner, when the second tensor to be accessed is a block data of the first tensor, the access information includes: a position of the second tensor in the first tensor and a size of the second tensor. Determining the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed includes: determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.
[0006] In a possible implementation, the method further includes: performing a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, so as to obtain a check result of the second tensor; in the case where the check result indicates that a first part of the second tensor is within the first tensor and a second part of the second tensor is outside the first tensor, determining address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; the tensor information of the first tensor includes a padding flag, and reading the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor, including: reading the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor, padding the second part of the second tensor according to the padding flag to obtain the second part of the second tensor; writing the second part of the second tensor to the local memory.
[0007] In a possible implementation, when the second tensor to be accessed is the covered data corresponding to any one or more elements in the convolution kernel in the first tensor, the access information includes: the unfolded matrix coordinates and the convolution kernel element coordinates, the unfolded matrix is a matrix formed by sequentially converting the covered data corresponding to the convolution kernel receptive field into corresponding column vectors or row vectors in the first tensor and arranging them row by row or column by column according to the description information, the description information includes convolution description information and deconvolution description information, and determining the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed includes: obtaining the mapping relationship between the unfolded matrix and the first tensor; determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates.
[0008] In a possible implementation, determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates includes: determining the second tensor to be accessed from the unfolded matrix according to the unfolded matrix coordinates and the convolution kernel element coordinates; determining multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolded matrix and the first tensor; determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.
[0009] In a possible implementation, reading the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor includes: when the size of the second tensor is greater than the read bandwidth, according to the address information of the second tensor, splitting a first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests, each second memory access request being used to read split data of different parts of the second tensor from the global memory to the local memory; and serially writing multiple split data constituting the second tensor in the global memory to the local memory according to the multiple second memory access requests.
[0010] In a possible implementation, the tensor information of the first tensor is stored in the global memory, and the access information of the second tensor to be accessed is stored in a register.
[0011] In a possible implementation, the first tensor includes feature data in a deep learning task, and the feature data includes at least one of image feature data, speech feature data, and text feature data.
[0012] According to one aspect of the present disclosure, a data processing apparatus is provided, including: a determination module configured to determine address information of a second tensor according to tensor information of a first tensor and access information of the second tensor to be accessed; a reading module configured to read the second tensor from a global memory storing the first tensor to a local memory according to the address information of the second tensor; and a triggering module configured to trigger an interruption in response to the local memory receiving the second tensor.
[0013] In a possible implementation, the determination module is configured to: when the second tensor to be accessed is block data of the first tensor, and the access information includes the position of the second tensor in the first tensor and the size of the second tensor, determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.
[0014] In a possible implementation, the determining module is further configured to: perform a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, so as to obtain a check result of the second tensor; in the case that the check result indicates that a first part of the second tensor is within the first tensor and a second part of the second tensor is outside the first tensor, determine the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; the reading module is configured to: in the case that the tensor information of the first tensor includes a padding flag, read the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor, and pad the second part of the second tensor according to the padding flag to obtain the second part of the second tensor; write the second part of the second tensor to the local memory.
[0015] In a possible implementation, the determining module is configured to: when the second tensor to be accessed is the covered data corresponding to any one or more elements in the convolution kernel in the first tensor, and the access information includes the unfolded matrix coordinates and the convolution kernel element coordinates, obtain the mapping relationship between the unfolded matrix and the first tensor; determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates; wherein, the unfolded matrix is a matrix formed by sequentially converting the covered data corresponding to the convolution kernel receptive field into corresponding column vectors or row vectors according to the description information in the first tensor and arranging them row by row or column by column, and the description information includes convolution description information and deconvolution description information.
[0016] In a possible implementation, determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates includes: determining the second tensor to be accessed from the unfolded matrix according to the unfolded matrix coordinates and the convolution kernel element coordinates; determining a plurality of elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolded matrix and the first tensor; determining the address information of the plurality of elements corresponding to the second tensor in the first tensor as the address information of the second tensor.
[0017] In a possible implementation, the reading module is configured to: when the size of the second tensor is greater than the reading bandwidth, split a first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests according to the address information of the second tensor, where each second memory access request is for reading split data of different parts of the second tensor from the global memory to the local memory; and serially write multiple split data that constitute the second tensor in the global memory to the local memory according to the multiple second memory access requests.
[0018] In a possible implementation, the tensor information of the first tensor is stored in the global memory, and the access information of the second tensor to be accessed is stored in a register.
[0019] In a possible implementation, the first tensor includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
[0020] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.
[0021] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0022] According to one aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes to implement the above method.
[0023] In an embodiment of the present disclosure, according to the tensor information of the first tensor and the access information of the second tensor to be accessed, the address information of the second tensor is determined, and according to the address information of the second tensor, the second tensor is read from the global memory storing the first tensor to the local memory; in response to the local memory receiving the second tensor, an interruption is triggered. In this way, the second tensor can be quickly read from the global memory storing the first tensor to the local memory according to the tensor information of the first tensor and the access information of the second tensor to be accessed, improving the efficiency and flexibility of tensor storage, and moreover, each time the local memory receives a second tensor, an interruption can be triggered to realize the asynchrony of the tensor reading process.
[0024] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0026] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure.
[0027] Figure 2 A schematic diagram showing a hardware architecture according to an embodiment of the present disclosure.
[0028] Figure 3 A schematic diagram showing a method for determining address information of a second tensor according to an embodiment of the present disclosure.
[0029] Figure 4 A schematic diagram showing another method for determining address information of a second tensor according to an embodiment of the present disclosure.
[0030] Figure 5 A schematic diagram showing a mapping relationship between a first tensor and an unfolded matrix according to an embodiment of the present disclosure.
[0031] Figure 6 A block diagram showing a data processing apparatus according to an embodiment of the present disclosure.
[0032] Figure 7 A block diagram showing an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0034] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0035] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C.
[0036] In addition, to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0037] With the wide application of deep learning and neural networks, data is often stored in the form of tensors (Tensors), and tensors can generalize vectors and matrices to any dimension. In related technologies, a processor (such as including a General-Purpose Computing on Graphics Processing Units (GPGPU)) needs to decompose the storage space storing tensors when accessing the storage space storing tensors. Each thread of the processor accesses a part of it, and during the process of each thread accessing the part of the storage space it is responsible for, a large number of address calculation instructions and boundary check instructions need to be added. This way of accessing the tensor storage space has low efficiency. At the same time, since a single instruction has a limit on the storage space size when accessing the storage space, accessing tensors occupying a large storage space requires the cooperation of multiple instructions, which is not flexible enough during use and also affects the improvement of memory access performance.
[0038] In view of this, in order to improve the efficiency and flexibility of tensor storage, an embodiment of the present disclosure provides a data processing method. Figure 1 The flowchart showing the data processing method according to an embodiment of the present disclosure is as Figure 1 shown, and the data processing method includes:
[0039] In step S11, according to the tensor information of the first tensor and the access information of the second tensor to be accessed, determine the address information of the second tensor;
[0040] In step S12, according to the address information of the second tensor, read the second tensor from the global memory storing the first tensor to the local memory;
[0041] In step S13, in response to the local memory receiving the second tensor, trigger an interrupt.
[0042] The data processing method according to the embodiments of the present disclosure can quickly read a second tensor from a global memory storing a first tensor to a local memory according to the tensor information of the first tensor and the access information of the second tensor to be accessed, improving the efficiency and flexibility of tensor storage. Moreover, each time the local memory receives a second tensor, an interrupt can be triggered to implement the asynchrony of the tensor reading process.
[0043] In a possible implementation manner, the data processing method according to the embodiments of the present disclosure can be executed by a tensor storage engine, which is a hardware module for implementing data transfer in a processor chip. Herein, the processor chip can be a newly designed one or an improved one based on an existing processor chip. The types of processor chips can include, but are not limited to: Central Processing Unit (CPU), Graphic Processing Unit (GPU), General-Purpose Computing on Graphics Processing Units (GPGPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Tensor Processing Unit (TPU), Field Programmable Gate Array (FPGA) or other programmable logic devices, and can also include a microprocessor or a processor of other conventional processors.
[0044] In a possible implementation manner, one or more tensor storage engines can be included in the processor chip, and the present disclosure does not limit the number of tensor storage engines in the processor chip.
[0045] In an example, for a single-core and single-thread processor chip, a tensor storage engine can be set in the processor chip to perform data transfer processing on the first tensor stored in the global memory outside the processor chip, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine inside the processor chip.
[0046] In an example, for a multi-core and multi-threaded processor chip, multiple tensor storage engines can be set in the processor chip. For example, assume that the processor chip includes N (N≥2) computing cores (Cores), and each computing core can execute M (M≥2) threads in parallel. A tensor storage engine can be set for each computing core respectively. The M threads in each computing core can share a tensor storage engine and a local memory. The tensor storage engine is used to perform data transfer processing on a first tensor stored in a global memory shared by multiple computing cores, and store a second tensor determined based on the first tensor in the local memory within the computing core. Alternatively, M tensor storage engines can also be set for each computing core respectively. Each thread in each computing core can correspond to a tensor storage engine and a local memory. The tensor storage engine is used to perform data transfer processing on the first tensor stored in the global memory shared by multiple computing cores, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine within the computing core.
[0047] In an example, for a multi-core and single-threaded processor chip, multiple tensor storage engines can be set in the processor chip. For example, assume that the processor chip includes N (N≥2) single-threaded computing cores (Cores). A tensor storage engine can be set for each computing core respectively. The tensor storage engine is used to perform data transfer processing on a first tensor stored in a global memory shared by multiple computing cores, and store the second tensor determined based on the first tensor in the local memory within the computing core.
[0048] In an example, for a single-core and multi-threaded processor chip, multiple tensor storage engines can be set in the processor chip. For example, assume that the processor chip can execute M (M≥2) threads in parallel. Each thread can correspond to a tensor storage engine and a local memory. Each tensor storage engine can be used to perform data transfer processing on a first tensor stored in a global memory shared by multiple threads, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine.
[0049] In a possible implementation Figure 2 shows a schematic diagram of a hardware architecture according to an embodiment of the present disclosure, as Figure 2 shown, the tensor storage engine is respectively connected to the global memory and the local memory, and can be used to perform read operations and / or write operations on the global memory and the local memory. Moreover, the tensor storage engine can also be used to receive tensor information of the first tensor and access information of the second tensor, so as to transfer the second tensor from the global memory storing the first tensor to the local memory, and trigger an interruption in response to the local memory receiving the second tensor.
[0050] The following is an exemplary description of the data processing method of the present disclosure embodiment implemented by the tensor storage engine.
[0051] In step S11, according to the tensor information of the first tensor and the access information of the second tensor to be accessed, the address information of the second tensor is determined.
[0052] In a possible implementation manner, the first tensor is a multi-dimensional tensor, which can be regarded as a multi-dimensional matrix. For example, the first tensor can be a two-dimensional tensor, a three-dimensional tensor, a four-dimensional tensor, or other tensors with more dimensions. The embodiments of the present application do not limit the number of dimensions of the first tensor.
[0053] In a possible implementation manner, the first tensor includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
[0054] For example, in the scenario of using a deep neural network to perform face recognition on a target object, the first tensor can be image feature data, and the image feature data of the target object (such as a face feature map) can be stored in the global memory; in the scenario of using a deep neural network to perform voice recognition on a target object, the first tensor can be voice feature data, and the voice feature data of the target object can be stored in the global memory; in the scenario of using a deep neural network to perform character recognition on a target document, the first tensor can be text feature data, and the text feature data of the target document can be stored in the global memory; the embodiments of the present application do not limit the type of the first tensor.
[0055] In a possible implementation manner, the tensor information of the first tensor may include the data structure information of the first tensor and the address information of the first tensor; wherein, the data structure information of the first tensor, for example, includes the dimensions of the first tensor, the size of the first tensor, the data type of the elements in the first tensor (such as integer type, single-precision floating-point type, double-precision floating-point type, character type, etc.) and other information for describing the first tensor; the address information of the first tensor, for example, includes the base address (Base Address) of the first tensor in the global memory, the addressing space, and other address-related information.
[0056] Among them, since the content of the tensor information of the first tensor is relatively large, the tensor information of the first tensor can be stored in the global memory.
[0057] In a possible implementation, when the second tensor to be accessed is the block data of the first tensor, the access information includes: the position of the second tensor in the first tensor, and the size of the second tensor. Step S11 may include: determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.
[0058] Exemplarily, Figure 3 FIG. shows a schematic diagram of determining the address information of a second tensor according to an embodiment of the present disclosure. As Figure 3 shown, the second tensor block to be accessed is the block data of the first tensor tensor, and the second tensor block can be the data in any local area of the first tensor tensor.
[0059] The first tensor tensor is a two-dimensional tensor. The tensor information of the first tensor tensor includes: the base address tensor_base_address for determining the storage position of the first tensor tensor in the global memory, the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, and the size tensor_dim[1] of the first tensor tensor in the vertical dimension. Among them, since the first tensor tensor is address-continuous in the lowest dimension (for example, the horizontal dimension or the vertical dimension), the tensor storage engine can access each element in the first tensor tensor according to the base address tensor_base_address of the first tensor tensor and the position of each element in the first tensor tensor.
[0060] As Figure 3 shown, the access information of the second tensor block includes: the position block_pos of the second tensor block in the first tensor tensor, the size block_dim[0] of the second tensor block in the horizontal dimension, and the size block_dim[1] of the second tensor block in the vertical dimension.
[0061] The tensor storage engine can determine the base address block_base_address of the second tensor block according to the base address tensor_base_address in the tensor information of the first tensor tensor and the position block_pos of the second tensor block in the first tensor tensor.
[0062] The tensor storage engine can determine that the size of the addressing space of the second tensor block is block_dim[0] × block_dim[1] according to the size block_dim[0] of the second tensor block in the horizontal dimension and the size block_dim[1] of the second tensor block in the vertical dimension.
[0063] The tensor storage engine can determine the address information of the second tensor block to be accessed in the global memory according to the base address block_base_address of the second tensor block and the size block_dim[0] × block_dim[1] of the addressing space of the second tensor block.
[0064] In this way, the tensor storage engine can quickly and accurately determine the address information of the second tensor (for example, any block data in the first tensor) from the address of the first tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed.
[0065] In the case where the second tensor to be accessed is the block data at the edge of the first tensor, there may be a part of the data of the second tensor that is not within the first tensor. In one possible implementation, the method further includes: first, performing a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed to obtain a check result of the second tensor, where the check result is used to determine whether the second tensor is within the first tensor; in the case where the check result indicates that the first part of the second tensor is within the first tensor and the second part of the second tensor is outside the first tensor, determining the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed.
[0066] Among them, the tensor information of the first tensor further includes a padding flag (such as the number 0, the number 1, etc.). The tensor storage engine can pad the second part of the second tensor according to the padding flag to obtain the second part of the second tensor, so that in the subsequent step S12, in addition to reading the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor; the second part of the second tensor will also be synchronously written into the local memory to make the local memory store the complete second tensor.
[0067] Figure 4 A schematic diagram showing a method for determining the address information of the first part of the second tensor according to an embodiment of the present disclosure. As Figure 4As shown, the tensor information of the first tensor tensor includes: the base address tensor_base_address for determining the storage location of the first tensor tensor in the global memory, the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, and the size tensor_dim[1] of the first tensor tensor in the vertical dimension. The access information of the second tensor block includes: the position block_pos of the second tensor block in the first tensor tensor, the size block_dim[0] of the second tensor block in the horizontal dimension, and the size block_dim[1] of the second tensor block in the vertical dimension. Among them, the position block_pos of the second tensor block in the first tensor tensor includes: the position block_pos[0] of the second tensor block in the horizontal dimension of the first tensor tensor, and the position block_pos[1] of the second tensor block in the vertical dimension of the first tensor tensor.
[0068] First, based on the tensor information of the first tensor tensor and the access information of the second tensor block to be accessed, the inspection result of the second tensor block can be obtained, and this inspection result is used to determine whether the second tensor block is within the first tensor tensor.
[0069] If the sum of the position block_pos[0] of the second tensor block in the horizontal dimension of the first tensor tensor and the size block_dim[0] of the second tensor block in the horizontal dimension is less than or equal to the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, that is: (block_pos[0]+block_dim[0])≤tensor_dim[0], and the sum of the position block_pos[1] of the second tensor block in the vertical dimension of the first tensor tensor and the size block_dim[1] of the second tensor block in the vertical dimension is less than or equal to the size tensor_dim[1] of the first tensor tensor in the vertical dimension, that is: (block_pos[1]+block_dim[1])≤tensor_dim[1], it can be determined that the second tensor block is within the first tensor tensor. The address information of the second tensor block can be determined by referring to the method shown above, which will not be elaborated here. Figure 3 The address information of the second tensor block can be determined by referring to the method shown above, which will not be elaborated here.
[0070] If the sum of the position block_pos[0] of the second tensor block in the horizontal dimension of the first tensor tensor and the size block_dim[0] of the second tensor block in the horizontal dimension is greater than the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, that is: (block_pos[0] + block_dim[0]) > tensor_dim[0], and / or, the sum of the position block_pos[1] of the second tensor block in the vertical dimension of the first tensor tensor and the size block_dim[1] of the second tensor block in the vertical dimension is greater than the size tensor_dim[1] of the first tensor tensor in the vertical dimension, that is: (block_pos[1] + block_dim[1]) > tensor_dim[1], it can be determined that there is a part of the data of the second tensor block outside the first tensor tensor. As Figure 4 shown, the first part of the second tensor block to be accessed is within the first tensor tensor, and the second part of the second tensor block is outside the first tensor tensor.
[0071] In the case where the first part of the second tensor block is within the first tensor tensor and the second part of the second tensor block is outside the first tensor tensor, the tensor storage engine can determine the address information of the first part of the second tensor block according to the tensor information of the first tensor tensor and the access information of the second tensor block to be accessed.
[0072] In the example, the tensor storage engine can determine the base address block_base_address of the second tensor block, that is, the base address block_base_address of the first part of the second tensor block, according to the base address tensor_base_address in the tensor information of the first tensor tensor and the position block_pos of the second tensor block in the first tensor tensor.
[0073] Optionally, when (block_pos[0] + block_dim[0]) > tensor_dim[0] and (block_pos[1] + block_dim[1]) > tensor_dim[1], the tensor storage engine can determine the size of the addressing space of the first part of the second tensor block according to the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, the size tensor_dim[1] of the first tensor tensor in the vertical dimension, the position block_pos[0] of the second tensor block in the horizontal dimension of the first tensor tensor, and the position block_pos[1] of the second tensor block in the vertical dimension of the first tensor tensor, that is: (tensor_dim[0] - block_pos[0]) × (tensor_dim[1] - block_pos[1]).
[0074] The tensor storage engine can determine the address information in the global memory of the first part of the second tensor block to be accessed according to the base address block_base_address of the second tensor block and the size (tensor_dim[0] - block_pos[0]) × (tensor_dim[1] - block_pos[1]) of the addressing space of the first part of the second tensor block.
[0075] Optionally, when (block_pos[0] + block_dim[0]) > tensor_dim[0] and (block_pos[1] + block_dim[1]) ≤ tensor_dim[1] (see Figure 4 ), the tensor storage engine can determine the size of the addressing space of the first part of the second tensor block according to the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, the position block_pos[0] of the second tensor block in the horizontal dimension of the first tensor tensor, and the size block_dim[1] of the second tensor block in the vertical dimension, that is: (tensor_dim[0] - block_pos[0]) × block_dim[1].
[0076] The tensor storage engine can determine the address information in the global memory of the first part of the second tensor block to be accessed based on the base address block_base_address of the second tensor block and the size of the addressing space of the first part of the second tensor block, which is (tensor_dim[0] - block_pos[0]) × block_dim[1].
[0077] Optionally, when (block_pos[0] + block_dim[0]) ≤ tensor_dim[0] and (block_pos[1] + block_dim[1]) > tensor_dim[1], the tensor storage engine can determine the size of the addressing space of the first part of the second tensor block according to the size tensor_dim[1] of the first tensor tensor in the vertical dimension, the position block_pos[1] of the second tensor block in the vertical dimension of the first tensor tensor, and the size block_dim[0] of the second tensor block in the horizontal dimension, that is: block_dim[0] × (tensor_dim[1] - block_pos[1]).
[0078] The tensor storage engine can determine the address information in the global memory of the first part of the second tensor block to be accessed based on the base address block_base_address of the second tensor block and the size of the addressing space of the first part of the second tensor block, which is block_dim[0] × (tensor_dim[1] - block_pos[1]).
[0079] Among them, the tensor information of the first tensor tensor also includes a padding flag (such as the number 0, the number 1, etc.). The tensor storage engine can pad the second part of the second tensor block according to the padding flag to obtain the second part of the second tensor block, so that in the subsequent step S12, in addition to reading the first part of the second tensor block from the global memory storing the first tensor tensor to the local memory according to the address information of the first part of the second tensor block; the padded second part of the second tensor block will also be synchronously written into the local memory to make the local memory store the complete second tensor block.
[0080] In this way, the address information of the second tensor can be determined from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor. The second tensor can be a block of data obtained by arbitrarily cropping the first tensor, which is conducive to the subsequent tensor storage engine accessing any block of data in the first tensor that occupies a large space from the global memory through a single instruction, improving the applicability of the data processing method according to the embodiments of the present disclosure. Moreover, the tensor storage engine can also implement automatic address calculation and boundary checking for the block data, and at the same time, it can implement the filling process with any constant value for out-of-bounds access.
[0081] In a possible implementation manner, the convolution kernel can traverse the entire first tensor in a sliding manner on the first tensor. To facilitate the convolution operation and / or deconvolution operation of the first tensor and further improve the applicability of the data processing method according to the embodiments of the present disclosure, when the second tensor to be accessed is the covering data corresponding to any one or more elements in the convolution kernel in the first tensor, the access information includes: the expanded matrix coordinates and the convolution kernel element coordinates. The expanded matrix is a matrix formed by arranging, in rows or columns, the column vectors or row vectors corresponding to the covering data of the convolution kernel receptive field in the first tensor according to the description information in the first tensor. The description information includes convolution description information and deconvolution description information. Step S11 may include: obtaining the mapping relationship between the expanded matrix and the first tensor; determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the expanded matrix coordinates, and the convolution kernel element coordinates.
[0082] Among them, the tensor information of the first tensor may include the data structure information of the first tensor (such as including the dimensions of the first tensor, the size of the first tensor, the data type of the elements in the first tensor, etc.) and the address information of the first tensor (such as including the base address of the first tensor in the global memory, the addressing space, etc.). The mapping relationship is used to determine the address information of each element in the expanded matrix in the global storage space; the expanded matrix coordinates and the convolution kernel element coordinates are used to locate the second tensor in the expanded matrix.
[0083] Among them, the description information is used to provide an unfolding method for unfolding the original matrix into an unfolding matrix suitable for the convolution kernel. The description information includes the convolution kernel size, the stride of the convolution kernel, the padding number, the dilation rate, etc. Among them, if the description information is deconvolution description information and the stride in the deconvolution description information is greater than 1, when converting the covered data corresponding to the receptive field of the convolution kernel into corresponding column vectors or row vectors and arranging them into an unfolding matrix row by row or column by column, a constant value filling operation can be performed between each element in the covered data according to the padding identifier (such as the number 0, the number 1, etc.) included in the tensor information of the first tensor.
[0084] Figure 5 A schematic diagram showing the mapping relationship between the first tensor and the unfolding matrix according to an embodiment of the present disclosure. As Figure 5 shown, the tensor storage engine can know from the tensor information of the obtained first tensor that the first tensor is a four-dimensional tensor N×Cx×H×W (see Figure 5 the two cuboids on the upper side, whose size is 2×2×5×4), where N represents the dimension where the batch is located, Cx represents the dimension where the channel is located, H represents the dimension where the height is located, and W represents the dimension where the width is located.
[0085] As Figure 5 shown, the digital labels 0 to 79 in the first tensor are used to distinguish the elements of the first tensor, and each element occupies 128 bytes (hexadecimal representation: 80HB) of storage space in the global memory. Thus, assuming that the base address of the first tensor in the global memory is address, the address information of element 0 of the first tensor in the global memory is address to (address + 007FH), the address information of element 1 of the first tensor in the global memory is (address + 0080H) to (address + 00FFH), the address information of element 2 of the first tensor in the global memory is (address + 0100H) to (address + 017FH), and so on. The address information of element 79 of the first tensor in the global memory is (address + 2780H) to (address + 27FFH).
[0086] Figure 5 The two-dimensional matrix shown in the lower part is the unfolding matrix of the first tensor. This unfolding matrix is obtained by sequentially converting the covered data corresponding to the receptive field of the convolution kernel into corresponding row vectors according to the convolution description information (for example, the convolution kernel size is 3×3, the stride of the convolution kernel is 1, and the padding number is 0) in the first tensor and arranging them into a two-dimensional matrix row by row or column by column.
[0087] As shown Figure 5 in the figure, assume that P represents the position where the convolutional kernel slides in the height direction, Q represents the position where the convolutional kernel slides in the width direction, Cx represents the position where the convolutional kernel slides in the channel direction, and N represents whether the convolutional kernel slides in the N1×Cx×H×W part of the first tensor or in the N2×Cx×H×W part of the first tensor.
[0088] As shown Figure 5 in the gray part, for the first time, the convolutional kernel slides to the Q0 position in the width direction, to the P0 position in the height direction, and to the Cx = 0 position in the width direction. The covered data corresponding to the receptive field of the convolutional kernel in the N1×Cx×H×W part of the first tensor is which can be converted to [0 1 2 4 5 6 8 9 10], and the covered data corresponding to the receptive field of the convolutional kernel in the N2×Cx×H×W part of the first tensor is which can be converted to [40 41 42 44 45 46 48 49 50].
[0089] For the second time, the convolutional kernel slides to the Q1 position in the width direction, remains at the P0 position in the height direction, and slides to the Cx = 0 position in the width direction. The covered data corresponding to the receptive field of the convolutional kernel in the N1×Cx×H×W part of the first tensor is which can be converted to [1 2 3 5 6 7 9 10 11], and the covered data corresponding to the receptive field of the convolutional kernel in the N2×Cx×H×W part of the first tensor is which can be converted to [41 42 43 45 46 47 49 50 51].
[0090] For the third time, the convolutional kernel slides to the Q0 position in the width direction, remains at the P1 position in the height direction, and slides to the Cx = 0 position in the width direction. The covered data corresponding to the receptive field of the convolutional kernel in the N1×Cx×H×W part of the first tensor is which can be converted to [4 5 6 8 9 10 12 13 14], and the covered data corresponding to the receptive field of the convolutional kernel in the N2×Cx×H×W part of the first tensor is which can be converted to [44 45 46 48 49 50 52 53 54].
[0091] And so on, until the 12th time, the convolutional kernel slides to the Q1 position in the width direction, to the P2 position in the height direction, and to the Cx = 1 position in the width direction. The covered data corresponding to the receptive field of the convolutional kernel in the N1×Cx×H×W part of the first tensor is It can be converted into [29 30 31 33 34 35 37 38 39], and the covered data corresponding to the receptive field of the convolution kernel in the N2×Cx×H×W part of the first tensor is It can be converted into [697071 73 74 75 77 78 79].
[0092] After converting the row vectors of the covered data for these 12 times, arranging them row by row or column by column results in an unfolded matrix as shown in Figure 5 the unfolded matrix shown. The mapping relationship between the unfolded matrix and the first tensor can be obtained according to the corresponding relationship between the elements with the same digital labels in the first tensor and the unfolded matrix. In the example, the tensor storage engine can directly obtain the mapping relationship between the user-input unfolded matrix and the first tensor; or, the tensor storage engine can first perform the unfolding process on the first tensor read from the global memory as described above according to the tensor information and description information of the first tensor obtained, determine the unfolded matrix, and then obtain the mapping relationship between the unfolded matrix and the first tensor according to the corresponding relationship between the first tensor and the unfolded matrix.
[0093] It should be understood that Figure 5 what is shown is the case where the description information is convolution description information. If the description information is deconvolution description information, reference can be made to the above introduction, and in the case where the stride of the convolution kernel in the deconvolution description information is greater than 1, a constant value filling operation is performed between each element in the unfolded matrix, which will not be elaborated here.
[0094] In a possible implementation manner, determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates includes: determining the second tensor to be accessed from the unfolded matrix according to the unfolded matrix coordinates and the convolution kernel element coordinates; determining multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolded matrix and the first tensor; and determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.
[0095] In the unfolded matrix shown in Figure 5 NPQ can be used as the coordinate of the unfolded matrix in the vertical dimension to locate a certain row in the unfolded matrix. For example, N1P0Q0 represents the first row, N1P0Q1 represents the second row, N1P1Q0 represents the third row, N1P1Q1 represents the fourth row, N1P2Q0 represents the fifth row, N1P2Q1 represents the sixth row, N2P0Q0 represents the seventh row, N2P0Q1 represents the eighth row, N2P1Q0 represents the ninth row, N2P1Q1 represents the tenth row, N2P2Q0 represents the eleventh row, and N2P2Q1 represents the twelfth row.
[0096] Cx can be used as the coordinate of the unfolded matrix in the horizontal dimension. Cx = 0 is used to locate the leftmost 8 columns of the unfolded matrix, and Cx = 1 is used to locate the rightmost 8 columns of the unfolded matrix.
[0097] RS is the coordinate of the convolutional kernel element. R represents the coordinate of the vertical dimension of the convolutional kernel, used to locate a certain row of the convolutional kernel, and S represents the coordinate of the horizontal dimension of the convolutional kernel, used to locate a certain column of the convolutional kernel. For example, R = 0, S = 0 represents the convolutional kernel element in the first column of the first row, R = 0, S = 1 represents the convolutional kernel element in the second column of the first row, R = 0, S = 2 represents the convolutional kernel element in the third column of the first row, R = 1, S = 0 represents the convolutional kernel element in the first column of the second row, R = 1, S = 1 represents the convolutional kernel element in the second column of the first row, R = 1, S = 2 represents the convolutional kernel element in the third column of the second row, R = 2, S = 0 represents the convolutional kernel element in the first column of the third row, R = 2, S = 1 represents the convolutional kernel element in the second column of the third row, and R = 2, S = 2 represents the convolutional kernel element in the third column of the third row.
[0098] According to the coordinate Cx of the unfolded matrix in the horizontal dimension and the convolutional kernel element coordinate RS, any one or more columns of the second tensor can be located in the unfolded matrix, that is, the covered data corresponding to any one or more elements in the convolutional kernel in the first tensor.
[0099] For example, Cx = 0, R = 0, S = 0 represents the first column in the unfolded matrix, that is, the covered data corresponding to the element in the first column of the first row of the convolutional kernel after sliding multiple times (e.g., 12 times) in the first tensor; Cx = 0, R = 0, S = 1 represents the second column in the unfolded matrix, that is, the covered data corresponding to the element in the second column of the first row of the convolutional kernel after sliding multiple times (e.g., 12 times) in the first tensor; and so on. Cx = 1, R = 2, S = 2 represents the eighteenth column (the last column) in the unfolded matrix, that is, the covered data corresponding to the element in the third column of the third row of the convolutional kernel after sliding multiple times (e.g., 12 times) in the first tensor.
[0100] After locating any one or more columns of the second tensor in the unfolded matrix, according to the mapping relationship between the unfolded matrix and the first tensor, multiple elements corresponding to the second tensor in the first tensor can be determined, and the address information of multiple elements corresponding to the second tensor in the first tensor can be determined as the address information of the second tensor.
[0101] For example, assume that the second tensor is the first column in the unfolded matrix: [0 1 4 5 8 9 40 41 44 45 48 49]. The covered data of the element in the first row and first column of the corresponding convolution kernel in the first tensor can be obtained from the first tensor according to the mapping relationship between the unfolded matrix and the first tensor. The elements in the first tensor with the same element numbers as those in the first column of the unfolded matrix are retrieved, and the addresses of the elements [0 1 4 5 8 9 40 41 44 45 48 49] in the first tensor are determined as the address information of the second tensor.
[0102] Another example, assume that the second tensor is the second column in the unfolded matrix: [1 2 5 6 9 10 41 42 45 46 49 50]. The covered data of the element in the first row and second column of the corresponding convolution kernel in the first tensor can be obtained from the first tensor according to the mapping relationship between the unfolded matrix and the first tensor. The elements in the first tensor with the same element numbers as those in the second column of the unfolded matrix are retrieved, and the addresses of the elements [1 2 5 6 9 10 41 42 45 46 49 50] in the first tensor are determined as the address information of the second tensor.
[0103] In practical applications, the above process can be encapsulated as a TensorMemoryEngine instruction, which can be used to access one or more columns in the unfolded matrix (corresponding to the elements of the first tensor covered by one or more convolution kernel elements). This is beneficial for the tensor storage engine to access the data covered by any one or more convolution kernel elements in the first tensor, which occupies a large space in the global memory, through a single instruction. The pseudo-code is as follows:
[0104]
[0105] In this way, the tensor storage engine does not need to perform the process of unfolding the first tensor into an unfolded matrix, and even less need to occupy a large amount of storage space to record the unfolded matrix of the first tensor. In the scenario where the unfolded matrix is not obtained, the tensor storage engine can calculate the coordinates of any element in the unfolded matrix in the first tensor (such as the coordinates of the elements of the first tensor covered by one or more convolution kernel elements) according to the unfolded matrix coordinates NPQCx and the convolution kernel element coordinates RS. This is beneficial for using a small amount of hardware resources to simplify the software implementation logic (such as software implementation of convolution and deconvolution operations).
[0106] In this way, the tensor storage engine can quickly and accurately determine the address information of the second tensor (such as the covered data corresponding to any one or more elements in the convolution kernel in the first tensor) from the address of the first tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed (such as unfolded matrix coordinates, convolution kernel element coordinates).
[0107] In the data processing method of the embodiments of the present disclosure, the second tensor may be the block data of the first tensor, or the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, or the data formed by filling constant values between the coverage data and each element in the coverage data, and it supports the block loading of the first tensor, the block loading of the expanded data applicable to convolution operations on the first tensor, and the block loading of the expanded data applicable to deconvolution operations on the first tensor, unifying multiple loading modes and being more user-friendly for programming.
[0108] In a possible implementation manner, the tensor information of the first tensor is stored in the global memory, and the access information of the second tensor to be accessed is stored in the register.
[0109] Since the content of the tensor information of the first tensor is large and the tensor information of the first tensor does not change easily, the tensor information of the first tensor can be stored in the global memory. Since the second tensor to be accessed is the block data of any area in the first tensor, or the second tensor to be accessed is the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, and the access information of the second tensor is changeable, the access information of the second tensor can be stored in the register for easy change during program operation.
[0110] This way of storing the tensor information of the unchanging first tensor in the global memory and storing the access information of the changing second tensor to be accessed in the register places various information in different places according to whether the information is likely to change, which is beneficial to improving the flexibility and processing efficiency of data processing.
[0111] Compared with the synchronous memory access method in the related art, it is necessary to load data into the register file. For example, for the single-instruction multiple-thread (SIMT) architecture of a general-purpose computing on graphics processing units (GPGPU), the registers of each thread are isolated from each other and cannot be shared, and it is also necessary to move the data from the register to the shared storage space. However, in the embodiments of the present disclosure, since the tensor storage engine is connected to the register storing the access information of the second tensor, the tensor storage engine can access the register at any time, so that multiple threads corresponding to the tensor storage engine can share the register, and moreover, the processing method of the embodiments of the present disclosure can also adopt the register passing method to improve the flexibility of data processing.
[0112] The address information of the second tensor is determined in step S11. In step S12, the second tensor can be read from the global memory storing the first tensor to the local memory according to the address information of the second tensor. Among them, the local memory is closer to the tensor storage engine than the global memory, and the tensor storage engine can have higher efficiency for read and write operations on the local memory. The tensor storage engine can perform data migration processing on the first tensor stored in the global memory, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine.
[0113] In a possible implementation manner, step S12 may include: when the size of the second tensor is greater than the read bandwidth, according to the address information of the second tensor, split the first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests, and each second memory access request is used to read the split data of different parts of the second tensor from the global memory to the local memory; according to the multiple second memory access requests, serially write the multiple split data constituting the second tensor in the global memory into the local memory. Among them, the read bandwidth refers to the ability to read data from the memory per unit time, and the read bandwidth of the memory can be expressed by the number of bits (bits) or the number of bytes (Bytes) per transmission.
[0114] Exemplarily, assume that the read bandwidth of the global memory is 128 bytes, and the first memory access request is to access a 1024-bit second tensor, whose corresponding address is 0. However, the read bandwidth supports a maximum 128-bit memory access request. The first memory access request can be split into eight second memory access requests, namely: second memory access request 0 to second memory access request 7. Among them, the second memory access request 0 is used to read 128-bit split data 0 of the second tensor from address 0 in the global memory to the local memory; the second memory access request 0 is used to read 128-bit split data 0 belonging to the second tensor from address 0 in the global memory to the local memory; the second memory access request 1 is used to read 128-bit split data 1 belonging to the second tensor from address 128 in the global memory to the local memory; the second memory access request 2 is used to read 128-bit split data 2 belonging to the second tensor from address 256 in the global memory to the local memory; the second memory access request 3 is used to read 128-bit split data 3 belonging to the second tensor from address 384 in the global memory to the local memory; the second memory access request 4 is used to read 128-bit split data 4 belonging to the second tensor from address 512 in the global memory to the local memory; the second memory access request 5 is used to read 128-bit split data 5 belonging to the second tensor from address 640 in the global memory to the local memory; the second memory access request 6 is used to read 128-bit split data 6 belonging to the second tensor from address 768 in the global memory to the local memory; the second memory access request 7 is used to read 128-bit split data 7 belonging to the second tensor from address 896 in the global memory to the local memory. Then, according to the second memory access request 0 to the second memory access request 7, the split data 0 to split data 7 that constitute the second tensor in the global memory can be serially written into the local memory. It should be understood that the present disclosure does not specifically limit the size of the second tensor and the size of the read bandwidth, and can be set according to the actual application scenario.
[0115] In this way, if the size of the second tensor in the global memory exceeds the maximum granularity supported by the storage system, it can be split for serial processing, improving the applicability of the data processing method in the embodiments of the present disclosure. Moreover, in the embodiments of the present disclosure, the read bandwidth of the global memory is less than the read bandwidth of the local memory. When reading data (such as the second tensor) from the global memory to the local memory, the local memory can be reused multiple times, reducing the bandwidth pressure on the global memory.
[0116] In step S12, the second tensor is read from the global memory storing the first tensor to the local memory. In step S13, in response to the local memory receiving the second tensor, an interrupt is triggered. In step S12, the tensor storage engine transfers the second tensor from the global memory to the local memory, and the time consumed may be relatively long. Before waiting for the local memory to receive the second tensor and trigger an interrupt, the tensor storage engine can asynchronously execute other instructions to further improve the processing efficiency.
[0117] In this way, by setting up an interrupt wake-up mechanism, it is not necessary to follow the pipeline mode (for example, the next instruction can only be executed after the data is transferred to the destination). Instead, after the current instruction is sent out, other instructions can be executed, and wait for the current instruction to complete and send a signal to the processor core to trigger an interrupt to wake up the current instruction, which is more efficient.
[0118] The data processing method of the embodiments of the present disclosure can quickly read the second tensor from the global memory storing the first tensor to the local memory according to the tensor information of the first tensor and the access information of the second tensor to be accessed, improving the efficiency and flexibility of tensor storage. Moreover, every time the local memory receives the second tensor, an interrupt can be triggered to realize the asynchrony of the tensor reading process.
[0119] Among them, the second tensor of the embodiments of the present disclosure can be the block data of the first tensor, or the covered data corresponding to any one or more elements in the convolution kernel in the first tensor, or the data formed by filling constant values between the covered data and each element in the covered data, supporting the block loading of the first tensor, the block loading of the expanded data suitable for convolution operation on the first tensor, and the block loading of the expanded data suitable for deconvolution operation on the first tensor, unifying multiple loading modes, which is more user-friendly for programming.
[0120] Among them, the processing method of the embodiments of the present disclosure can store the tensor information of the unchanged first tensor in the global memory and store the access information of the variable second tensor to be accessed in the register. By placing various information in different places according to whether the information is likely to change, it is beneficial to improve the flexibility and processing efficiency of data processing.
[0121] It can be understood that the above-mentioned method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0122] In addition, the present disclosure also provides a data processing device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any one of the data processing methods provided by the present disclosure. For the corresponding technical solutions and descriptions, please refer to the corresponding records in the method section, which will not be elaborated herein.
[0123] Figure 6 The block diagram of a data processing device according to an embodiment of the present disclosure is shown, as Figure 6 shown, the device includes:
[0124] A determination module 61, configured to determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed;
[0125] A reading module 62, configured to read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor;
[0126] A triggering module 63, configured to trigger an interrupt in response to the local memory receiving the second tensor.
[0127] In a possible implementation manner, the determination module 61 is configured to: when the second tensor to be accessed is the block data of the first tensor, and the access information includes the position of the second tensor in the first tensor and the size of the second tensor, determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.
[0128] In a possible implementation manner, the determination module 61 is further configured to: perform a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, and obtain a check result of the second tensor; when the check result indicates that the first part of the second tensor is within the first tensor and the second part of the second tensor is outside the first tensor, determine the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; the reading module 62 is configured to: when the tensor information of the first tensor includes a padding flag, read the first part of the second tensor from the global memory storing the first tensor to the local memory, perform padding on the second part of the second tensor according to the padding flag to obtain the second part of the second tensor; write the second part of the second tensor into the local memory.
[0129] In a possible implementation, the determining module 61 is configured to: when the second tensor to be accessed is the covering data corresponding to any one or more elements in the convolution kernel in the first tensor, and the access information includes the coordinates of the unfolded matrix and the coordinates of the convolution kernel elements, obtain the mapping relationship between the unfolded matrix and the first tensor; determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the coordinates of the unfolded matrix, and the coordinates of the convolution kernel elements; wherein, the unfolded matrix is a matrix formed by arranging, row by row or column by column, the covering data corresponding to the receptive field of the convolution kernel into corresponding column vectors or row vectors in the first tensor according to the description information in sequence, and the description information includes convolution description information and deconvolution description information.
[0130] In a possible implementation, determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the coordinates of the unfolded matrix, and the coordinates of the convolution kernel elements includes: determining the second tensor to be accessed from the unfolded matrix according to the coordinates of the unfolded matrix and the coordinates of the convolution kernel elements; determining multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolded matrix and the first tensor; and determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.
[0131] In a possible implementation, the reading module 62 is configured to: when the size of the second tensor is greater than the reading bandwidth, split a first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests according to the address information of the second tensor, where each second memory access request is used to read split data of different parts of the second tensor from the global memory to the local memory; and serially write the multiple split data constituting the second tensor in the global memory into the local memory according to the multiple second memory access requests.
[0132] In a possible implementation, the tensor information of the first tensor is stored in the global memory, and the access information of the second tensor to be accessed is stored in the register.
[0133] In a possible implementation, the first tensor includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
[0134] This method has a specific technical association with the internal structure of a computer system and can solve technical problems such as how to improve the computing efficiency or execution effect of hardware (including reducing the amount of data storage, reducing the amount of data transmission, and increasing the hardware processing speed), thereby obtaining a technical effect of improving the internal performance of the computer system in line with natural laws.
[0135] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0136] The embodiments of the present disclosure also propose a computer-readable storage medium storing computer program instructions, and when the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0137] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.
[0138] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.
[0139] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0140] The embodiments of the present disclosure also provide an electronic device, which can be provided as a terminal, a server, or other forms of devices. For example, a User Equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a Personal Digital Assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.
[0141] Figure 7 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Refer to Figure 7, the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0142] The electronic device 1900 may further include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.
[0143] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0144] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0145] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0146] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0147] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computer (e.g., via the Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0148] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0149] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture comprising instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0150] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0151] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0152] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0153] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses can be referred to each other. For the sake of brevity, they are not elaborated herein.
[0154] Those skilled in the art can understand that in the above methods of the specific implementation manners, the writing order of the steps does not mean a strict execution order that constitutes any limitation on the implementation process. The specific execution order of the steps should be determined according to their functions and possible internal logics.
[0155] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A data processing method, characterized in that, Including: Determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed; Read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor; Trigger an interrupt in response to the local memory receiving the second tensor.
2. The method according to claim 1, characterized in that When the second tensor to be accessed is the block data of the first tensor, the access information includes: the position of the second tensor in the first tensor and the size of the second tensor. Determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, including: Determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.
3. The method according to claim 2, wherein The method further includes: Perform a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, and obtain the check result of the second tensor; When the check result indicates that the first part of the second tensor is within the first tensor and the second part of the second tensor is outside the first tensor, determine the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; The tensor information of the first tensor includes a padding flag. Reading the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor includes: Read the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor; Pad the second part of the second tensor according to the padding flag to obtain the second part of the second tensor; Write the second part of the second tensor to the local memory.
4. The method according to claim 1, wherein When the second tensor to be accessed is the covering data corresponding to any one or more elements in the convolution kernel in the first tensor, the access information includes: the unfolded matrix coordinates and the convolution kernel element coordinates. The unfolded matrix is a matrix formed by sequentially converting the covering data corresponding to the convolution kernel receptive field into corresponding column vectors or row vectors in the first tensor and arranging them row by row or column by column. The description information includes convolution description information and deconvolution description information. Determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, including: Obtain the mapping relationship between the unfolded matrix and the first tensor; Determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates.
5. The method according to claim 4, characterized in that Determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates, including: Determine the second tensor to be accessed from the expanded matrix according to the expanded matrix coordinates and the convolutional kernel element coordinates; Determine multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the expanded matrix and the first tensor; Determine the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.
6. The method according to any one of claims 1-5, characterized in that, Read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor, including: When the size of the second tensor is greater than the read bandwidth, split the first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests according to the address information of the second tensor, and each second memory access request is used to read split data of different parts of the second tensor from the global memory to the local memory; Write the multiple split data constituting the second tensor in the global memory to the local memory serially according to the multiple second memory access requests.
7. The method according to any one of claims 1-5, characterized in that, The tensor information of the first tensor is stored in the global memory, and the access information of the second tensor to be accessed is stored in the register.
8. The method according to any one of claims 1-5, characterized in that, The first tensor includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
9. A data processing device, characterized in that, including: A determination module, configured to determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed; A reading module, configured to read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor; A triggering module, configured to trigger an interrupt in response to the local memory receiving the second tensor.
10. An electronic device, characterized in that, including: A memory and a processor; Computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 8 is executed.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, including computer-readable code, or a computer-readable storage medium carrying computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Accessing prologue and epilogue data
CN109324827A
Tensor processor
CN110033085A
Memory devices and methods which may facilitate tensor memory access
CN112470133A
Data stream cache of neural network tensor processor
CN112860596A
Stride slice operator processing method and device based on a mercuric chloride AI processor
CN113722269A