Data processing method and apparatus, electronic device, and storage medium

By determining the address information of the tensor and reading it asynchronously to local memory, the problem of low access efficiency of tensor storage space is solved, and the data processing efficiency and flexibility of the artificial intelligence chip are improved.

WO2025139521A1PCT designated stage expired Publication Date: 2025-07-03MOORE THREADS TECH CO LTD

Patent Information

Application Number
PCT/CN2024/134081
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the access efficiency of tensor storage space is low and the flexibility is insufficient, which affects the data flow design and performance improvement of artificial intelligence chips.

Method used

By determining the address information of the second tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed, the address information of the second tensor to be accessed, and reading it from the global memory into the local memory, an interrupt is triggered to implement the asynchronous reading process.

Benefits of technology

It improves the efficiency and flexibility of tensor storage, realizes fast and asynchronous data access, and is suitable for a variety of data processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024134081_03072025_PF_FP_ABST
    Figure CN2024134081_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data processing method and apparatus, an electronic device, and a storage medium. The method comprises: on the basis of tensor information of a first tensor and access information of a second tensor to be accessed, determining address information of said second tensor, and on the basis of the address information of said second tensor, reading said second tensor from a global memory storing the first tensor into a local memory; and in response to the local memory receiving said second tensor, triggering an interrupt. The embodiments of the present disclosure can improve the efficiency and flexibility of tensor storage, and each time the local memory receives the second tensor, an interrupt can be triggered, achieving asynchronization of the tensor reading process.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method and device, electronic device and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311865153.8 and application name “Data processing method and device, electronic device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of computer technology, and in particular to a data processing method and device, an electronic device, and a storage medium. Background Art

[0003] As the application areas of artificial intelligence (AI) technology continue to expand, the performance requirements for AI chips are also increasing. Storage performance is one of the bottlenecks in AI chip performance. There is a desire to further improve the data flow design of AI chips, enabling them to access data more efficiently and flexibly. Summary of the Invention

[0004] The present disclosure proposes a data processing technology solution.

[0005] According to one aspect of the present disclosure, a data processing method is provided, including: determining address information of a first tensor based on tensor information of the second tensor and access information of a second tensor to be accessed; reading the second tensor from a global memory storing the first tensor to a local memory based on the address information of the second tensor; and triggering an interrupt in response to the local memory receiving the second tensor.

[0006] In one possible implementation, when the second tensor to be accessed is the block data of the first tensor, the access information includes: the position of the second tensor in the first tensor, and the size of the second tensor. According to the tensor information of the first tensor and the access information of the second tensor to be accessed, the address information of the second tensor is determined, including: according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor, the address information of the second tensor is determined from the address of the first tensor.

[0007] In one possible implementation, the method further includes: performing a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, to obtain a check result of the second tensor; when the check result indicates that the first part of the second tensor is located within the first tensor and the second part of the second tensor is located outside the first tensor, determining the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; the tensor information of the first tensor includes a fill identifier, and reading the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor, including: reading the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor, filling the second part of the second tensor according to the fill identifier to obtain the second part of the second tensor; and writing the second part of the second tensor into the local memory.

[0008] In one possible implementation, when the second tensor to be accessed is the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, the access information includes: expansion matrix coordinates, convolution kernel element coordinates, and the expansion matrix is ​​in the first tensor, and converts the coverage data corresponding to the convolution kernel receptive field into corresponding column vectors or row vectors in sequence according to the description information, and arranges them in rows or columns. The description information includes convolution description information and deconvolution description information. According to the tensor information of the first tensor and the access information of the second tensor to be accessed, the address information of the second tensor is determined, including: obtaining the mapping relationship between the expansion matrix and the first tensor; and determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the expansion matrix coordinates, and the convolution kernel element coordinates.

[0009] In a possible implementation, the address information of the second tensor is determined from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolding matrix coordinates, and the convolution kernel element coordinates, including: determining the second tensor to be accessed from the unfolding matrix according to the unfolding matrix coordinates and the convolution kernel element coordinates; determining the multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolding matrix and the first tensor; and determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.

[0010] In one possible implementation, according to the address information of the second tensor, the second tensor is read from the global memory storing the first tensor to the local memory, including: when the size of the second tensor is greater than the read bandwidth, according to the address information of the second tensor, the first memory access request for reading the second tensor from the global memory to the local memory is split into multiple second memory access requests, each second memory access request is used to read different parts of the split data of the second tensor from the global memory to the local memory; according to the multiple second memory access requests, the multiple split data constituting the second tensor in the global memory are serially written to the local memory.

[0011] In a possible implementation, the tensor information of the first tensor is stored in a global memory, and the access information of the second tensor to be accessed is stored in a register.

[0012] In one possible implementation, the first tensor includes feature data in a deep learning task, where the feature data includes at least one of image feature data, speech feature data, and text feature data.

[0013] According to one aspect of the present disclosure, a data processing device is provided, including: a determination module for determining address information of a first tensor based on tensor information of the second tensor and access information of a second tensor to be accessed; a reading module for reading the second tensor from a global memory storing the first tensor to a local memory based on the address information of the second tensor; and a triggering module for triggering an interrupt in response to the local memory receiving the second tensor.

[0014] In one possible implementation, the determination module is used to: when the second tensor to be accessed is block data of the first tensor, and the access information includes the position of the second tensor in the first tensor and the size of the second tensor, determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.

[0015] In one possible implementation, the determination module is further used to: perform a boundary check on the second tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed, to obtain a check result of the second tensor; when the check result indicates that the first part of the second tensor is located within the first tensor and the second part of the second tensor is located outside the first tensor, determine the address information of the first part of the second tensor based on the tensor information of the first tensor and the access information of the second tensor; the reading module is used to: when the tensor information of the first tensor includes a fill identifier, read the first part of the second tensor from the global memory storing the first tensor to the local memory based on the address information of the first part of the second tensor, fill the second part of the second tensor based on the fill identifier to obtain the second part of the second tensor; and write the second part of the second tensor into the local memory.

[0016] In one possible implementation, the determination module is used to: when the second tensor to be accessed is the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, and the access information includes the expansion matrix coordinates and the convolution kernel element coordinates, obtain the mapping relationship between the expansion matrix and the first tensor; determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the expansion matrix coordinates, and the convolution kernel element coordinates; wherein the expansion matrix is ​​in the first tensor, and the coverage data corresponding to the convolution kernel receptive field is converted into corresponding column vectors or row vectors in sequence according to the description information, and the matrix is ​​arranged in rows or columns, and the description information includes convolution description information and deconvolution description information.

[0017] In a possible implementation, the address information of the second tensor is determined from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolding matrix coordinates, and the convolution kernel element coordinates, including: determining the second tensor to be accessed from the unfolding matrix according to the unfolding matrix coordinates and the convolution kernel element coordinates; determining the multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolding matrix and the first tensor; and determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.

[0018] In one possible implementation, the reading module is used to: when the size of the second tensor is greater than the read bandwidth, split the first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests according to the address information of the second tensor, each second memory access request being used to read different parts of the split data in the second tensor from the global memory to the local memory; and according to the multiple second memory access requests, write the multiple split data constituting the second tensor in the global memory into the local memory serially.

[0019] In a possible implementation, the tensor information of the first tensor is stored in a global memory, and the access information of the second tensor to be accessed is stored in a register.

[0020] In one possible implementation, the first tensor includes feature data in a deep learning task, where the feature data includes at least one of image feature data, speech feature data, and text feature data.

[0021] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0022] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.

[0023] According to one aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes to implement the above method.

[0024] In an embodiment of the present disclosure, the address information of the second tensor is determined based on the tensor information of the first tensor and the access information of the second tensor to be accessed, and the second tensor is read from the global memory storing the first tensor to the local memory based on the address information of the second tensor; in response to the local memory receiving the second tensor, an interrupt is triggered. In this way, the second tensor can be quickly read from the global memory storing the first tensor to the local memory based on the tensor information of the first tensor and the access information of the second tensor to be accessed, thereby improving the efficiency and flexibility of tensor storage. In addition, each time the local memory receives the second tensor, an interrupt can be triggered to achieve asynchronous tensor reading.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0027] FIG1 shows a flow chart of a data processing method according to an embodiment of the present disclosure.

[0028] FIG2 shows a schematic diagram of a hardware architecture according to an embodiment of the present disclosure.

[0029] FIG3 shows a schematic diagram of determining address information of a second tensor according to an embodiment of the present disclosure.

[0030] FIG4 shows another schematic diagram of determining address information of a second tensor according to an embodiment of the present disclosure.

[0031] FIG5 is a schematic diagram showing a mapping relationship between a first tensor and an unfolding matrix according to an embodiment of the present disclosure.

[0032] FIG6 shows a block diagram of a data processing device according to an embodiment of the present disclosure.

[0033] FIG7 shows a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0035] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0036] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0037] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0038] With the widespread application of deep learning and neural networks, data is often stored in the form of tensors, which can generalize vectors and matrices to any dimension. In related technologies, processors (such as general-purpose computing on graphics processing units (GPGPU)) need to decompose the storage space where tensors are stored in order to access it. Each thread of the processor accesses a part of it, and in the process of each thread accessing the part of the storage space it is responsible for, a large number of address calculation instructions and boundary checking instructions need to be added. This method of accessing tensor storage space is less efficient. At the same time, since a single instruction accessing storage space is limited in size, accessing a tensor that occupies a large storage space requires the coordinated use of multiple instructions, which is not flexible enough during use and also affects the improvement of memory access performance.

[0039] In view of this, in order to improve the efficiency and flexibility of tensor storage, an embodiment of the present disclosure provides a data processing method. FIG1 shows a flowchart of the data processing method according to an embodiment of the present disclosure. As shown in FIG1 , the data processing method includes:

[0040] In step S11, address information of the second tensor is determined according to the tensor information of the first tensor and the access information of the second tensor to be accessed;

[0041] In step S12, according to the address information of the second tensor, the second tensor is read from the global memory storing the first tensor to the local memory;

[0042] In step S13 , in response to the local memory receiving the second tensor, an interrupt is triggered.

[0043] The data processing method of the embodiment of the present disclosure can quickly read the second tensor from the global memory storing the first tensor to the local memory based on the tensor information of the first tensor and the access information of the second tensor to be accessed, thereby improving the efficiency and flexibility of tensor storage. Moreover, each time the local memory receives the second tensor, an interrupt can be triggered to realize the asynchronous tensor reading process.

[0044] In one possible implementation, the data processing method of the embodiment of the present disclosure can be executed by a tensor storage engine, which is a hardware module in a processor chip for implementing data movement, wherein the processor chip can be a newly designed one or an improved one of an existing processor chip, and the types of processor chips may include but are not limited to: central processing unit (CPU), graphics processing unit (GPU), general-purpose computing on graphics processing units (GPGPU), digital signal processor (DSP), application specific integrated circuit (ASIC), tensor processing unit (TPU), field programmable gate array (FPGA) or other programmable logic devices, and may also include processors of microprocessors or other conventional processors.

[0045] In one possible implementation, the processor chip may include one or more tensor storage engines. The present disclosure does not limit the number of tensor storage engines in the processor chip.

[0046] In the example, for a single-core, single-threaded processor chip, a tensor storage engine can be set up in the processor chip to perform data movement processing on the first tensor stored in the global memory outside the processor chip, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine in the processor chip.

[0047] In the example, for a multi-core and multi-threaded processor chip, multiple tensor storage engines can be set in the processor chip. For example, assuming that the processor chip includes N (N≥2) computing cores (Core), each computing core can execute M (M≥2) threads in parallel, a tensor storage engine can be set for each computing core respectively, and the M threads in each computing core can share a tensor storage engine, and a local memory, the tensor storage engine is used to perform data movement processing on the first tensor stored in the global memory shared by multiple computing cores, and store the second tensor determined based on the first tensor in the local memory within the computing core. Alternatively, M tensor storage engines can be set for each computing core respectively, each thread in each computing core can correspond to a tensor storage engine, and a local memory, the tensor storage engine is used to perform data movement processing on the first tensor stored in the global memory shared by multiple computing cores, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine in the computing core.

[0048] In this example, for a multi-core single-threaded processor chip, multiple tensor storage engines can be set up in the processor chip. For example, assuming that the processor chip includes N (N≥2) single-threaded computing cores, a tensor storage engine can be set up for each computing core. The tensor storage engine is used to move data of the first tensor stored in the global memory shared by multiple computing cores, and store the second tensor determined based on the first tensor in the local memory within the computing core.

[0049] In this example, for a single-core multi-threaded processor chip, multiple tensor storage engines can be set up on the processor chip. For example, assuming that the processor chip can execute M (M≥2) threads in parallel, each thread can correspond to a tensor storage engine and a local memory. Each tensor storage engine can be used to move data of a first tensor stored in a global memory shared by multiple threads, and store a second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine.

[0050] In one possible implementation, Figure 2 shows a schematic diagram of the hardware architecture according to an embodiment of the present disclosure. As shown in Figure 2, the tensor storage engine is connected to the global memory and the local memory, respectively, and can be used to perform read operations and / or write operations on the global memory and the local memory. In addition, the tensor storage engine can also be used to receive tensor information of the first tensor and access information of the second tensor, so as to facilitate the transfer of the second tensor from the global memory storing the first tensor to the local memory, and trigger an interrupt in response to the local memory receiving the second tensor.

[0051] The following is an exemplary description of the data processing method implemented by the tensor storage engine according to the embodiment of the present disclosure.

[0052] In step S11 , address information of the second tensor to be accessed is determined according to the tensor information of the first tensor and the access information of the second tensor to be accessed.

[0053] In one possible implementation, the first tensor is a multidimensional tensor, which can be regarded as a multidimensional matrix. For example, the first tensor can be a two-dimensional tensor, a three-dimensional tensor, a four-dimensional tensor or other multi-dimensional tensor data. The embodiments of the present application do not limit the dimension of the first tensor.

[0054] In one possible implementation, the first tensor includes feature data in a deep learning task, where the feature data includes at least one of image feature data, speech feature data, and text feature data.

[0055] For example, in a scenario where a deep neural network is used to perform face recognition on a target object, the first tensor may be image feature data, and the image feature data of the target object (e.g., a face feature map) may be stored in a global memory; in a scenario where a deep neural network is used to perform speech recognition on a target object, the first tensor may be speech feature data, and the speech feature data of the target object may be stored in a global memory; in a scenario where a deep neural network is used to perform text recognition on a target document, the first tensor may be text feature data, and the text feature data of the target document may be stored in a global memory; the embodiments of the present application do not impose any restrictions on the type of the first tensor.

[0056] In one possible implementation, the tensor information of the first tensor may include data structure information of the first tensor and address information of the first tensor; wherein the data structure information of the first tensor includes, for example, the dimension of the first tensor, the size of the first tensor, the data type of the elements in the first tensor (for example, integer type, single-precision floating-point type, double-precision floating-point type, character type, etc.), and other information used to describe the first tensor; the address information of the first tensor includes, for example, the base address (Base Address) of the first tensor in the global memory, the addressing space, and other address-related information.

[0057] Since the tensor information of the first tensor has more content, the tensor information of the first tensor can be stored in the global memory.

[0058] In one possible implementation, when the second tensor to be accessed is the block data of the first tensor, the access information includes: the position of the second tensor in the first tensor, and the size of the second tensor. Step S11 may include: determining the address information of the second tensor from the address of the first tensor based on the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.

[0059] For example, Figure 3 illustrates a schematic diagram of determining address information of a second tensor according to an embodiment of the present disclosure. As shown in Figure 3, the second tensor block to be accessed is a block of data of the first tensor tensor, and the second tensor block can be data within any local area of ​​the first tensor tensor.

[0060] The first tensor tensor is a two-dimensional tensor. The tensor information of the first tensor tensor includes: the base address tensor_base_address used to determine the storage location of the first tensor tensor in the global memory, the size of the first tensor tensor in the horizontal dimension tensor_dim[0], and the size of the first tensor tensor in the vertical dimension tensor_dim[1]. Among them, since the first tensor tensor has a continuous address in the lowest dimension (such as the horizontal dimension or the vertical dimension), the tensor storage engine can access each element in the first tensor tensor according to the base address tensor_base_address of the first tensor tensor and the position of each element in the first tensor tensor.

[0061] As shown in FIG3 , the access information of the second tensor block includes: the position block_pos of the second tensor block in the first tensor tensor, the size block_dim[0] of the second tensor block in the horizontal dimension, and the size block_dim[1] of the second tensor block in the vertical dimension.

[0062] The tensor storage engine can determine the base address block_base_address of the second tensor block according to the base address tensor_base_address in the tensor information of the first tensor tensor and the position block_pos of the second tensor block in the first tensor tensor.

[0063] The tensor storage engine can determine the size of the addressing space of the second tensor block as block_dim[0]×block_dim[1] based on the size of the second tensor block in the horizontal dimension block_dim[0] and the size of the second tensor block in the vertical dimension block_dim[1].

[0064] The tensor storage engine can determine the address information of the second tensor block to be accessed in the global memory according to the base address block_base_address of the second tensor block and the size block_dim[0]×block_dim[1] of the addressing space of the second tensor block.

[0065] In this way, the tensor storage engine can quickly and accurately determine the address information of the second tensor (for example, any block data in the first tensor) from the address of the first tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed.

[0066] In the case where the second tensor to be accessed is the block data at the edge of the first tensor, there may be some area data of the second tensor that is not within the first tensor. In a possible implementation, the method further includes: firstly performing a boundary check on the second tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed, and obtaining a check result of the second tensor, wherein the check result is used to determine whether the second tensor is within the first tensor; when the check result indicates that the first part of the second tensor is within the first tensor and the second part of the second tensor is outside the first tensor, determining the address information of the first part of the second tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed.

[0067] Among them, the tensor information of the first tensor also includes a filling identifier (for example, the number 0, the number 1, etc.). The tensor storage engine can fill the second part of the second tensor according to the filling identifier to obtain the second part of the second tensor, so that in the subsequent step S12, in addition to reading the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor; the second part of the second tensor will also be synchronously written to the local memory so that the local memory stores the complete second tensor.

[0068] FIG4 shows a schematic diagram of determining the address information of the first part of the second tensor according to an embodiment of the present disclosure. As shown in FIG4 , the tensor information of the first tensor tensor includes: a base address tensor_base_address for determining the storage location of the first tensor tensor in the global memory, the size of the first tensor tensor in the horizontal dimension tensor_dim[0], and the size of the first tensor tensor in the vertical dimension tensor_dim[1]. The access information of the second tensor block includes: the position of the second tensor block in the first tensor tensor block block_pos, the size of the second tensor block in the horizontal dimension block_dim[0], and the size of the second tensor block in the vertical dimension block_dim[1]. Among them, the position of the second tensor block in the first tensor tensor block block_pos includes: the position of the second tensor block in the horizontal dimension block_pos[0] in the first tensor tensor, and the position of the second tensor block in the vertical dimension block_pos[1].

[0069] The inspection result of the second tensor block can be obtained based on the tensor information of the first tensor tensor and the access information of the second tensor block to be accessed. The inspection result is used to determine whether the second tensor block is located within the first tensor tensor.

[0070] If the sum of the horizontal position block_pos[0] of the second tensor block in the first tensor tensor and the size block_dim[0] of the second tensor block in the horizontal dimension is less than or equal to the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, that is: (block_pos[0]+block_dim[0])≤tensor_dim[0], and the vertical position block_pos[1] of the second tensor block in the first tensor tensor and the size block_dim[1] of the second tensor block in the vertical dimension is less than or equal to the size tensor_dim[1] of the first tensor tensor in the vertical dimension, that is: (block_pos[1]+block_dim[1])≤tensor_dim[1], it can be determined that the second tensor block is located within the first tensor tensor. The address information of the second tensor block can be determined by referring to the method shown in Figure 3 above, which will not be repeated here.

[0071] If the sum of the horizontal position of the second tensor block in the first tensor tensor, block_pos[0], and the horizontal size of the second tensor block, block_dim[0], is greater than the horizontal size of the first tensor tensor, tensor_dim[0], that is, (block_pos[0]+block_dim[0])>tensor_dim[0], and / or the sum of the vertical position of the second tensor block in the first tensor tensor, block_pos[1], and the vertical size of the second tensor block, block_dim[1], is greater than the vertical size of the first tensor tensor, tensor_dim[1], that is, (block_pos[1]+block_dim[1])>tensor_dim[1], it can be determined that some area of ​​the second tensor block is not within the first tensor tensor. As shown in Figure 4, the first part of the second tensor block to be accessed is within the first tensor tensor, and the second part of the second tensor block is outside the first tensor tensor.

[0072] When the first part of the second tensor block is located within the first tensor tensor and the second part of the second tensor block is located outside the first tensor tensor, the tensor storage engine can determine the address information of the first part of the second tensor block based on the tensor information of the first tensor tensor and the access information of the second tensor block to be accessed.

[0073] In the example, the tensor storage engine can determine the base address block_base_address of the second tensor block, that is, the base address block_base_address of the first part of the second tensor block, based on the base address tensor_base_address in the tensor information of the first tensor tensor and the position block_pos of the second tensor block in the first tensor tensor.

[0074] Optionally, when (block_pos[0]+block_dim[0])>tensor_dim[0], and (block_pos[1]+block_dim[1])>tensor_dim[1], the tensor storage engine can determine the size of the addressing space of the first part of the second tensor block according to the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, the size tensor_dim[1] of the first tensor tensor in the vertical dimension, the position block_pos[0] of the second tensor block in the horizontal dimension in the first tensor tensor, and the position block_pos[1] of the second tensor block in the vertical dimension in the first tensor tensor, that is, (tensor_dim[0]-block_pos[0])×(tensor_dim[1]-block_pos[1]).

[0075] The tensor storage engine can determine the address information of the first part of the second tensor block to be accessed in the global memory based on the base address block_base_address of the second tensor block and the size of the addressing space of the first part of the second tensor block (tensor_dim[0]-block_pos[0])×(tensor_dim[1]-block_pos[1]).

[0076] Optionally, when (block_pos[0]+block_dim[0])>tensor_dim[0], and (block_pos[1]+block_dim[1])≤tensor_dim[1] (see Figure 4), the tensor storage engine can determine the size of the addressing space of the first part of the second tensor block according to the size tensor_dim[0] of the first tensor tensor in the horizontal dimension, the position block_pos[0] of the second tensor block in the horizontal dimension in the first tensor tensor, and the size block_dim[1] of the second tensor block in the vertical dimension, that is: (tensor_dim[0]-block_pos[0])×block_dim[1].

[0077] The tensor storage engine can determine the address information of the first part of the second tensor block to be accessed in the global memory according to the base address block_base_address of the second tensor block and the size of the addressing space of the first part of the second tensor block (tensor_dim[0]-block_pos[0])×block_dim[1].

[0078] Optionally, when (block_pos[0]+block_dim[0])≤tensor_dim[0] and (block_pos[1]+block_dim[1])>tensor_dim[1], the tensor storage engine can determine the size of the addressing space of the first part of the second tensor block according to the vertical dimension size tensor_dim[1] of the first tensor tensor, the vertical dimension position block_pos[1] of the second tensor block in the first tensor tensor, and the horizontal dimension size block_dim[0] of the second tensor block, that is, block_dim[0]×(tensor_dim[1]-block_pos[1]).

[0079] The tensor storage engine can determine the address information of the first part of the second tensor block to be accessed in the global memory according to the base address block_base_address of the second tensor block and the size block_dim[0]×(tensor_dim[1]-block_pos[1]) of the addressing space of the first part of the second tensor block.

[0080] Among them, the tensor information of the first tensor tensor also includes a filling identifier (such as the number 0, the number 1, etc.). The tensor storage engine can fill the second part of the second tensor block according to the filling identifier to obtain the second part of the second tensor block, so that in the subsequent step S12, in addition to reading the first part of the second tensor block from the global memory storing the first tensor tensor to the local memory according to the address information of the first part of the second tensor block; the second part of the filled second tensor block will also be synchronously written to the local memory so that the local memory stores the complete second tensor block.

[0081] In this way, the address information of the second tensor can be determined from the address of the first tensor based on the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor. The second tensor can be the block data obtained by cutting the first tensor to any size, which is conducive to the subsequent tensor storage engine to access any block data in the first tensor that occupies a large space from the global memory through a single instruction, thereby improving the applicability of the data processing method of the embodiment of the present disclosure. In addition, the tensor storage engine can also realize automatic address calculation and boundary checking for block data, and can realize filling processing of arbitrary constant values ​​for out-of-bounds access.

[0082] In one possible implementation, the convolution kernel can traverse the entire first tensor by sliding on the first tensor. In order to facilitate the convolution operation and / or deconvolution operation of the first tensor and further improve the applicability of the data processing method of the embodiment of the present disclosure, when the second tensor to be accessed is the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, the access information includes: the coordinates of the expansion matrix and the coordinates of the convolution kernel elements. The expansion matrix is ​​in the first tensor, and the coverage data corresponding to the convolution kernel receptive field is converted into corresponding column vectors or row vectors in sequence according to the description information, and arranged in rows or columns. The description information includes convolution description information and deconvolution description information. Step S11 may include: obtaining the mapping relationship between the expansion matrix and the first tensor; determining the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the expansion matrix coordinates, and the convolution kernel element coordinates.

[0083] The tensor information of the first tensor may include data structure information of the first tensor (e.g., the dimension of the first tensor, the size of the first tensor, the data type of the elements in the first tensor, etc.) and address information of the first tensor (e.g., the base address of the first tensor in the global memory, the addressing space, etc.). The mapping relationship is used to determine the address information of each element in the unfolded matrix in the global memory space; the unfolded matrix coordinates and the convolution kernel element coordinates are used to locate the second tensor in the unfolded matrix.

[0084] The description information is used to provide an expansion method for expanding the original matrix into an expanded matrix suitable for the convolution kernel. The description information includes the convolution kernel size, the stride of the convolution kernel, the padding amount, and the dilation_rate. If the description information is deconvolution description information, and the stride in the deconvolution description information is greater than 1, in the process of converting the coverage data corresponding to the receptive field of the convolution kernel into the corresponding column vector or row vector and arranging them into an expanded matrix in rows or columns, a constant value padding operation can be performed between each element in the coverage data according to the padding identifier (such as the number 0, the number 1, etc.) included in the tensor information of the first tensor.

[0085] Figure 5 shows a schematic diagram of the mapping relationship between the first tensor and the unfolded matrix according to an embodiment of the present disclosure. As shown in Figure 5, based on the tensor information of the first tensor obtained, the tensor storage engine knows that the first tensor is a four-dimensional tensor N×Cx×H×W (see the two rectangular blocks on the upper side of Figure 5, whose dimensions are 2×2×5×4), where N represents the batch dimension, Cx represents the channel dimension, H represents the height dimension, and W represents the width dimension.

[0086] As shown in Figure 5, the numerical labels 0 to 79 in the first tensor are used to distinguish the elements of the first tensor, and each element occupies 128 bytes (80 HB in hexadecimal) of storage space in the global memory. Thus, assuming that the base address of the first tensor in the global memory is address, the address information of element 0 of the first tensor in the global memory is address to (address+007FH), the address information of element 1 of the first tensor in the global memory is (address+0080H) to (address+00FFH), the address information of element 2 of the first tensor in the global memory is (address+0100H) to (address+017FH), and so on. The address information of element 79 of the first tensor in the global memory is (address+2780H) to (address+27FFH).

[0087] The two-dimensional matrix on the lower side of Figure 5 is the expanded matrix of the first tensor. The expanded matrix is ​​in the first tensor. According to the convolution description information (for example, the convolution kernel size is 3×3, the stride of the convolution kernel is 1, and the padding is 0), the coverage data corresponding to the receptive field of the convolution kernel is converted into corresponding row vectors in the order in which the convolution kernel slides, and the two-dimensional matrix is ​​arranged in rows or columns.

[0088] As shown in Figure 5, assume that P represents the position where the convolution kernel slides in the height direction, Q represents the position where the convolution kernel slides in the width direction, Cx represents the position where the convolution kernel slides in the channel direction, and N represents whether the convolution kernel slides in the N1×Cx×H×W part of the first tensor or the N2×Cx×H×W part of the first tensor.

[0089] As shown in Figure 5, the first convolution kernel slides to the Q0 position in the width direction, slides to the P0 position in the height direction, and slides to the Cx=0 position in the width direction. The coverage data corresponding to the N1×Cx×H×W part of the convolution kernel receptive field in the first tensor is It can be converted to [0 1 2 4 5 6 8 9 10], and the coverage data corresponding to the N2×Cx×H×W part of the convolution kernel receptive field in the first tensor is This can be converted to [40 41 42 44 45 46 48 49 50].

[0090] The second convolution kernel slides to the Q1 position in the width direction, remains at the P0 position in the height direction, and slides to the Cx=0 position in the width direction. The coverage data corresponding to the N1×Cx×H×W part of the convolution kernel receptive field in the first tensor is It can be converted to [1 2 3 5 6 7 9 10 11], and the coverage data corresponding to the N2×Cx×H×W part of the convolution kernel receptive field in the first tensor is This can be converted to [41 42 43 45 46 47 49 50 51].

[0091] The third convolution kernel slides to the Q0 position in the width direction, remains at the P1 position in the height direction, and slides to the Cx=0 position in the width direction. The coverage data corresponding to the N1×Cx×H×W part of the convolution kernel receptive field in the first tensor is It can be converted to [4 5 6 8 9 10 12 13 14], and the coverage data corresponding to the N2×Cx×H×W part of the convolution kernel receptive field in the first tensor is This can be converted to [44 45 46 48 49 50 52 53 54].

[0092] And so on, until the 12th convolution kernel slides to the Q1 position in the width direction, slides to the P2 position in the height direction, and slides to the Cx=1 position in the width direction. The coverage data corresponding to the N1×Cx×H×W part of the convolution kernel receptive field in the first tensor is It can be converted to [29 30 31 33 34 35 37 38 39], and the coverage data corresponding to the N2×Cx×H×W part of the convolution kernel receptive field in the first tensor is This can be converted to [69 70 71 73 74 75 77 78 79].

[0093] The row vectors after the 12 coverage data conversions are arranged in rows or columns to obtain the expanded matrix shown in Figure 5. The mapping relationship between the expanded matrix and the first tensor can be obtained based on the correspondence between the first tensor and the elements with the same numerical labels in the expanded matrix. In the example, the tensor storage engine can directly obtain the mapping relationship between the expanded matrix input by the user and the first tensor; or, the tensor storage engine can first perform the expansion processing as described above on the first tensor read from the global memory based on the tensor information and description information of the obtained first tensor, determine the expanded matrix, and then obtain the mapping relationship between the expanded matrix and the first tensor based on the correspondence between the first tensor and the expanded matrix.

[0094] It should be understood that Figure 5 shows the case where the description information is convolution description information. If the description information is deconvolution description information, you can refer to the above introduction, and when the stride of the convolution kernel in the deconvolution description information is greater than 1, a constant value filling operation is performed between each element in the unfolded matrix, which will not be repeated here.

[0095] In a possible implementation, the address information of the second tensor is determined from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolding matrix coordinates, and the convolution kernel element coordinates, including: determining the second tensor to be accessed from the unfolding matrix according to the unfolding matrix coordinates and the convolution kernel element coordinates; determining the multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolding matrix and the first tensor; and determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.

[0096] In the unfolded matrix shown in Figure 5, NPQ can be used as the coordinate of the unfolded matrix in the vertical dimension to locate a row in the unfolded matrix. For example, N1P0Q0 represents the first row, N1P0Q1 represents the second row, N1P1Q0 represents the third row, N1P1Q1 represents the fourth row, N1P2Q0 represents the fifth row, N1P2Q1 represents the sixth row, N2P0Q0 represents the seventh row, N2P0Q1 represents the eighth row, N2P1Q0 represents the ninth row, N2P1Q1 represents the tenth row, N2P2Q0 represents the eleventh row, and N2P2Q1 represents the twelfth row.

[0097] Cx may be used as the coordinate of the unfolded matrix in the horizontal dimension, where Cx=0 is used to locate the 8 columns on the left side of the unfolded matrix, and Cx=1 is used to locate the 8 columns on the right side of the unfolded matrix.

[0098] RS is the coordinate of the convolution kernel element, R represents the coordinate of the vertical dimension of the convolution kernel, which is used to locate a row of the convolution kernel, and S represents the coordinate of the horizontal dimension of the convolution kernel, which is used to locate a column of the convolution kernel. For example, R=0, S=0 represents the convolution kernel element in the first column of the first row, R=0, S=1 represents the convolution kernel element in the second column of the first row, R=0, S=2 represents the convolution kernel element in the third column of the first row, R=1, S=0 represents the convolution kernel element in the first column of the second row, R=1, S=1 represents the convolution kernel element in the second column of the first row, R=1, S=2 represents the convolution kernel element in the third column of the second row, R=2, S=0 represents the convolution kernel element in the first column of the third row, R=2, S=1 represents the convolution kernel element in the second column of the third row, R=2, S=2 represents the convolution kernel element in the third column of the third row.

[0099] According to the coordinates Cx of the unfolded matrix in the horizontal dimension and the coordinates RS of the convolution kernel elements, we can locate any one or more columns of the second tensor in the unfolded matrix, that is, the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor.

[0100] For example, Cx=0, R=0, S=0 represents the first column in the unfolded matrix, that is, the first row and first column elements in the convolution kernel slide multiple times (for example, 12 times) to the corresponding coverage data in the first tensor; Cx=0, R=0, S=1 represents the second column in the unfolded matrix, that is, the first row and second column elements in the convolution kernel slide multiple times (for example, 12 times) to the corresponding coverage data in the first tensor; and so on, Cx=1, R=2, S=2 represents the eighteenth column (the last column) in the unfolded matrix, that is, the third row and third column elements in the convolution kernel slide multiple times (for example, 12 times) to the corresponding coverage data in the first tensor.

[0101] By locating any one or more columns of the second tensor in the unfolded matrix, the multiple elements corresponding to the second tensor in the first tensor can be determined based on the mapping relationship between the unfolded matrix and the first tensor, and the address information of the multiple elements corresponding to the second tensor in the first tensor can be determined as the address information of the second tensor.

[0102] For example, assuming that the second tensor is the first column [0 1 4 5 8 9 40 41 44 45 48 49] in the unfolded matrix, which corresponds to the coverage data of the first row and first column elements in the convolution kernel in the first tensor, we can obtain the elements with the same number as the elements in the first column of the unfolded matrix from the first tensor according to the mapping relationship between the unfolded matrix and the first tensor, and determine the addresses of the elements [0 1 4 5 8 9 40 41 44 45 48 49] in the first tensor as the address information of the second tensor.

[0103] For another example, assuming that the second tensor is the second column [1 2 5 6 9 10 41 42 45 46 49 50] in the unfolded matrix, which corresponds to the coverage data of the first row and second column elements in the convolution kernel in the first tensor, we can obtain the elements with the same number as the elements in the second column of the unfolded matrix from the first tensor according to the mapping relationship between the unfolded matrix and the first tensor, and determine the addresses of the elements [1 2 5 6 9 10 41 42 45 46 49 50] in the first tensor as the address information of the second tensor.

[0104] In practical applications, the above process can be encapsulated as a TensorMemoryEngine instruction, which can be used to access one or more columns in the unfolded matrix (corresponding to the elements of the first tensor covered by one or more convolution kernel elements). This is beneficial for the tensor storage engine to access the coverage data corresponding to any one or more convolution kernel elements in the first tensor that occupies a large space from the global memory through a single instruction. The pseudo code is as follows:

[0105] In this way, the tensor storage engine does not need to execute the process of expanding the first tensor into an expanded matrix, and does not need to occupy a large amount of storage space to record the expanded matrix of the first tensor. In the scenario where the expanded matrix is ​​not obtained, the tensor storage engine can calculate the coordinates of any element in the expanded matrix in the first tensor (for example, the coordinates of the elements of the first tensor covered by one or more convolution kernel elements) based on the expanded matrix coordinates NPQCx and the convolution kernel element coordinates RS, which is conducive to using a small amount of hardware resources to simplify the software implementation logic (for example, software implementation of convolution and deconvolution operations).

[0106] In this way, the tensor storage engine can quickly and accurately determine the address information of the second tensor (such as the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor) from the address of the first tensor based on the tensor information of the first tensor and the access information of the second tensor to be accessed (such as the unfolded matrix coordinates, the convolution kernel element coordinates).

[0107] In the data processing method of the embodiment of the present invention, the second tensor can be the block data of the first tensor, or the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, or the data composed of coverage data and constant value filling between the elements in the coverage data. It can support the block loading of the first tensor, the block loading of the expanded data suitable for convolution operation of the first tensor, and the block loading of the expanded data suitable for deconvolution operation of the first tensor, unifying multiple loading modes and making it more user-friendly for programming.

[0108] In a possible implementation, the tensor information of the first tensor is stored in a global memory, and the access information of the second tensor to be accessed is stored in a register.

[0109] Since the first tensor contains a lot of information and its information is not easily changed, the first tensor's information can be stored in the global memory. Since the second tensor to be accessed is the block data of any area in the first tensor, or the second tensor to be accessed is the corresponding coverage data of any one or more elements in the convolution kernel in the first tensor, the access information of the second tensor is variable, and the access information of the second tensor can be stored in a register to facilitate its modification during program execution.

[0110] This method of storing the tensor information of the unchanged first tensor in the global memory and the access information of the changing second tensor to be accessed in the register places various information in different places according to whether the information is prone to change, which is conducive to improving the flexibility and efficiency of data processing.

[0111] Compared with the synchronous memory access method in the related art, data needs to be loaded into the register file. For example, for the General-Purpose Computing on Graphics Processing Units (GPGPU) single instruction multiple thread (SIMT) architecture, the registers of each thread are isolated from each other and cannot be shared. Data also needs to be moved from the register to the shared storage space. In the embodiment of the present disclosure, since the tensor storage engine is connected to the register that stores the second tensor access information, the tensor storage engine can access the register at any time, so that multiple threads corresponding to the tensor storage engine can share the register. In addition, the processing method of the embodiment of the present disclosure can also adopt the register transfer method to improve the flexibility of data processing.

[0112] In step S11, the address information of the second tensor is determined. In step S12, the second tensor can be read from the global memory storing the first tensor to the local memory according to the address information of the second tensor. The local memory is closer to the tensor storage engine than the global memory, and the tensor storage engine can have higher efficiency in reading and writing operations on the local memory. The tensor storage engine can perform data movement processing on the first tensor stored in the global memory, and store the second tensor determined based on the first tensor in the local memory corresponding to the tensor storage engine.

[0113] In one possible implementation, step S12 may include: when the size of the second tensor is greater than the read bandwidth, splitting the first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests based on the address information of the second tensor, each second memory access request being used to read a different portion of the split data of the second tensor from the global memory to the local memory; and serially writing the multiple split data constituting the second tensor in the global memory to the local memory based on the multiple second memory access requests. The read bandwidth refers to the ability to read data from the memory per unit time, and the read bandwidth of the memory can be expressed in terms of the number of bits transferred per time or the number of bytes transferred per time.

[0114] For example, assuming that the read bandwidth of the global memory is 128 bytes, the first memory access request is to access the 1024-bit second tensor, and its corresponding address is 0, but the read bandwidth supports a maximum of 128-bit memory access requests. The first memory access request can be split into eight second memory access requests, namely: second memory access request 0 to second memory access request 7, wherein the second memory access request 0 is used to read the 128-bit split data 0 in the second tensor from address 0 in the global memory to the local memory; the second memory access request 0 is used to read the 128-bit split data 0 belonging to the second tensor from address 0 in the global memory to the local memory; the second memory access request 1 is used to read the 128-bit split data 1 belonging to the second tensor from address 128 in the global memory to the local memory; the second memory access request 2 is used to read the 128-bit split data 1 belonging to the second tensor from address 256 in the global memory, Read the 128-bit split data 2 belonging to the second tensor to the local memory; a second memory access request 3 is used to read the 128-bit split data 3 belonging to the second tensor from address 384 in the global memory to the local memory; a second memory access request 4 is used to read the 128-bit split data 4 belonging to the second tensor from address 512 in the global memory to the local memory; a second memory access request 5 is used to read the 128-bit split data 5 belonging to the second tensor from address 640 in the global memory to the local memory; a second memory access request 6 is used to read the 128-bit split data 6 belonging to the second tensor from address 768 in the global memory to the local memory; a second memory access request 7 is used to read the 128-bit split data 7 belonging to the second tensor from address 896 in the global memory to the local memory. Then, according to the second memory access requests 0 to 7, the split data 0 to 7 constituting the second tensor in the global memory can be serially written to the local memory. It should be understood that the present disclosure does not impose any specific restrictions on the size of the second tensor and the size of the reading bandwidth, which can be set according to the actual application scenario.

[0115] In this way, if the size of the second tensor in the global memory exceeds the maximum granularity supported by the storage system, it can be split and processed serially, thereby improving the applicability of the data processing method of the embodiment of the present disclosure. In addition, in the embodiment of the present disclosure, the read bandwidth of the global memory is smaller than the read bandwidth of the local memory. When reading data (such as the second tensor) from the global memory to the local memory, the local memory can be reused multiple times, reducing the bandwidth pressure on the global memory.

[0116] In step S12, the second tensor is read from the global memory storing the first tensor to the local memory. In step S13, an interrupt is triggered in response to the local memory receiving the second tensor. In step S12, the tensor storage engine performs the process of transferring the second tensor from the global memory to the local memory, which may take a long time. Before waiting for the local memory to receive the second tensor and trigger the interrupt, the tensor storage engine can asynchronously execute other instructions to further improve processing efficiency.

[0117] In this way, by setting up an interrupt wake-up mechanism, you do not need to follow the pipeline method (for example, the next instruction can be executed only after the data is moved to the destination). Instead, you can execute other instructions after the current instruction is issued, and wait for the current instruction to be completed to send a signal to the processor core to trigger an interrupt, waking up the current instruction, which will be more efficient.

[0118] The data processing method of the embodiment of the present disclosure can quickly read the second tensor from the global memory storing the first tensor to the local memory based on the tensor information of the first tensor and the access information of the second tensor to be accessed, thereby improving the efficiency and flexibility of tensor storage. Moreover, each time the local memory receives the second tensor, an interrupt can be triggered to realize the asynchronous tensor reading process.

[0119] Among them, the second tensor of the embodiment of the present disclosure can be the block data of the first tensor, or the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, or the data composed of coverage data and constant value filling between each element in the coverage data. It can support the block loading of the first tensor, the block loading of the expanded data suitable for convolution operation of the first tensor, and the block loading of the expanded data suitable for deconvolution operation of the first tensor, unifying multiple loading modes and making it more user-friendly for programming.

[0120] Among them, the processing method of the embodiment of the present disclosure can store the tensor information of the unchanged first tensor in the global memory, and store the access information of the changing second tensor to be accessed in the register. This method of placing various information in different places according to whether the information is easy to change is conducive to improving the flexibility and efficiency of data processing.

[0121] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0122] In addition, the present disclosure also provides a data processing device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any data processing method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.

[0123] FIG6 shows a block diagram of a data processing device according to an embodiment of the present disclosure. As shown in FIG6 , the device includes:

[0124] a determination module 61, configured to determine address information of the second tensor to be accessed based on the tensor information of the first tensor and the access information of the second tensor to be accessed;

[0125] a reading module 62, configured to read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor;

[0126] The triggering module 63 is configured to trigger an interrupt in response to the local memory receiving the second tensor.

[0127] In one possible implementation, the determination module 61 is used to: when the second tensor to be accessed is the block data of the first tensor, and the access information includes the position of the second tensor in the first tensor and the size of the second tensor, determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.

[0128] In one possible implementation, the determination module 61 is further used to: perform a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, to obtain a check result of the second tensor; when the check result indicates that the first part of the second tensor is located within the first tensor and the second part of the second tensor is located outside the first tensor, determine the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; the reading module 62 is used to: when the tensor information of the first tensor includes a fill identifier, read the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor, fill the second part of the second tensor according to the fill identifier to obtain the second part of the second tensor; and write the second part of the second tensor into the local memory.

[0129] In one possible implementation, the determination module 61 is used to: when the second tensor to be accessed is the coverage data corresponding to any one or more elements in the convolution kernel in the first tensor, and the access information includes the expansion matrix coordinates and the convolution kernel element coordinates, obtain the mapping relationship between the expansion matrix and the first tensor; determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the expansion matrix coordinates, and the convolution kernel element coordinates; wherein the expansion matrix is ​​in the first tensor, and the coverage data corresponding to the convolution kernel receptive field is converted into corresponding column vectors or row vectors in sequence according to the description information, and the matrix is ​​arranged in rows or columns, and the description information includes convolution description information and deconvolution description information.

[0130] In a possible implementation, the address information of the second tensor is determined from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolding matrix coordinates, and the convolution kernel element coordinates, including: determining the second tensor to be accessed from the unfolding matrix according to the unfolding matrix coordinates and the convolution kernel element coordinates; determining the multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolding matrix and the first tensor; and determining the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.

[0131] In one possible implementation, the reading module 62 is used to: when the size of the second tensor is greater than the read bandwidth, split the first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests according to the address information of the second tensor, each second memory access request is used to read different parts of the split data in the second tensor from the global memory to the local memory; according to the multiple second memory access requests, write the multiple split data constituting the second tensor in the global memory into the local memory serially.

[0132] In a possible implementation, the tensor information of the first tensor is stored in a global memory, and the access information of the second tensor to be accessed is stored in a register.

[0133] In one possible implementation, the first tensor includes feature data in a deep learning task, where the feature data includes at least one of image feature data, speech feature data, and text feature data.

[0134] This method has a specific technical connection with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware computing efficiency or execution effect (including reducing the amount of data storage, reducing the amount of data transmission, increasing the hardware processing speed, etc.), thereby obtaining the technical effect of improving the internal performance of the computer system in accordance with the laws of nature.

[0135] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0136] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0137] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0138] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0139] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0140] The present disclosure also provides an electronic device, which can be provided as a terminal, a server, or other device, such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc.

[0141] FIG7 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to FIG7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions that can be executed by the processing component 1922, such as an application. The application stored in the memory 1932 can include one or more modules, each of which corresponds to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0142] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (Mac OS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ) or similar.

[0143] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0144] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0145] Computer-readable storage media can be a tangible device that can hold and store the instructions used by the instruction execution device. Computer-readable storage media can be, for example, (but not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. Computer-readable storage media used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0146] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0147] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0148] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0149] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0150] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0151] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0152] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0153] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0154] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0155] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A data processing method, characterized in that Including: Determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed; Read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor; Trigger an interrupt in response to the local memory receiving the second tensor.

2. The method according to claim 1, characterized in that When the second tensor to be accessed is the block data of the first tensor, the access information includes: the position of the second tensor in the first tensor and the size of the second tensor. Determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, including: Determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the position of the second tensor in the first tensor, and the size of the second tensor.

3. The method according to claim 2, wherein The method further includes: Perform a boundary check on the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, and obtain the check result of the second tensor; When the check result indicates that the first part of the second tensor is within the first tensor and the second part of the second tensor is outside the first tensor, determine the address information of the first part of the second tensor according to the tensor information of the first tensor and the access information of the second tensor; The tensor information of the first tensor includes a padding flag. Reading the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor includes: Read the first part of the second tensor from the global memory storing the first tensor to the local memory according to the address information of the first part of the second tensor; Pad the second part of the second tensor according to the padding flag to obtain the second part of the second tensor; Write the second part of the second tensor into the local memory.

4. The method according to claim 1, characterized in that, When the second tensor to be accessed is the covering data corresponding to any one or more elements in the convolution kernel in the first tensor, the access information includes: the unfolded matrix coordinates and the convolution kernel element coordinates. The unfolded matrix is a matrix formed by sequentially converting the covering data corresponding to the convolution kernel receptive field into corresponding column vectors or row vectors in the first tensor and arranging them row by row or column by column. The description information includes convolution description information and deconvolution description information. Determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed, including: Obtain the mapping relationship between the unfolded matrix and the first tensor; Determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates.

5. The method according to claim 4, wherein Determine the address information of the second tensor from the address of the first tensor according to the tensor information of the first tensor, the mapping relationship, the unfolded matrix coordinates, and the convolution kernel element coordinates, including: Determine the second tensor to be accessed from the unfolded matrix according to the unfolded matrix coordinates and the convolutional kernel element coordinates; Determine multiple elements corresponding to the second tensor in the first tensor according to the mapping relationship between the unfolded matrix and the first tensor; Determine the address information of the multiple elements corresponding to the second tensor in the first tensor as the address information of the second tensor.

6. The method according to any one of claims 1-5, characterized in that, Read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor, including: When the size of the second tensor is greater than the read bandwidth, split the first memory access request for reading the second tensor from the global memory to the local memory into multiple second memory access requests according to the address information of the second tensor, and each second memory access request is used to read split data of different parts of the second tensor from the global memory to the local memory; Write the multiple split data constituting the second tensor in the global memory to the local memory serially according to the multiple second memory access requests.

7. The method according to any one of claims 1-5, characterized in that, The tensor information of the first tensor is stored in the global memory, and the access information of the second tensor to be accessed is stored in the register.

8. The method according to any one of claims 1-5, characterized in that, The first tensor includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.

9. A data processing device, characterized in that, Including: A determination module, configured to determine the address information of the second tensor according to the tensor information of the first tensor and the access information of the second tensor to be accessed; A reading module, configured to read the second tensor from the global memory storing the first tensor to the local memory according to the address information of the second tensor; A triggering module, configured to trigger an interrupt in response to the local memory receiving the second tensor.

10. An electronic device, characterized in that, Including: A memory and a processor; Computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 8 is executed.

11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

12. A computer program product, including computer-readable code, or a computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Memory devices and methods which may facilitate tensor memory access

    CN112470133A

  • Tensor processing method, device and equipment and computer readable storage medium

    CN114764489A

  • Calculation resource allocation method and device for tensor calculation graph and readable storage medium

    CN116483550A

  • Tensor-based memory access

    US11321092B1

  • Operation Accelerator, Processing Method, and Related Device

    US20210224125A1

Cited By

  • Memory data exchange transmission processing method and transmission device

    CN121743247A