Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By obtaining the calculation diagram in the data processing method, determining the core operators, and determining the arrangement method of each operator through iterative operations combined with the arrangement method of adjacent operators, the efficiency problem caused by failure to consider the relationship between operators in the prior art is solved, and the optimal arrangement method for the calculation diagram is realized, which improves data processing efficiency and system performance.

CN119903020BActive Publication Date: 2025-06-17SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510308282.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

When setting the operator arrangement method, the prior art fails to fully consider the relationship between operators, resulting in the inability to obtain the optimal data arrangement or memory access method for the entire computing graph, which affects data processing efficiency and system performance.

Method used

By obtaining the calculation diagram, the core operator is determined, and the arrangement method of each operator is determined through iterative operations. In each iteration, the starting operator is selected from the core operator, and the arrangement method is determined based on its read and write overhead under various preset arrangement methods, and the arrangement method of adjacent operators is propagated forward and backward, and the arrangement method of all operators is gradually determined.

Benefits of technology

It realizes the operator arrangement method that is optimal for the entire computing graph, improves data processing efficiency and system performance, and avoids data handling overhead caused by poor arrangement methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903020B_ABST
    Figure CN119903020B_ABST
Patent Text Reader

Abstract

A data processing method, apparatus, electronic device, and storage medium. The data processing method includes: obtaining a computational graph; determining at least one core operator based on the computational graph; determining the arrangement modes respectively corresponding to a plurality of operators through at least one round of iterative operations; wherein, in each round of iterative operation: selecting one or more core operators from the at least one core operator as starting operators; determining the arrangement mode corresponding to each starting operator according to the read-write overheads of each starting operator under multiple preset arrangement modes; starting from the starting operators, combining the arrangement mode corresponding to the starting operators, and propagating forward and backward according to the dependency relationship to sequentially obtain the arrangement modes corresponding to each operator, wherein, in the process of sequentially obtaining the arrangement modes corresponding to each operator, obtaining the arrangement mode corresponding to the current operator by combining the arrangement modes corresponding to adjacent operators, and the adjacent operators are operators that are adjacent to the current operator according to the dependency relationship and whose arrangement modes have been determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data processing method, a data processing apparatus, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] A tensor is a multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. For example, a scalar can be regarded as a 0-dimensional tensor, a vector can be regarded as a 1-dimensional tensor, a matrix can be regarded as a 2-dimensional tensor, and a tensor can have any number of dimensions. Tensor operations are widely used in processors such as parallel processors.

[0003] With the development of artificial intelligence and machine learning, new requirements are put forward for many parallel processor devices represented by parallel processors (such as multi-core processors, digital signal processors, etc.). In general computing, the computing units of parallel processors require a large amount of data, and this data is generally stored in the storage components of parallel processors. For example, the storage components can be memory. There are various data arrangement methods for the placement of tensor elements in the memory, and the processor provides various memory access methods for tensors. The data arrangement method and the memory access method have an important impact on the data transfer efficiency of tensors. Summary of the Invention

[0004] At least one embodiment of the present disclosure provides a data processing method, including: obtaining a computation graph, where the computation graph includes a plurality of operators and the dependency relationships between the plurality of operators; determining at least one core operator based on the computation graph; determining the arrangement methods respectively corresponding to the plurality of operators through at least one round of iterative operations, where the arrangement method corresponding to each operator includes the arrangement method of the input tensor and the output tensor of the operator, and the arrangement method is used to indicate the layout method of the tensor in the storage component; where, in each round of iterative operations: selecting one or more core operators from the at least one core operator as starting operators; determining the arrangement method corresponding to each starting operator according to the read-write overheads of each starting operator under a plurality of preset arrangement methods; starting from the starting operator, combining the arrangement method corresponding to the starting operator, and propagating forward and backward according to the dependency relationships to sequentially obtain the arrangement methods corresponding to each operator, where, in the process of sequentially obtaining the arrangement methods corresponding to each operator, the arrangement method corresponding to the current operator is obtained by combining the arrangement methods corresponding to adjacent operators, and the adjacent operator is an operator that is adjacent to the current operator according to the dependency relationships and whose arrangement method has been determined.

[0005] For example, in the data processing method provided by at least one embodiment of the present disclosure, based on the computational graph, determining at least one core operator may include: selecting the at least one core operator from the multiple operators based on a predetermined condition; wherein, the predetermined condition includes at least one of the following: a fused operator, an operator with a relatively high computational volume when the multiple operators are sorted in descending order of computational volume, an operator of a predetermined functional type, wherein the fused operator is an operator obtained by fusing multiple operators through an optimization strategy.

[0006] For example, in the data processing method provided by at least one embodiment of the present disclosure, selecting one or more core operators from the at least one core operator as starting operators includes: dividing the at least one core operator into at least one priority level according to the importance degree; selecting one or more core operators belonging to the same priority level as the starting operators in each round of iterative operation; wherein, the at least one round of iterative operation is sequentially executed in descending order of the importance degree of the at least one priority level, the propagation stops when reaching the core operator, and this round of iterative operation ends when both the forward propagation and the backward propagation starting from the at least one starting operator reach the core operator.

[0007] For example, in the data processing method provided by at least one embodiment of the present disclosure, determining the arrangement mode corresponding to each starting operator according to the read / write overheads of each starting operator in multiple preset arrangement modes includes: for a target tensor whose arrangement mode has not been determined among the starting operators: in the first execution of the iterative operation, selecting the preset arrangement mode with the smallest read / write overhead from the read / write overheads of the target tensor in the multiple preset arrangement modes as the arrangement mode of the target tensor, wherein the target tensor includes an input tensor or an output tensor; in the execution of a non-first iterative operation, in response to the tensor corresponding to the target tensor not having its arrangement mode determined, selecting the preset arrangement mode with the smallest read / write overhead from the read / write overheads of the target tensor in the multiple preset arrangement modes as the arrangement mode of the target tensor, wherein the corresponding tensor is loaded as the target tensor, or the target tensor is loaded as the corresponding tensor; in response to the tensor corresponding to the target tensor having its arrangement mode determined, selecting the preset arrangement mode with the smallest read / write overhead from the multiple read / write overheads corresponding to the target tensor as the arrangement mode of the target tensor, wherein the multiple read / write overheads include the read / write overhead of the target tensor in the first preset arrangement mode and the read / write overheads when the target tensor is converted between multiple arrangement modes, the first preset arrangement mode is the arrangement mode determined to be used by the corresponding tensor, and the multiple arrangement mode conversions include the conversion of the first preset arrangement mode to multiple other preset arrangement modes.

[0008] For example, in the data processing method provided by at least one embodiment of the present disclosure, selecting the preset arrangement method with the smallest read / write overhead from the read / write overheads of the target tensor under the multiple preset arrangement methods as the arrangement method of the target tensor includes: excluding at least one preset arrangement method from the multiple preset arrangement methods according to the operator characteristics of the starting operator; selecting the preset arrangement method with the smallest read / write overhead from the read / write overheads of the starting operator under the remaining preset arrangement methods as the arrangement method of the target tensor.

[0009] For example, in the data processing method provided by at least one embodiment of the present disclosure, starting from the starting operator, combining the arrangement method corresponding to the starting operator, and propagating forward and backward according to the dependency relationship to sequentially obtain the arrangement methods corresponding to each operator, including: combining the dependency relationship, performing topological sorting on the computational graph according to a preset topological sorting type; starting from the starting operator, sequentially performing propagation operations on each operator forward and backward according to the result of the topological sorting; where the propagation operation includes: determining the current operator, where when propagating backward starting from the starting operator, the current operator is the operator that uses the output tensor of the adjacent operator as the input tensor of the current operator, and when propagating forward starting from the starting operator, the current operator is the operator that outputs the output tensor to the adjacent operator as the input tensor of the adjacent operator; in response to the current operator not being a core operator, determining the arrangement method corresponding to the current operator based on the arrangement method corresponding to the adjacent operator; in response to the current operator being a core operator, stopping the current round of iterative operations.

[0010] For example, in the data processing method provided by at least one embodiment of the present disclosure, determining the arrangement method corresponding to the current operator based on the arrangement method corresponding to the adjacent operator includes: in response to using the output tensor of the adjacent operator as the input tensor of the current operator: selecting the preset arrangement method with the smallest read / write overhead from the first read / write overhead corresponding to the input tensor of the current operator and the second read / write overhead corresponding to the input tensor of the current operator as the arrangement method of the input tensor of the current operator, where the first read / write overhead is the read / write overhead in the first preset arrangement method, and the output tensor of the adjacent operator is determined to use the first preset arrangement method, and the second read / write overhead is the read / write overhead when converting the first preset arrangement method to at least one other preset arrangement method; selecting the preset arrangement method with the smallest read / write overhead from the read / write overheads of the output tensor of the current operator under the multiple preset arrangement methods as the arrangement method of the output tensor of the current operator.

[0011] For example, in the data processing method provided by at least one embodiment of the present disclosure, determining the arrangement mode corresponding to the current operator based on the arrangement mode corresponding to the adjacent operator includes: in response to the output tensor of the current operator being used as the input tensor of the adjacent operator: selecting, from the third read / write cost corresponding to the output tensor of the current operator and the fourth read / write cost corresponding to the output tensor of the current operator, the preset arrangement mode with the smallest read / write cost as the arrangement mode of the output tensor of the current operator, where the third read / write cost is the read / write cost in the second preset arrangement mode, the input tensor of the adjacent operator is determined to use the second preset arrangement mode, and the fourth read / write cost is the read / write cost when converting at least one other preset arrangement mode to the second preset arrangement mode; selecting, from the read / write costs of the input tensor of the current operator in the multiple preset arrangement modes, the preset arrangement mode with the smallest read / write cost as the arrangement mode of the input tensor of the current operator.

[0012] For example, in the data processing method provided by at least one embodiment of the present disclosure, in response to the number of executions of the iterative operation reaching a preset threshold and there being at least one operator in the computation graph whose arrangement mode has not been determined, determining the arrangement mode of the at least one operator based on a predetermined rule, where the predetermined rule specifies a condition requirement and a corresponding arrangement mode, and an operator that meets the corresponding condition requirement uses the corresponding arrangement mode.

[0013] For example, in the data processing method provided by at least one embodiment of the present disclosure, before determining the arrangement modes corresponding to the multiple operators through at least one round of iterative operations, the data processing method further includes: determining the shape and size of the tensors related to the multiple operators; for each operator, obtaining the read / write costs of the related tensors in the multiple preset arrangement modes according to the shape and size of the tensors related to the operator.

[0014] For example, in the data processing method provided by at least one embodiment of the present disclosure, the related tensors include an input tensor and an output tensor, and determining the shape and size of the tensors related to the multiple operators includes: deriving the shape and size of the input tensor and the output tensor of each operator according to the computation graph and the shape and size of the input activation tensor input to the computation graph.

[0015] For example, in the data processing method provided by at least one embodiment of the present disclosure, obtaining the read / write overheads of the relevant tensors under the multiple preset arrangement manners according to the shape dimensions of the tensors related to the operator includes: determining the amount of data to be transmitted for the relevant tensors under each preset arrangement manner according to the shape dimensions of the tensors related to the operator; and obtaining the read / write overheads of the relevant tensors under each preset arrangement manner based on the amount of data to be transmitted and the hardware information, where the hardware information includes the data transmission bandwidth under each preset arrangement manner.

[0016] For example, in the data processing method provided by at least one embodiment of the present disclosure, before determining the arrangement manners respectively corresponding to the multiple operators through at least one round of iterative operations, the data processing method further includes: obtaining the read / write overheads of the relevant tensors when the arrangement manner is converted according to the shape dimensions of the tensors related to the operator, where the arrangement manner conversion includes that the input tensor of the operator is stored in a storage component in a first preset arrangement manner and is loaded in a second preset arrangement manner when the input tensor is loaded from the storage component, and the first preset arrangement manner is different from the second preset arrangement manner.

[0017] For example, in the data processing method provided by at least one embodiment of the present disclosure, the arrangement manner includes at least one of a memory access manner and a data arrangement manner, and the data arrangement manner indicates the storage order and dimension arrangement of data in a storage component.

[0018] At least one embodiment of the present disclosure provides a data processing apparatus, including: an acquisition unit configured to acquire a computation graph, where the computation graph includes a plurality of operators and dependencies between the plurality of operators; a core operator determination unit configured to determine at least one core operator based on the computation graph; an iterative operation unit configured to determine the arrangement modes respectively corresponding to the plurality of operators through at least one round of iterative operations, where the arrangement mode corresponding to each operator includes the arrangement mode of the input tensor and the output tensor of the operator, and the arrangement mode includes at least one of a memory access mode and a data arrangement mode, and the data arrangement mode indicates the storage order and dimensional arrangement of data in a storage component; where, in each round of iterative operation: select one or more core operators from the at least one core operator as starting operators; determine the arrangement mode corresponding to each starting operator according to the read / write overheads of each starting operator in a plurality of preset arrangement modes; starting from the starting operator, in combination with the arrangement mode corresponding to the starting operator, propagate forward and backward according to the dependencies to sequentially obtain the arrangement modes corresponding to the respective operators, where, in the process of sequentially obtaining the arrangement modes corresponding to the respective operators, obtain the arrangement mode corresponding to the current operator by combining the arrangement modes corresponding to adjacent operators, and the adjacent operator is an operator adjacent to the current operator according to the dependencies and whose arrangement mode has been determined.

[0019] At least one embodiment of the present disclosure provides an electronic device, including: a memory storing computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, where the computer-executable instructions, when run by the processor, implement the data processing method according to at least one embodiment of the present disclosure.

[0020] At least one embodiment of the present disclosure provides a non-transient computer-readable storage medium, where the non-transient computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the data processing method according to at least one embodiment of the present disclosure.

[0021] In the data processing method provided by at least one embodiment of the present disclosure, first select starting operators from the core operators, determine the arrangement modes corresponding to the starting operators, and then use the starting operators as starting points for propagating the arrangement modes. Moreover, when determining the arrangement modes corresponding to the respective operators, consider the arrangement modes of adjacent operators in combination with the computation graph, that is, when determining the arrangement mode of each operator, not only consider the operator itself, but also combine the dependencies between the operators and the arrangement modes of adjacent operators to determine the arrangement mode of the current operator, so that the obtained arrangement modes of the respective operators are not only suitable for the operators themselves, but can obtain the optimal arrangement modes of the respective operators for the entire computation graph, improving data processing efficiency and system performance. Description of the Drawings

[0022] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0023] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU);

[0024] Figure 2A It is a schematic structure of a tensor;

[0025] Figure 2B It is a schematic diagram of the NCHW storage format;

[0026] Figure 2C It is a schematic diagram of the NHWC storage format;

[0027] Figure 2D It is a schematic diagram of the NC / 32HW32 storage format;

[0028] Figure 3 It is a schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure;

[0029] Figure 4A It shows a schematic diagram of a directed acyclic graph of a neural network;

[0030] Figure 4B It shows according to Figure 4A The implementation process of obtaining a computation graph by operator fusion according to the shown directed acyclic graph;

[0031] Figure 5 It is a schematic flowchart of the iterative operation provided by at least one embodiment of the present disclosure;

[0032] Figure 6 It is a schematic diagram of the sorting result after preset topological sorting shown in an embodiment of the present disclosure;

[0033] Figure 7 It is a schematic diagram of the execution process of the data processing method provided by at least one embodiment of the present disclosure;

[0034] Figure 8 It is a schematic block diagram of the data processing device provided by at least one embodiment of the present disclosure;

[0035] Figure 9 It is a schematic block diagram of the electronic device provided by at least one embodiment of the present disclosure;

[0036] Figure 10 It is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present disclosure with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0038] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should have the ordinary meaning understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits the detailed descriptions of some known functions and known components.

[0039] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU).

[0040] As Figure 1 shown, a general-purpose graphics processing unit is actually an array of programmable multi-processors. For example, the programmable multi-processors can be Streaming Processor Clusters (SPCs), such as including Figure 1 the Streaming Processor Cluster 1 shown, ..., the Streaming Processor Cluster M, where M is a positive integer greater than 1. In the general-purpose graphics processing unit, 1 Streaming Processor Cluster processes one computing task, or multiple Streaming Processor Clusters process one computing task. Data sharing between multiple Streaming Processor Clusters is achieved through global caches or global memory.

[0041] As Figure 1 shown, taking the Streaming Processor Cluster 1 as an example, 1 Streaming Processor Cluster includes multiple computing units, such as Figure 1The computing units 1, 2, …, N in it, where N is a positive integer. Each computing unit (Compute Unit, abbreviated as CU) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores or computing nuclei), and each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The computing cores are used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in it) and shared memory, which are used to hierarchically store the source data and destination data related to the computing tasks. The shared memory in a computing unit is used to share data among the cores of the computing unit.

[0042] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processing unit (or called a parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in it). All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0043] In each computing unit, a thread bundle scheduling / distribution module ( Figure 1 not shown in it) schedules and allocates the thread bundles so that the multiple computing cores in the computing unit can run the thread bundles. According to the number of computing cores in the computing unit, the multiple thread bundles in a thread block can be executed simultaneously or time-shared. The multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the middle-level cache or global cache or global memory for read / write operations, etc.

[0044] The elements of a tensor are arranged in various formats in the memory (such as Figure 1 the memory of it).

[0045] For example, Figure 2A is a schematic structure of a tensor. Suppose the tensor can be expressed as N×C×H×W, where N represents the batch size, for example, 1, C represents the number of feature maps (i.e., the number of channels), for example, 64, H represents the image height, for example, 5, and W represents the image width, for example, 4. The pixel elements of the tensor are represented as 0, 1, 2, 3, …, and so on. Below, taking Figure 2AThe tensors shown describe different layout formats.

[0046] For example, the data layout format may include NCHW (channels-height-width). Figure 2B Schematic diagram of the storage format for NCHW.

[0047] For example, for NCHW, refer to Figure 2B , starting from the first channel (c = 0), the elements in this channel are continuously arranged in row-major order, that is, in the order of 0, 1, 2, 3,..., 19, and then continue with the second channel (c = 1) and subsequent channels until all the elements of all channels are arranged. If N > 1, continue with the next batch.

[0048] For example, the data layout format can also include NHWC. Figure 2C Schematic diagram of the storage format for NHWC.

[0049] For example, for NHWC (height-width-channels), as Figure 2C shown, starting from the first element of the first channel (c = 0) (the element 0 in Figure 2C ), then store the first element of the second channel (c = 1) (the element 20 in Figure 2C ), and so on until all the first elements of all channels are laid out. For example, after the first element of the 64th channel (c = 63) (the element 1260 in Figure 2C ), select the second element of the first channel (c = 0) (the element 1 in Figure 2C ), then store the second element of the second channel (c = 1) (the element 21 in Figure 2C ), and so on until all the second elements of all channels are laid out, and so on.

[0050] For example, the data layout format can also include NC / xHWx (interleaved mode), where x can take values such as 8, 16, 32, etc. according to needs.

[0051] NC / xHWx is similar to NHWC, but there is a key difference. In the memory layout of NC / xHWx, C channels are divided into C / x groups, with each group having x channels: the first group consists of channels c = 0 to c = x - 1, the second group consists of channels c = x to c = 2x - 1, and each group is arranged in NHWC format.

[0052] Figure 2D Schematic diagram of the storage format for NC / 32HW32.

[0053] As Figure 2DAs shown, 64 channels are divided into two groups, with 32 channels in each group. The first group consists of channels c = 0 to c = 31, and the second group consists of channels c = 32 to c = 63. Then each group is arranged in NHWC format.

[0054] Of course, there can be more other data layout formats, such as BLOCK_ROW_MAJOR (block-row major order), BLOCK_COL_MAJOR (block-column major order), etc. The present disclosure will not list them one by one.

[0055] The processor can also have multiple memory access modes. For example, typical memory access modes include Unified Memory Access (UMA for short) and Non-Uniform Memory Access (NUMA for short). Taking Figure 1 the shown graphics processor architecture as an example, assuming that the streaming processor cluster is used as the computing core, if unified memory access is used, each streaming processor cluster can access all storage spaces; if non-uniform memory access is used, each streaming processor cluster only accesses the storage space closest to itself, so the data transfer efficiency is higher.

[0056] Different types of memories (such as registers, caches, main memories, and auxiliary storages) have different access speeds. Registers and caches are fast but have small capacities; while auxiliary storages such as hard disks are slow but have large capacities. Through the storage hierarchy, a balance can be achieved between speed and capacity.

[0057] The convolution kernel tensor is represented as N×Cin×Kh×Kw, where N represents the number of batches, Cin represents the number of input channels, Kh represents the height of the convolution kernel, and Kw represents the width of the convolution kernel. Taking conventional convolution calculation as an example, convolution usually performs reduction operations in three dimensions of Cin, Kh, and Kw. Assuming that the input activation tensor is arranged in NCHW data layout, the data required for one convolution may not be continuous. Especially when the W of the input activation tensor is very large, it is easy to cause cache misses. Additionally, if SIMD (Single Instruction Multiple Data) instructions are used for calculation, assuming that SIMD instructions can process 4 data at a time, due to the discontinuity of the data, the acceleration effect of SIMD instructions cannot be achieved. However, if the NC / 4HW4 layout is adopted, SIMD instructions can be used for acceleration. Therefore, the accurate selection of data layout and memory access modes has a crucial impact on data processing efficiency and data transfer efficiency.

[0058] A neural network can be regarded as a directed acyclic graph (DAG) composed of many computational nodes, where each node corresponds to an operator. In a neural network, an operator usually refers to the basic mathematical operations or operations used in the model network layer. These operators are used to construct each layer and component of the network to achieve data transfer, transformation, and calculation. They are the basic building blocks of the network model, defining the structure and operation process of the model, including input, output, and intermediate calculations. In the model, the connection relationships between operators form a directed graph, reflecting the calculation order of different operations in the model. By combining these operators, a complex and powerful neural network model can be constructed to handle various complex tasks and data. For example, the following are some common neural network operators: linear layer, convolutional layer, pooling layer, recurrent layer, activation function layer, etc.

[0059] The inventors found that currently, when setting the arrangement method of operators, only the impact of a certain data arrangement method or memory arrangement method on the calculation efficiency is considered. For example, the arrangement method of a certain operator is manually specified according to experience, without considering the impact of the mutual relationship between operators on the data arrangement method or memory access method setting. In fact, a neural network is often composed of various operators, and different operators have different requirements for the data arrangement method or memory access method. This setting method will result in the inability to obtain the optimal data arrangement method or memory access method for the operator in the entire computational graph.

[0060] In addition, during the flow of tensors between operators, conversion due to different arrangement methods will incur additional overheads, such as causing additional data movement. However, it is very likely that the conversion of the arrangement method of a certain operator is the optimal choice for the entire computational graph. The current operator arrangement method setting cannot automatically obtain whether and how to perform the arrangement method conversion, nor can it obtain the optimal operator arrangement method choice for the entire computational graph.

[0061] At least one embodiment of the present disclosure provides a data processing method, a data processing device, an electronic device, and a non-transitory computer-readable storage medium. The data processing method includes: obtaining a computational graph, where the computational graph includes a plurality of operators and the dependencies between the plurality of operators; determining at least one core operator based on the computational graph; determining the arrangement modes respectively corresponding to the plurality of operators through at least one round of iterative operations, where the arrangement mode corresponding to each operator includes the arrangement mode of the input tensor of the operator and the arrangement mode of the output tensor, and the arrangement mode is used to indicate the layout mode of the tensor in the storage component; where, in each round of iterative operations: selecting one or more core operators from the at least one core operator as starting operators; determining the arrangement mode corresponding to each starting operator according to the read-write overhead of each starting operator in a plurality of preset arrangement modes; starting from the starting operator, combining the arrangement mode corresponding to the starting operator, and propagating forward and backward according to the dependencies to sequentially obtain the arrangement modes corresponding to each operator, where, in the process of sequentially obtaining the arrangement modes corresponding to each operator, obtaining the arrangement mode corresponding to the current operator by combining the arrangement modes corresponding to adjacent operators, and the adjacent operator is an operator that is adjacent to the current operator according to the dependencies and whose arrangement mode has been determined.

[0062] In the data processing method provided by at least one embodiment of the present disclosure, first, starting operators are selected from the core operators, the arrangement mode corresponding to the starting operator is determined, and then the arrangement mode is propagated starting from the starting operator. Moreover, when determining the arrangement modes corresponding to each operator, the arrangement modes of adjacent operators are considered in combination with the dependencies between the operators provided by the computational graph, that is, when determining the arrangement mode of each operator, not only the operator itself is considered, but also the dependencies between the operators and the arrangement modes of adjacent operators are combined to determine the arrangement mode of the current operator. Therefore, the obtained arrangement modes of each operator are not only suitable for the operator itself, but also the arrangement modes that are optimal for the overall operation efficiency of the entire computational graph, effectively improving the data processing efficiency and system performance.

[0063] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0064] Figure 3 It is a schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure.

[0065] As Figure 3 shown, the data processing method provided by at least one embodiment of the present disclosure includes at least steps S10 - S30.

[0066] First, in step S10, a computational graph is obtained.

[0067] For example, the computational graph includes a plurality of operators and the dependencies between the plurality of operators.

[0068] A neural network can be regarded as a directed acyclic graph (DAG) composed of many computational nodes, and each node corresponds to an operator. Specifically, Figure 4A shows a schematic diagram of the directed acyclic graph of the neural network. As Figure 4A shown, a directed acyclic graph composed of 13 computational nodes is shown, that is, the corresponding neural network includes 13 operators. In addition, the nodes are connected by lines, and the lines represent the data dependency relationship and data flow direction between the computational nodes. For example, the output data of node 1 flows to node 2, and the output data of node 2 flows to node 3 and node 6, and so on. From this, it can be determined that there is a data dependency relationship between node 1 and node 2, and there is a data dependency relationship between node 2 and node 6 as well as node 3. It can be understood that, Figure 4A the network structure shown in

[0069] is only schematic, and the embodiments of the present disclosure do not specifically limit the structure of the neural network. Figure 4A Next, based on the directed acyclic graph shown in Figure 4B it is possible to perform operator fusion to obtain a fused computational graph. Figure 4A shows the implementation process of obtaining a computational graph by performing operator fusion according to the directed acyclic graph shown in

[0070] In the related art, for the calculation of a neural network, in order to more efficiently execute the calculation process of the operators in the network, operator fusion can be performed according to some rules and patterns, that is, multiple operators are fused into one fused operator. As Figure 4B shown, operator 1 and operator 2 can be fused into a fused operator A. Similarly, operator 3, operator 4, and operator 5 can be fused into a fused operator C, and operator 6, operator 7, operator 8, and operator 9 can be fused into a fused operator B, and so on. In addition, Figure 4B the computational graph in

[0071] also includes single operators that do not fuse with other operators. For example, operator D and operator E. According to the above rules of operator fusion, a computational graph composed of operators, fused operators, and the lines between them can be obtained. It can be understood that the embodiments of the present disclosure do not limit the specific rules for performing operator fusion, and conventional methods in the related art can be used.

[0072] Of course, in the present invention, the computational graph can be a computational graph after operator fusion. For example, Figure 4B the computational graph shown inFigure 4A The computational graph shown is not specifically limited in the present disclosure.

[0073] The data processing method of the present invention can be applied to a variety of application fields. According to different application fields, the computational graph can be used to implement different functions.

[0074] For example, the data processing method can be used in computer vision, and the computational graph can be a computational graph obtained from neural networks in fields such as image recognition, video analysis, image segmentation, object detection, face recognition, and medical image analysis.

[0075] For example, the data processing method can be used in natural language processing for computer vision, and the computational graph can be a computational graph obtained from neural networks in fields such as machine translation, sentiment analysis, speech recognition, and text generation.

[0076] For example, the data processing method can be used in speech recognition and synthesis, and the computational graph can be a computational graph obtained from neural networks in fields such as speech enhancement and speech recognition.

[0077] For example, the data processing method can also be applied to fields such as recommendation systems, medical diagnosis, finance, autonomous driving, and robotics. The computational graph can be a computational graph obtained from neural networks in the corresponding technical fields, which will not be elaborated one by one here. Step S20: Based on the computational graph, determine at least one core operator.

[0078] For example, in some embodiments, step S20 may include: selecting at least one core operator from multiple operators based on a predetermined condition; where the predetermined condition includes at least one of the following: a fused operator, an operator with a higher computational volume ranking when multiple operators are sorted in descending order of computational volume, an operator of a predetermined functional type, where the fused operator is an operator obtained by fusing multiple operators through an optimization strategy.

[0079] In the present disclosure, an operator with a relatively large computational volume and high computational complexity can be selected as the core operator. The arrangement method of the core operator has a more important impact on the overall operation efficiency of the entire computational graph. For example, the core operator can be selected according to the degree of influence of the arrangement method of the operator on the overall operation efficiency of the entire computational graph. For example, an operator with the largest computational volume can be selected as the core operator. For example, a fused operator can be selected as the core operator, and the number of core operators is not specifically limited.

[0080] For example, a fused operator can be selected as the core operator. The fused operator itself is obtained by fusing multiple operators and belongs to a complex operator with a relatively large computational volume or a relatively complex structure.

[0081] For example, operators with a relatively large amount of computation can be selected. For example, among multiple operators sorted in descending order of computational complexity, the operators ranked at the top in terms of computational complexity, such as the top 30% of the sorted operators, the top 50% of the sorted operators, etc. The present disclosure does not make specific limitations in this regard.

[0082] For example, operators of a predetermined functional type can be selected. For example, the predetermined functional type of operators can be other functional types of operators except for RELU (activation function) operators and POINTWISE (pointwise operation) operators.

[0083] Of course, those skilled in the art can also set other appropriate conditions to select the core operators according to actual experience. The present disclosure does not make specific limitations in this regard.

[0084] Step S30, determine the layout modes respectively corresponding to the multiple operators through at least one round of iterative operations.

[0085] For example, the layout mode corresponding to each operator includes the layout mode of the input tensor of the operator and the layout mode of the output tensor.

[0086] For example, the layout mode here includes at least one of a memory access mode and a data layout mode. For example, the data layout mode indicates the storage order and dimensional arrangement of data in a storage component, such as NCHW, NHWC, etc. listed in the present disclosure. The memory access mode is, for example, UMA or NUMA, etc.

[0087] For example, before the iterative operation, the read and write overheads of tensors (input tensors and output tensors) related to each operator in multiple preset layout modes can be obtained in advance, so that real-time calculation is not required during the iterative operation, avoiding repeated calculation and accelerating the iterative process. The read and write overheads include the time and resources consumed by reading data from memory to a computing unit and writing the calculation result back to memory. The layout mode of the operator is determined according to the read and write overheads. The read and write operations may become a performance bottleneck, and the read and write overheads directly indicate the computing efficiency. The layout mode with the optimal computing efficiency can be obtained by determining the layout mode according to the read and write overheads.

[0088] For example, the preset layout mode can be a combination of a memory access mode and a data layout mode. For example, one preset layout mode can be NCHW + UMA, and one preset layout mode can be NHWC + UMA, etc. Those skilled in the art can set according to needs. The present disclosure does not make specific limitations in this regard.

[0089] For example, the preset layout mode can be only a data layout mode or only a memory access mode. The present disclosure does not make specific limitations in this regard.

[0090] For example, before step S30, the data processing method provided by at least one embodiment of the present disclosure further includes: determining the shape and size of tensors related to multiple operators; for each operator, obtaining the read / write overheads of the related tensors under multiple preset arrangement manners according to the shape and size of the tensors related to the operator.

[0091] For example, determining the shape and size of tensors related to multiple operators may include: deriving the shape and size of the input tensors and output tensors of each operator according to the computational graph and the shape and size of the input activation tensors of the input computational graph.

[0092] For example, by traversing the computational graph and combining the inference function of each operator, the shape and size of the tensors flowing in the computational graph can be derived.

[0093] For example, for a matrix multiply-accumulate operator, which is usually expressed as [M, K] × [K, N] = [M, N], where M, N, and K are positive integers representing dimension sizes, the shape and size of the output tensor can be derived from the shape and size of the input tensors.

[0094] For example, for a unary pointwise operator, the shape and size of the output tensor are the same as those of the input tensor.

[0095] For example, for a binary pointwise operator, the shape and size of the output tensor are related to the broadcast operation of a certain dimension.

[0096] For example, for a reduction operator, the shape and size of the output tensor are obtained according to the dimension (reduction axis) specified during the reduction operation.

[0097] For example, for a convolution operator, the shape and size of the output tensor are obtained according to the convolution calculation formula.

[0098] For other types of operators, similar methods can be used for derivation, and the present disclosure will not list them one by one.

[0099] Thus, according to the computational graph and the shape and size of the input activation tensors of the input computational graph, each operator can be traversed to sequentially derive the shape and size of the input tensors and output tensors of each operator.

[0100] After obtaining the shape and size of the input tensors and output tensors of each operator, the read / write overheads of the input tensors and output tensors of each operator under multiple preset arrangement manners can be obtained.

[0101] For example, multiple preset arrangement methods may include multiple data arrangement methods, such as NCHW, NHWC, N(C / x)HWCx, BLOCK_COL_MAJOR, COL_MAJOR, etc. Multiple preset arrangement methods may also include memory access methods, such as UMA or NUMA, etc. For example, multiple preset arrangement methods may include combinations of multiple data arrangement methods and memory access methods, such as NCHW+UMA, NCHW+NUMA, etc.

[0102] For example, in some embodiments, according to the shape and size of the tensor related to the operator, obtaining the read / write overhead of the related tensor in multiple preset arrangement methods may include: determining the amount of data to be transmitted for the related tensor in each preset arrangement method according to the shape and size of the tensor related to the operator; and obtaining the read / write overhead of the related tensor in each preset arrangement method based on the amount of data to be transmitted and the hardware information, where the hardware information includes the data transmission bandwidth in each preset arrangement method.

[0103] For example, the following formula can be used to calculate the read / write overhead of each preset arrangement method of the operator:

[0104] Cost = M / B

[0105] Where M is the amount of data of the tensor in a certain preset arrangement method, B is the data transfer bandwidth in this preset arrangement method, and Cost represents the read / write overhead of the tensor in this preset arrangement method.

[0106] Even if the shape and size of the tensor are the same, the amount of data in different preset arrangement methods may be different. For example, the amount of data of BLOCK_COL_MAJOR is larger than that of COL_MAJOR because of padding processing due to alignment requirements.

[0107] The data transfer bandwidths of different preset arrangement methods are also different. For example, the data transfer bandwidth of BLOCK_COL_MAJOR is higher than that of COL_MAJOR.

[0108] For example, in some other embodiments, the read / write overhead of the operator can be modeled. The input of the model is information such as the shape and size of the operator, and the output of the model is the read / write overhead brought by data transfer. For example, machine learning, neural networks, etc. can be used for modeling. The present disclosure does not specifically limit the model and the modeling process. For example, complex operators, such as fused operators, can be modeled separately, and for other operators, the above calculation formula can be used to obtain the read / write overhead.

[0109] For example, in some embodiments, the data processing method provided by at least one embodiment of the present disclosure further includes: obtaining the read / write overhead when the relevant tensors undergo a layout conversion according to the shape and size of the tensors related to the operator, where the layout conversion includes storing the input tensors of the operator in a first preset layout in the storage component and loading the input tensors in a second preset layout when loading from the storage component, and the first preset layout and the second preset layout are different.

[0110] For example, the following formula can be used to calculate the read / write overhead when the operator undergoes a layout conversion:

[0111] Cost = M1 / B1 + M2 / B2

[0112] Where M1 is the data volume of the tensor in the first preset layout, B1 is the data transfer bandwidth in the first preset layout, M2 is the data volume of the tensor in the second preset layout, B2 is the data transfer bandwidth in the second preset layout, and Cost represents the read / write overhead when the tensor is converted from the first preset layout to the second preset layout (or from the second preset layout to the first preset layout).

[0113] For example, generally, the layout conversion will occur for the input tensors. For the input tensors, the read / write overheads when the layout is converted and not converted can be calculated, and for the output tensors, only the read / write overheads in each preset layout can be calculated to reduce the calculation amount.

[0114] Thus, through the above process, the read / write overheads of the input tensors and output tensors of each operator in different preset layouts and layout conversion modes can be obtained.

[0115] For example, in some other embodiments, after deriving the shape and size of each operator, before the start of a round of iterative operations, the read / write overheads of some operators can be calculated first, and then before the start of the next round of iterative operations, the read / write overheads of some operators can be calculated again. The present disclosure does not make specific limitations on this.

[0116] Figure 5 It is a schematic flowchart of the iterative operation provided by at least one embodiment of the present disclosure. As Figure 5 shown, the iterative operation may include steps S301 - S303.

[0117] In step S301, select one or more core operators from at least one core operator as the starting operator.

[0118] In step S302, determine the corresponding layout for each starting operator according to the read / write overheads of each starting operator in multiple preset layouts.

[0119] In step S303, starting from the starting operator, combined with the arrangement corresponding to the starting operator, propagate forward and backward according to the dependency relationship to obtain the arrangements corresponding to each operator in turn.

[0120] For example, in the process of obtaining the arrangements corresponding to each operator in turn, obtain the arrangement corresponding to the current operator by combining the arrangements corresponding to adjacent operators. The adjacent operator is an operator that is adjacent to the current operator according to the dependency relationship and whose arrangement has been determined.

[0121] Here, for forward propagation, the current operator is an operator that is adjacent and before the adjacent operator according to the dependency relationship. For example, the output tensor of the current operator is output to the adjacent operator as an input tensor. For backward propagation, the current operator is an operator that is adjacent and after the adjacent operator according to the dependency relationship. For example, the output tensor of the adjacent operator is output to the current operator as an input tensor. Taking Figure 4A as an example, assume the starting operator is operator 7. The forward propagation direction is operator 7 -> operator 6 -> operator 2..., and the backward propagation direction is operator 7 -> operator 9 -> operator 10...

[0122] Therefore, when determining the arrangements corresponding to each operator, consider the arrangements of adjacent operators in combination with the computational graph. That is, when determining the arrangement of each operator, not only consider the operator itself, but also combine the dependency relationship between operators and the arrangements of adjacent operators to determine the arrangement of the current operator. Thus, the arrangements obtained for each operator are not only suitable for the operator itself, but can obtain the optimal arrangement for each operator in the entire computational graph, improving data processing efficiency and system performance.

[0123] For example, stop when the propagation reaches the core operator, and end this round of iterative operation when both the forward propagation and backward propagation starting from at least one starting operator reach the core operator. For example, when there are multiple starting points, it can be that both the forward and backward propagations starting from each starting point reach the core operator, ending this round of iterative operation and starting the next round of iterative operation. For example, the propagation reaching the core operator here means that if the current operator is the core operator, it is considered that the propagation reaches the core operator and this round of iterative operation ends.

[0124] For example, the propagation here can be understood as propagating the arrangement corresponding to the starting operator to other operators. Since in the process of obtaining the arrangements corresponding to each operator in turn, the arrangement corresponding to the current operator is obtained by combining the arrangements corresponding to adjacent operators, so in fact each operator will combine and refer to the arrangement of the starting operator when determining its own arrangement, realizing the propagation of the arrangement corresponding to the starting operator.

[0125] For example, end the iterative operation in response to multiple operators having all determined their corresponding arrangements.

[0126] For example, the layout corresponding to each operator can be determined through one or more rounds of iterative operations. For example, if the layout settings for all operators have not been completed and the number of iterations does not exceed a preset threshold, the iterative operation continues to be executed. For example, if the number of executions of the iterative operation reaches the preset threshold and there is at least one operator in the computational graph for which the layout has not been determined, the layout of the at least one operator is determined based on a predetermined rule. For example, the predetermined rule can be a heuristic rule, which specifies the conditional requirements and the corresponding layout, and the operator that meets the corresponding conditional requirements uses the corresponding layout.

[0127] Some possible predetermined rules are listed below non - restrictively, and the present disclosure is not limited thereto.

[0128] For example, taking a graphics processing unit as an example, using the layout of blocks is more efficient.

[0129] For example, for a three - dimensional tensor, set it to the layout of BLOCK_COL_MAJOR + NUMA; for a linear operator, the second input tensor is the weight, and it needs to be set to UMA.

[0130] For a four - dimensional activation tensor, set it to the layout of BLOCK_NCHW (the data blocks are organized in the NCHW format) + NUMA, and for a four - dimensional convolution kernel tensor, set it to the layout of BLOCK_CONV_WEIGHT (the layout is block - convolution weight) + UMA.

[0131] The predetermined rule can be set according to the hardware, the shape and size of the tensor, the type of operator, etc. The predetermined rule changes continuously with the update of the hardware. Those skilled in the art can set it according to actual needs, and the present disclosure does not make specific limitations.

[0132] The specific process of each round of iterative operation is described in detail below.

[0133] For example, in some embodiments, step S301 may include: dividing at least one core operator into at least one priority level according to the importance; selecting one or more core operators belonging to the same priority level as the starting operator in each round of iterative operation; wherein, at least one round of iterative operation is executed in descending order of the importance of at least one priority level, the propagation stops when reaching the core operator, and this round of iterative operation ends when both the forward propagation and the backward propagation starting from at least one starting operator reach the core operator.

[0134] For example, the core operators can be divided into one or more priority levels according to the importance, such as according to the amount of computation, etc. In each round of iterative operation, all the core operators in a certain priority level are used as the starting operators for this round of iterative operation.

[0135] For example, the order of iterative operations can be determined according to the priority from high to low. For example, in the first iterative operation, one or more core operators with the highest priority are selected as the starting operators for this round of iterative operation. When the propagation reaches the core operators, this round of iterative operation ends. In the second iterative operation, one or more core operators with the second highest priority are selected as the starting operators for this round of iterative operation. When the propagation reaches the core operators, this round of iterative operation ends, and so on.

[0136] Thus, the arrangement of operators with a high degree of importance is determined first. For example, operators with a large amount of computation or complex structures, etc., which have a greater impact on the overall operation efficiency of the computational graph, can have their arrangements determined first. These arrangements of operators have a greater impact on the overall operation efficiency of the computational graph. Determining these arrangements first can more accurately obtain the optimal arrangements of each operator for the entire computational graph.

[0137] For example, in some other embodiments, it is also possible to set that there is only one priority for the core operators, such as in the case where the number of core operators is small. In this case, the corresponding arrangements of each operator can be obtained through one round of iterative operation.

[0138] For example, in some other embodiments, the core operators can also be divided into different priorities according to other conditions besides the degree of importance. The present disclosure does not make specific limitations on this.

[0139] For example, in some other embodiments, in each iteration operation, any core operator in any priority can be selected as the starting operator. The present disclosure does not make specific limitations on this.

[0140] For example, in some embodiments, step S302 may include: in response to the arrangement of the tensor corresponding to the target tensor not being determined, selecting the preset arrangement with the smallest read / write overhead from the read / write overheads of the target tensor under multiple preset arrangements as the arrangement of the target tensor, where loading the corresponding tensor as the target tensor, or loading the target tensor as the corresponding tensor; in response to the arrangement of the tensor corresponding to the target tensor having been determined, selecting the preset arrangement with the smallest read / write overhead from the multiple read / write overheads corresponding to the target tensor as the arrangement of the target tensor, where the multiple read / write overheads include the read / write overhead of the target tensor under the first preset arrangement and the read / write overhead when the target tensor is converted between multiple arrangements. The first preset arrangement is the arrangement determined for the corresponding tensor, and the conversion between multiple arrangements includes the conversion of the first preset arrangement to multiple other preset arrangements.

[0141] For example, when the iterative operation is performed for the first time, the arrangements of all operators have not been determined, and the input tensors and output tensors of the starting operators can be determined in the same way.

[0142] For example, taking the input tensor as an example, among the read-write overheads of the input tensor of the starting operator under multiple preset arrangement modes, the preset arrangement mode with the smallest read-write overhead can be selected as the arrangement mode of the input tensor of the starting operator. The read-write overheads of the input tensor under multiple preset arrangement modes can be determined with reference to the foregoing content and will not be elaborated here.

[0143] For example, when performing iterative operations for the second time and subsequent times, for the target tensor in the starting operator whose arrangement mode has not been determined yet, for example, the target tensor can be the input tensor or output tensor in the starting operator whose arrangement mode has not been determined yet, it is also necessary to consider the arrangement modes of other input tensors or output tensors whose arrangement modes have been determined. For example, it is also necessary to consider the read-write overhead brought by the arrangement mode conversion.

[0144] For example, if the corresponding tensor of the target tensor has not had its arrangement mode determined, among the read-write overheads of the target tensor under multiple preset arrangement modes, the preset arrangement mode with the smallest read-write overhead is selected as the arrangement mode of the target tensor. For example, assuming that the target tensor is the input tensor of Operator 1, the output tensor 1 of Operator 2 is output to Operator 1 as this target tensor, and at this time the corresponding tensor is output tensor 1. For example, assuming that the target tensor is the output tensor of Operator 1, the output tensor of Operator 1 is output to Operator 3 as input tensor 1, and at this time the corresponding tensor is input tensor 1.

[0145] For example, if the corresponding tensor of the target tensor has had its arrangement mode determined, among the multiple read-write overheads of the target tensor, the preset arrangement mode with the smallest read-write overhead is selected as the arrangement mode of the target tensor. Here, the multiple read-write overheads include the read-write overhead of the target tensor under the first preset arrangement mode and the read-write overhead when the target tensor is converted among multiple arrangement modes. The first preset arrangement mode is the arrangement mode determined for the corresponding tensor, and the conversion among multiple arrangement modes includes the conversion of the first preset arrangement mode to multiple other preset arrangement modes.

[0146] For example, assuming that the read-write overhead under the first preset arrangement mode is the smallest, it is determined that the target tensor uses the first preset arrangement mode. For example, assuming that the read-write overhead when the first preset arrangement mode is converted to the second preset arrangement mode is the smallest, it is determined that the target tensor uses the second preset arrangement mode.

[0147] Thus, in the present disclosure, for core operators with complex performance, such as matrix multiplication operators, etc., it is necessary to comprehensively consider the combination of arrangement modes of multiple input and output tensors to obtain a more accurate arrangement mode.

[0148] For example, in some embodiments, some preset arrangement modes with inevitably large read-write overheads can be excluded according to the operator characteristics of the starting operator, thereby reducing the comparison range and accelerating the execution of the iterative operation.

[0149] For example, to select the preset arrangement method with the smallest read / write overhead from the read / write overheads of the target tensor under multiple preset arrangement methods as the arrangement method of the target tensor may include: excluding at least one preset arrangement method from multiple preset arrangement methods according to the operator characteristics of the starting operator; and selecting the preset arrangement method with the smallest read / write overhead from the read / write overheads of the starting operator under the remaining preset arrangement methods as the arrangement method of the target tensor.

[0150] For example, for the convolution kernel tensor, the memory access mode adopts UMA. For example, for the tensor of the matrix multiplication operator, one of the tensors must be UMA. For example, for the reduction type operator, if the reduction operation is performed between multiple computing cores, the output tensor must adopt UMA. Then, selection can be made from the remaining preset arrangement methods. For example, taking the convolution kernel tensor as an example, the case where the memory access mode is UMA does not need to be considered. Thus, the number of read / write overheads to be compared can be reduced, the process of the iterative operation can be accelerated, and the efficiency can be improved.

[0151] For example, in some embodiments, step S303 may include: combining the dependency relationship to perform topological sorting on the computation graph according to a preset topological sorting type; starting from the starting operator, sequentially performing propagation operations on each operator forward and backward according to the result of the topological sorting; where the propagation operation includes: determining the current operator, where when propagating backward from the starting operator, the current operator is the operator that uses the output tensor of the adjacent operator as the input tensor of the current operator, and when propagating forward from the starting operator, the current operator is the operator that outputs the output tensor to the adjacent operator as the input tensor of the adjacent operator; in response to the current operator not being the core operator, determining the arrangement method corresponding to the current operator based on the arrangement method corresponding to the adjacent operator; and in response to the current operator being the core operator, stopping the current round of iterative operation.

[0152] For example, the preset topological sorting type may be breadth - first traversal. For example, other feasible topological sorting methods may also be adopted.

[0153] For example, determining the layout mode corresponding to the current operator based on the layout mode corresponding to the adjacent operator may include: in response to using the output tensor of the adjacent operator as the input tensor of the current operator: selecting, from the first read / write overhead corresponding to the input tensor of the current operator and the second read / write overhead corresponding to the input tensor of the current operator, the preset layout mode with the smallest read / write overhead as the layout mode of the input tensor of the current operator, where the first read / write overhead is the read / write overhead in the first preset layout mode, and it is determined that the output tensor of the adjacent operator uses the first preset layout mode, and the second read / write overhead is the read / write overhead when converting the first preset layout mode to at least one other preset layout mode; selecting, from the read / write overheads of the output tensor of the current operator in multiple preset layout modes, the preset layout mode with the smallest read / write overhead as the layout mode of the output tensor of the current operator.

[0154] For example, determining the layout mode corresponding to the current operator based on the layout mode corresponding to the adjacent operator may include: in response to using the output tensor of the current operator as the input tensor of the adjacent operator: selecting, from the third read / write overhead corresponding to the output tensor of the current operator and the fourth read / write overhead corresponding to the output tensor of the current operator, the preset layout mode with the smallest read / write overhead as the layout mode of the output tensor of the current operator, where the third read / write overhead is the read / write overhead in the second preset layout mode, and it is determined that the input tensor of the adjacent operator uses the second preset layout mode, and the fourth read / write overhead is the read / write overhead when converting at least one other preset layout mode to the second preset layout mode; selecting, from the read / write overheads of the input tensor of the current operator in multiple preset layout modes, the preset layout mode with the smallest read / write overhead as the layout mode of the input tensor of the current operator.

[0155] For example, the propagation process starts from the core operator and ends when the propagation reaches the core operator. For example, starting from the starting operator, the propagation operation is sequentially performed on each operator forward and backward according to the result of the topological sorting.

[0156] Figure 6 This is a schematic diagram of the sorting result after the preset topological sorting shown in an embodiment of the present disclosure. It should be noted that Figure 6 only three operators are shown, but actually the sorting result may also include more operators, which will not be repeated here.

[0157] For example, assume that operator OP2 is the starting operator, operator OP3 uses the output tensor t3 of operator OP2 as the input tensor t4, and the output tensor t1 of operator OP1 is output to operator OP2 as the input tensor t2. Here, tensor t1 and tensor t2 are actually the same tensor, and tensor t3 and tensor t4 are actually the same tensor. However, it is possible that operator OP1 stores the output tensor t1 in the storage component in a certain data layout format, and operator OP2 loads the input tensor t2 from the storage component in a certain data layout format.

[0158] In step S302, the layout methods of the input tensor t2 and the output tensor t3 of operator OP2 can be obtained. The specific process can refer to the description of step S302 and will not be elaborated here.

[0159] In step S303, when propagating backward starting from operator OP2, it is determined that the current operator is operator OP3.

[0160] Assume that operator OP3 is the core operator, then there is no need to propagate backward, and this round of iterative operation can be stopped. The specific stopping condition is that a core operator is also encountered during forward propagation.

[0161] Assume that operator OP3 is not the core operator, for example, it is a RELU - like operator. Based on the layout method corresponding to the starting operator OP2, determine the layout method corresponding to operator OP3.

[0162] Specifically, in step S302, the layout method used by the output tensor t3 of operator OP2 has been determined. For example, it is the preset layout method 1. From the first read - write cost cost1 corresponding to the input tensor t4 of operator OP3 and the second read - write cost cost2 corresponding to the input tensor t4 of operator OP3, select the preset layout method with the smallest read - write cost as the preset layout method of the input tensor t4 of operator OP3. Here, the first read - write cost cost1 is the read - write cost when the input tensor t4 of operator OP3 uses the preset layout method 1, and the second read - write cost is the read - write cost when converting the preset layout method 1 to other preset layout methods.

[0163] For example, assume that the first read - write cost cost1 is the smallest, then it is determined that the input tensor t4 of operator OP3 uses the preset layout method 1. Assume that the read - write cost of converting the preset layout method 1 to the preset layout method 2 is the smallest, then it is determined that the input tensor t4 of operator OP3 uses the preset layout method 2.

[0164] For example, for the output tensor t5 of operator OP3, from the read - write costs of the output tensor t5 of operator OP3 under multiple preset layout methods, select the preset layout method with the smallest read - write cost as the layout method of the output tensor t5 of operator OP3.

[0165] In step S303, when propagating forward starting from operator OP2, it is determined that the current operator is operator OP1.

[0166] Assume that operator OP1 is a core operator, then no further forward propagation is performed, and this round of iterative operation can be stopped. The specific stopping condition is that a core operator is also encountered during backward propagation.

[0167] Assume that operator OP1 is not a core operator, for example, it is a POINTWISE type operator. Determine the layout mode corresponding to operator OP1 based on the layout mode corresponding to the starting operator OP2.

[0168] Specifically, in step S302, the layout mode used by the input tensor t2 of operator OP2 has been determined, for example, it is the preset layout mode 3. From the third read / write cost cost3 corresponding to the output tensor t1 of operator OP1 and the fourth read / write cost cost4 corresponding to the output tensor t1 of operator OP1, select the preset layout mode with the smallest read / write cost as the layout mode of the output tensor t1 of operator OP1. Here, the third read / write cost cost3 is the read / write cost when the output tensor t1 of operator OP1 uses the preset layout mode 3, and the second read / write cost is the read / write cost for converting other preset layout modes to the preset layout mode 3.

[0169] For example, assume that the third read / write cost cost3 is the smallest, then determine that the output tensor t1 of operator OP1 uses the preset layout mode 3. Assume that the read / write cost for converting the preset layout mode 4 to the preset layout mode 3 is the smallest, then determine that the output tensor t1 of operator OP1 uses the preset layout mode 4.

[0170] For example, for the input tensor t0 of operator OP1, the preset layout mode with the smallest read / write cost can be selected from the read / write costs of the input tensor t0 of operator OP1 under multiple preset layout modes as the layout mode of the input tensor t0 of operator OP1. Of course, if there is another operator OP0 before operator OP1, and the output tensor t' of operator OP0 is output to operator OP1 as the input tensor t0, if the layout mode of the output tensor t' of operator OP1 has been determined, then a similar process as above can be referred to, and the layout mode of the input tensor t0 of operator OP1 can be determined considering the read / write cost of layout mode conversion, which will not be elaborated here.

[0171] After that, operator OP1 can continue to be propagated forward as an adjacent operator to determine the current operator, and operator OP3 can be propagated backward as an adjacent operator to determine the current operator. The specific process will not be elaborated here.

[0172] Therefore, when determining the arrangement of each operator, the arrangement corresponding to adjacent operators is considered, and the arrangement of the core operators that has a greater impact on the overall efficiency of the computational graph and is determined in advance is propagated through the above process. When determining the arrangement of each operator, the arrangement of the core operators and the dependency relationship between the operators are combined, and the optimal operator arrangement for the entire computational graph can be obtained.

[0173] Figure 7 It is a schematic diagram of the execution process of the data processing method provided by at least one embodiment of the present disclosure.

[0174] As Figure 7 shown, first, a computational graph is obtained. The relevant process of obtaining the computational graph can refer to the relevant description of step S10 above, and will not be elaborated here.

[0175] After obtaining the computational graph, according to the computational graph and the shape and size of the input activation tensor input to the computational graph, the shape and size of each operator are deduced one by one. The specific process can refer to the foregoing content and will not be elaborated here.

[0176] After that, based on the computational graph, core operators are determined. The specific process can refer to the relevant description of step S20 above and will not be elaborated here.

[0177] After that, at least one round of iterative operations is started, as Figure 7 shown by the dashed box part in

[0178] In each round of iterative operation, first, the read and write overheads of some operators are determined. The specific process can refer to the relevant content above and will not be elaborated here.

[0179] After that, a starting operator is selected. For example, the starting operator in the first round of iterative operation is the core operator with the highest priority, and the starting operator in the second round of iterative operation is the core operator with the second highest priority, and so on. The specific content can refer to the relevant description in step S301 above and will not be elaborated here.

[0180] After that, according to the read and write overheads of the starting operator in multiple preset arrangement ways, the arrangement way corresponding to the starting operator is determined. The specific process can refer to the relevant description in step S302 above and will not be elaborated here.

[0181] After that, starting from the starting operator, the arrangement way is propagated. Combining the arrangement way corresponding to the starting operator, it is propagated forward and backward according to the dependency relationship, and the arrangement ways corresponding to each operator are obtained in turn. The specific process can refer to the relevant description in step S303 above and will not be elaborated here.

[0182] After that, if both the forward and backward propagations reach the core operator, this round of iterative operation ends. If the number of iterative operations does not exceed the preset threshold, continue with the next round of iterative operation, and the specific process is as described above. If the number of iterative operations exceeds the preset threshold, determine whether the arrangement methods of all tensors have been determined. If so, end; if not, set the arrangement methods of the remaining tensors according to the predetermined rules, and then end.

[0183] In this data processing method, first, the core operator and its arrangement method are determined, and the core operator also has different priorities. Therefore, when determining the arrangement methods of each operator, the arrangement methods of some operators can be set according to the importance level first, and the arrangement methods of other operators are determined in sequence according to these arrangement methods. Moreover, when determining the arrangement methods of each operator, the corresponding arrangement methods of adjacent operators are considered, and the arrangement method of the core operator that has a greater impact on the overall efficiency of the computational graph and has been determined in advance is propagated through the above process. When determining the arrangement methods of each operator, the arrangement method of the core operator and the dependency relationship between operators are combined, and the optimal operator arrangement method for the entire computational graph can be obtained.

[0184] Figure 8 It is a schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure.

[0185] As Figure 8 shown, the data processing device 100 includes an acquisition unit 101, a core operator determination unit 102, and an iterative operation unit 103.

[0186] The acquisition unit 101 is configured to acquire a computational graph, where the computational graph includes a plurality of operators and the dependency relationships between the plurality of operators.

[0187] The core operator determination unit 102 is configured to determine at least one core operator based on the computational graph.

[0188] The iterative operation unit 103 is configured to determine the corresponding arrangement methods of the plurality of operators through at least one round of iterative operation, where the arrangement method corresponding to each operator includes the arrangement method of the input tensor and the output tensor of the operator, and the arrangement method includes at least one of a memory access method and a data arrangement method, and the data arrangement method indicates the storage order and dimension arrangement of data in the storage component.

[0189] For example, in each round of iterative operation: one or more core operators are selected from at least one core operator as the starting operator; according to the read / write overheads of each starting operator under multiple preset arrangement modes, the arrangement mode corresponding to each starting operator is determined; starting from the starting operator, in combination with the arrangement mode corresponding to the starting operator, and propagating forward and backward according to the dependency relationship, the arrangement modes corresponding to each operator are obtained in sequence. Among them, in the process of obtaining the arrangement modes corresponding to each operator in sequence, the arrangement mode corresponding to the current operator is obtained by combining the arrangement modes corresponding to adjacent operators, and the adjacent operator is an operator that is adjacent to the current operator according to the dependency relationship and whose arrangement mode has been determined.

[0190] For example, the acquisition unit 101, the core operator determination unit 102, and the iterative operation unit 103 include codes and programs stored in a memory. The acquisition unit 101, the core operator determination unit 102, and the iterative operation unit 103 are implemented as a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, for example. The processing unit can be a general-purpose processor and can also be a single-chip microcomputer, a microprocessor, a digital signal processor, a dedicated image processing chip, or a field programmable logic array, etc. The acquisition unit 101, the core operator determination unit 102, and the iterative operation unit 103 execute the codes and programs to implement some or all of the functions of the acquisition unit 101, the core operator determination unit 102, and the iterative operation unit 103 as described above. For example, the acquisition unit 101, the core operator determination unit 102, and the iterative operation unit 103 can be a circuit board or a combination of multiple circuit boards for implementing the functions as described above. In the embodiments of the present application, the combination of the one circuit board or multiple circuit boards may include: (1) one or more processors; (2) one or more non-transitory memories connected to the processors; and (3) firmware stored in the memory that can be executed by the processors.

[0191] It should be noted that the acquisition unit 101 can be used to implement Figure 3 the step S10 shown; the core operator determination unit 102 can be used to implement Figure 3 the step S20 shown; the iterative operation unit 103 can be used to implement Figure 3Step S30 shown. Therefore, for the specific description of the functions that can be implemented by the acquisition unit 101, reference can be made to the relevant description of step S10 in the embodiment of the above-mentioned data processing method. For the specific description of the functions that can be implemented by the core operator determination unit 102, reference can be made to the relevant description of step S20 in the embodiment of the above-mentioned data processing method. For the specific description of the functions that can be implemented by the iteration operation unit 103, reference can be made to the relevant description of step S30 in the embodiment of the above-mentioned data processing method. The repeated parts will not be repeated here. In addition, the data processing device 100 can achieve technical effects similar to those of the aforementioned data processing method, which will not be repeated here.

[0192] It should be noted that in at least one embodiment of the present disclosure, the data processing device 100 may include more or fewer circuits or units, and the connection relationship between the various circuits or units is not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or can be composed in other applicable ways.

[0193] For example, the data processing device 100 may be implemented in hardware, software, or a combination of hardware and software, and the present disclosure does not impose any specific limitations on this.

[0194] The data processing method provided by at least one embodiment of the present disclosure can achieve a technical effect similar to the data processing method described above, and will not be described in detail here.

[0195] The data processing method and data processing device provided in at least one embodiment of the present disclosure can be applied to different systems or devices, such as Figure 9 The electronic device 300 shown. The electronic device 300 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an AR device, a VR device, a vehicle terminal, etc., or a server, etc. For example, the hardware architecture of the electronic device 300 may include a graphics processor. The data processing method provided in at least one embodiment of the present disclosure can be applied to scenarios involving CPU, high performance computing (HPC) and artificial intelligence (AI) in the electronic device 300. Of course, the present disclosure is not limited to this, and any scenario, device, apparatus, etc. involving calculation graphs and model operations can adopt the data processing method or data processing apparatus provided in at least one embodiment of the present disclosure.

[0196] In some embodiments, the data processing device provided by at least one embodiment of the present disclosure may be a chip. For example, the chip is a System-on-a-Chip (SoC). The system-on-a-chip includes a processor, which may be a single-core processor or a multi-core processor, a memory, and an I / O interface, etc.

[0197] Figure 9 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 9 shown, the electronic device 300 is, for example, suitable for implementing the data processing method provided by the embodiments of the present disclosure. It should be noted that Figure 9 the components of the electronic device 300 shown are merely exemplary and not restrictive. According to actual application requirements, the electronic device 300 may further have other components.

[0198] As Figure 9 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.

[0199] For example, when the computer-readable instructions are run by the processing device 301, one or more steps in the data processing method according to any of the above embodiments may be executed. It should be noted that a detailed description of the processing process of the data processing method may refer to the relevant descriptions in the embodiments of the above data processing method.

[0200] For example, the memory may include any combination of one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include a random access memory (RAM) 303 and / or a cache memory, etc. For example, the computer-readable instructions may be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. The non-volatile memory may, for example, include a read-only memory (ROM) 302, a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage media, such as style images, and various data used and / or generated by the application programs, etc.

[0201] For example, the processing device 301, the read-only memory (ROM) 302, and the random access memory (RAM) 303 are connected to each other via a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0202] Typically, the following devices can be connected to the input / output (I / O) interface 305: input devices 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 can alternatively implement or have more or fewer devices. For example, the processing device 301 can control other components in the electronic device 300 to perform desired functions. The processing device 301 can be a device with data processing capabilities and / or program execution capabilities such as a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit (GPU). The central processing unit (CPU) can be of X86, ARM, RISC-V architecture, etc. The GPU can be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.

[0203] Figure 10 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 10 shown, the storage medium 400 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 401 can be non-temporarily stored on the storage medium 400. For example, when the computer-readable instructions 401 are executed by a processor, one or more steps of the data processing method described above can be executed.

[0204] For example, the storage medium 400 can be applied to the electronic device 300. For example, the storage medium 400 can include the storage device 308 in the electronic device 300.

[0205] For example, a storage device may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage media, and the processor may run the computer-readable instructions to implement various functions of the processor. Various application programs and various data may also be stored in the storage media.

[0206] For example, the storage medium may include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and may also be other applicable storage media.

[0207] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0208] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.

[0209] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0210] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0211] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0212] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.

[0213] For the present disclosure, the following points also need to be noted:

[0214] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0215] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0216] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A data processing method, comprising: Obtaining a computation graph, wherein the computation graph includes a plurality of operators and dependencies between the plurality of operators; Based on the computation graph, determining at least one core operator, wherein when determining the at least one core operator, a selection is made based on the degree of influence of the arrangement of the operators on the overall operating efficiency of the entire computation graph; Determine, through at least one round of iterative operations, the arrangement modes corresponding to the multiple operators, respectively, wherein the arrangement mode corresponding to each operator includes the arrangement mode of the input tensor of the operator and the arrangement mode of the output tensor of the operator, and the arrangement mode is used to indicate the layout mode of the tensor in the storage component; Among them, in each round of iteration operation: Selecting one or more core operators from the at least one core operator as a starting operator; Determine the arrangement mode corresponding to each starting operator according to the read and write overhead of each starting operator under multiple preset arrangement modes, wherein the multiple preset arrangement modes include multiple combinations of memory access modes and data arrangement modes; Taking the starting operator as the starting point, combined with the arrangement mode corresponding to the starting operator, propagate forward and backward according to the dependency relationship, and obtain the arrangement mode corresponding to each operator in turn, wherein, in the process of obtaining the arrangement mode corresponding to each operator in turn, the arrangement mode corresponding to the current operator is obtained in combination with the arrangement mode corresponding to the adjacent operator, and the adjacent operator is an operator that is adjacent to the current operator according to the dependency relationship and has a determined arrangement mode.

2. The data processing method according to claim 1, wherein: Based on the computation graph, determining at least one core operator includes: Selecting the at least one core operator from the plurality of operators based on a predetermined condition; The predetermined condition includes at least one of the following: a fusion operator, an operator with a higher computational amount when the plurality of operators are sorted from high to low in terms of computational amount, an operator of a predetermined functional type, The fusion operator is an operator obtained by fusing multiple operators through an optimization strategy.

3. The data processing method according to claim 1, wherein: Selecting one or more core operators from the at least one core operator as a starting operator comprises: Dividing the at least one core operator into at least one priority level according to the importance; In each round of iterative operation, one or more core operators of the same priority are selected as the starting operators; Among them, the at least one round of iterative operation is executed in descending order according to the importance of the at least one priority level, the forward and backward propagation stops when reaching the core operator, and the current round of iterative operation ends when the forward and backward propagation starting from the starting operator both reach the core operator.

4. The data processing method according to claim 1, wherein: According to the read and write overhead of each starting operator under multiple preset arrangements, the arrangement method corresponding to each starting operator is determined, including: For the target tensor whose arrangement has not been determined in the starting operator: When performing the first iteration operation, selecting a preset arrangement mode with the smallest read / write overhead from the read / write overhead of the target tensor under the multiple preset arrangement modes as the arrangement mode of the target tensor, wherein the target tensor includes an input tensor or an output tensor; When executing non-first iteration operations, In response to an undetermined arrangement of the tensor corresponding to the target tensor, a preset arrangement with the smallest read / write overhead is selected from the read / write overhead of the target tensor under the multiple preset arrangement modes as the arrangement mode of the target tensor, wherein in response to the target tensor being the input tensor, the corresponding tensor is loaded as the target tensor, or in response to the target tensor being the output tensor, the target tensor is loaded as the corresponding tensor; In response to the determined arrangement of the tensor corresponding to the target tensor, a preset arrangement with the smallest read / write overhead is selected as the arrangement of the target tensor from the multiple read / write overheads corresponding to the target tensor, wherein the multiple read / write overheads corresponding to the target tensor include the read / write overhead of the target tensor under a first preset arrangement and the read / write overhead of the target tensor when converted to multiple arrangement modes, the first preset arrangement being the arrangement determined to be used by the corresponding tensor, and the multiple arrangement conversions including the conversion of the first preset arrangement into multiple other preset arrangement modes.

5. The data processing method according to claim 4, wherein: From the read and write costs of the target tensor under the multiple preset arrangements, selecting a preset arrangement with the smallest read and write cost as the arrangement of the target tensor includes: Excluding at least one preset arrangement from the plurality of preset arrangements according to the operator characteristics of the starting operator; From the read and write costs of the starting operator under the remaining preset arrangements, select the preset arrangement with the smallest read and write cost as the arrangement of the target tensor.

6. The data processing method according to claim 1, wherein: Taking the starting operator as the starting point, combined with the arrangement mode corresponding to the starting operator, propagating forward and backward according to the dependency relationship, the arrangement modes corresponding to each operator are obtained in turn, including: In combination with the dependency relationship, topologically sorting the computation graph according to a preset topological sorting type; Taking the starting operator as the starting point, performing propagation operations on each operator forward and backward in sequence according to the result of the topological sorting; The propagation operation includes: Determine the current operator, wherein, when propagating backward from the starting operator, the current operator is an operator that uses the output tensor of the adjacent operator as the input tensor of the current operator, and when propagating forward from the starting operator, the current operator is an operator that outputs the output tensor to the adjacent operator as the input tensor of the adjacent operator; In response to the current operator not being the core operator, determining an arrangement mode corresponding to the current operator based on arrangement modes corresponding to the adjacent operators; In response to the current operator being the core operator, stopping the current round of iteration operation.

7. The data processing method according to claim 6, wherein: Determining the arrangement mode corresponding to the current operator based on the arrangement mode corresponding to the adjacent operators includes: In response to the output tensor of the adjacent operator being output to the current operator as the input tensor of the current operator: From the first read-write cost corresponding to the input tensor of the current operator and the second read-write cost corresponding to the input tensor of the current operator, select the preset arrangement mode with the smallest read-write cost as the arrangement mode of the input tensor of the current operator, wherein the first read-write cost is the read-write cost under the first preset arrangement mode, the output tensor of the adjacent operator is determined to use the first preset arrangement mode, and the second read-write cost is the read-write cost when converting the first preset arrangement mode to at least one other preset arrangement mode; From the read and write costs of the output tensor of the current operator under the multiple preset arrangements, select a preset arrangement with the smallest read and write cost as the arrangement of the output tensor of the current operator.

8. The data processing method according to claim 6, wherein: Determining the arrangement mode corresponding to the current operator based on the arrangement mode corresponding to the adjacent operators includes: In response to outputting the output tensor of the current operator to the adjacent operator as the input tensor of the adjacent operator: From the third read-write cost corresponding to the output tensor of the current operator and the fourth read-write cost corresponding to the output tensor of the current operator, select the preset arrangement mode with the smallest read-write cost as the arrangement mode of the output tensor of the current operator, wherein the third read-write cost is the read-write cost under the second preset arrangement mode, the input tensor of the adjacent operator is determined to use the second preset arrangement mode, and the fourth read-write cost is the read-write cost when converting at least one other preset arrangement mode to the second preset arrangement mode; From the read and write costs of the input tensor of the current operator under the multiple preset arrangements, select a preset arrangement with the smallest read and write cost as the arrangement of the input tensor of the current operator.

9. The data processing method according to claim 1, wherein: In response to the number of executions of the iterative operation reaching a preset threshold and there being at least one operator in the computation graph whose arrangement has not yet been determined, determining the arrangement of the at least one operator based on a predetermined rule, The predetermined rule specifies the condition requirements and the arrangement methods corresponding to the condition requirements, and the operators that meet the corresponding condition requirements use the arrangement methods corresponding to the corresponding condition requirements.

10. The data processing method according to any one of claims 1 to 9, wherein: Before determining the arrangement modes respectively corresponding to the multiple operators through at least one round of iterative operation, the data processing method further includes: Determine the shape and size of the tensors associated with the multiple operators; For each operator, according to the shape and size of the tensor associated with the operator, the read and write overhead of the tensor associated with the operator in the plurality of preset arrangements is obtained.

11. The data processing method according to claim 10, wherein: The operator-related tensors include the operator's input tensor and the operator's output tensor. Determining the shape and size of the tensors associated with the multiple operators includes: The shape size of the input tensor and the output tensor of each operator is derived according to the computation graph and the shape size of the input activation tensor input into the computation graph.

12. The data processing method according to claim 10, wherein: According to the shape and size of the tensor associated with the operator, the read and write overhead of the tensor associated with the operator in the plurality of preset arrangements is obtained, including: Determine, according to the shape and size of the tensor associated with the operator, the amount of data that needs to be transmitted for the tensor associated with the operator in each preset arrangement mode; Based on the amount of data to be transmitted and the hardware information, the read and write overhead of the operator-related tensors in each preset arrangement is obtained, wherein the hardware information includes the data transmission bandwidth in each preset arrangement.

13. The data processing method according to claim 10, wherein: Before determining the arrangement modes respectively corresponding to the multiple operators through at least one round of iterative operation, the data processing method further includes: According to the shape and size of the tensor associated with the operator, the read and write overhead of the tensor associated with the operator when an arrangement conversion occurs is obtained, wherein the arrangement conversion includes storing the input tensor of the operator in a first preset arrangement in the storage component, and loading the input tensor from the storage component in a second preset arrangement, and the first preset arrangement is different from the second preset arrangement.

14. The data processing method according to any one of claims 1 to 9, wherein: The arrangement method includes at least one of a memory access method and a data arrangement method. The data arrangement method indicates the storage order and dimensional arrangement of data in the storage component, and the memory access method indicates the memory organization method and the rules for accessing the memory.

15. A data processing device, comprising: An acquisition unit configured to acquire a computation graph, wherein the computation graph includes a plurality of operators and dependencies between the plurality of operators; A core operator determination unit is configured to determine at least one core operator based on the computation graph, wherein when determining the at least one core operator, a selection is made based on the degree of influence of the arrangement of the operators on the overall operating efficiency of the entire computation graph; an iterative operation unit, configured to determine, through at least one round of iterative operation, the arrangement modes corresponding to the plurality of operators, respectively, wherein the arrangement mode corresponding to each operator includes an arrangement mode of the input tensor of the operator and an arrangement mode of the output tensor of the operator, and the arrangement mode is used to indicate a layout mode of the tensor in the storage component; Among them, in each round of iteration operation: Selecting one or more core operators from the at least one core operator as a starting operator; Determine the arrangement mode corresponding to each starting operator according to the read and write overhead of each starting operator under multiple preset arrangement modes, wherein the multiple preset arrangement modes include multiple combinations of memory access modes and data arrangement modes; Taking the starting operator as the starting point, combined with the arrangement mode corresponding to the starting operator, propagate forward and backward according to the dependency relationship, and obtain the arrangement mode corresponding to each operator in turn, wherein, in the process of obtaining the arrangement mode corresponding to each operator in turn, the arrangement mode corresponding to the current operator is obtained in combination with the arrangement mode corresponding to the adjacent operator, and the adjacent operator is an operator that is adjacent to the current operator according to the dependency relationship and has a determined arrangement mode.

16. An electronic device, comprising: A memory non-transitorily stores computer executable instructions; a processor configured to execute the computer executable instructions, Wherein, when the computer executable instructions are executed by the processor, the data processing method according to any one of claims 1-14 is implemented.

17. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the data processing method according to any one of claims 1 to 14 is implemented.

Citation Information

Patent Citations

  • Model conversion method, model conversion equipment and storage medium

    CN114896950A

  • Data layout optimization method and device, electronic equipment and storage medium

    CN116909571A