Dimension conversion chip, method, device, storage medium and program product
Through the dimensional conversion chip and method, the tensor arrangement format is adjusted according to the differences in the neural network layer, which solves the problem that AI hardware cannot flexibly adapt to different network layers, and improves the data processing efficiency and accuracy of the neural network model.
Patent Information
- Application Number
- PCT/CN2024/138007
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2024-12-10
- Publication Date
- 2025-07-31
AI Technical Summary
Existing AI hardware only supports dimension transposition of specific dimensions in neural network models, and cannot flexibly adapt to changes in different network layers, resulting in poor operation of neural network models.
Provide a dimension conversion chip and method, through information acquisition circuits, parameter determination circuits and dimension conversion circuits, flexibly set dimension conversion instructions and adjust the arrangement format of tensors according to the differences in the number and order of dimensions of the neural network layer.
It realizes flexible adjustment of tensor layout format in neural network model, improves data processing efficiency and accuracy, and avoids pre-set limitations.
Smart Images

Figure CN2024138007_31072025_PF_FP_ABST
Abstract
Description
Dimension conversion chip, method, device, storage medium and program product
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 22, 2024, with application number 202410086998.5, and invention name “Dimensional transposition chip, method, device, storage medium and program product”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a dimension conversion chip, method, device, storage medium, and program product. Background Art
[0003] Neural networks are computational models that mimic the structure and function of biological neural networks and are widely used in various fields. Neural networks are typically stored as tensors, and the layout of tensors can vary between different layers of a neural network model.
[0004] Dimension transposition is very important in neural network operations. It is mainly used to meet the requirements of different neural network layers for tensor arrangement formats. Usually, AI hardware instructions support the dimension transposition process in a specific dimension (such as four or three dimensions). After obtaining the tensor, the tensor is transposed using pre-set transposition parameters to achieve the dimension transposition of the tensor in that dimension.
[0005] Technical content
[0006] The present invention provides a dimensionality conversion chip, method, device, storage medium, and program product that can flexibly determine the parameters of dimensionality conversion instructions based on the dimensionality conversion requirements of different neural network layers during data processing, and based on the differences in the number of dimensions and / or dimensional order between the source and target layout formats, thereby enabling targeted adjustments to the layout format of tensors. The technical solution is as follows.
[0007] In one aspect, a dimension conversion chip is provided, comprising:
[0008] An information acquisition circuit is used to acquire dimension conversion information and send the dimension conversion information to the parameter determination circuit; the dimension conversion information is used to instruct the tensor generated during the operation of the neural network model to be converted from the first arrangement format to the second arrangement format;
[0009] The parameter determination circuit is configured to receive the dimension conversion information; determine parameters of the dimension conversion instruction based on a dimension difference between the first arrangement format and the second arrangement format, and send the parameters to the dimension conversion circuit; the dimension difference includes at least one of a difference in the number of dimensions and a difference in the order of dimensions;
[0010] The dimension conversion circuit is used to receive the parameters; read the tensor from the high-bandwidth memory; execute the dimension conversion instruction according to the parameters to perform dimension conversion on the tensor in the first arrangement format to obtain a tensor in a second arrangement format, and the tensor in the second arrangement format is stored in the high-bandwidth memory for input into the neural network layer in the neural network model.
[0011] In another aspect, a dimension transposition method is provided, which is executed by a data handling engine. The method includes:
[0012] Obtaining dimension conversion information; the dimension conversion information is used to instruct the tensor generated during the operation of the neural network model to be converted from a first arrangement format to a second arrangement format;
[0013] Determining parameters of a dimension conversion instruction based on a dimensionality difference between the first arrangement format and the second arrangement format; the dimensionality difference includes at least one of a difference in the number of dimensions and a difference in the order of dimensions;
[0014] The tensor is read from the high-bandwidth memory; the dimension conversion instruction is executed according to the parameter to perform dimension conversion on the tensor in the first arrangement format to obtain a tensor in a second arrangement format, and the tensor in the second arrangement format is stored in the high-bandwidth memory for input into the neural network layer in the neural network model.
[0015] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the dimension conversion method.
[0016] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the dimensional conversion method as described in any of the above embodiments of the present application.
[0017] In another aspect, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the dimensionality conversion method.
[0018] BRIEF DESCRIPTION OF THE DRAWINGS
[0019] FIG1 is a schematic diagram of an implementation environment provided by some exemplary embodiments of the present application;
[0020] FIG2 is a schematic structural diagram of a dimension conversion chip provided by some exemplary embodiments of the present application;
[0021] FIG3 is a schematic structural diagram of a dimension conversion chip provided by other exemplary embodiments of the present application;
[0022] FIG4 is a schematic diagram of a parallel axis process provided by some exemplary embodiments of the present application;
[0023] FIG5 is a schematic structural diagram of a dimension conversion chip provided in some further exemplary embodiments of the present application;
[0024] FIG6 is a schematic diagram of a format transposition parameter table provided in some exemplary embodiments of the present application;
[0025] FIG7 is a schematic structural diagram of a dimension conversion chip provided by some further exemplary embodiments of the present application;
[0026] FIG8 is a flowchart of a dimension conversion method provided by some exemplary embodiments of the present application;
[0027] FIG9 is a schematic diagram of dimension transposition provided by some exemplary embodiments of the present application;
[0028] FIG10 is a schematic diagram of a single path transposition provided by some exemplary embodiments of the present application;
[0029] FIG11 is a schematic diagram of multiple path-based transpositions provided by some exemplary embodiments of the present application;
[0030] FIG12 is a schematic diagram of a dimension conversion chip performing dimension transposition processing provided by some exemplary embodiments of the present application;
[0031] FIG13 is a schematic diagram of an application scenario of dimension transposition processing provided by some exemplary embodiments of the present application;
[0032] FIG14 is a structural block diagram of a server provided by some exemplary embodiments of the present application. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0034] Usually AI hardware supports the dimension transposition process within a specific dimension (such as four dimensions or three dimensions, etc.). For example, the data transfer engine DTE (data transfer engine) Permute instruction supports 4 dimensions. After obtaining the tensor, the tensor is dimensionally transposed using the transposition parameters pre-set for different network layers. However, the above-mentioned transposition parameters, as pre-set contents, are relatively rigid, and when a neural network model needs to add other network layers based on the improvement of the network structure, if the other network layers need to support conversions between different dimensions, the fixed transposition parameters cannot flexibly adapt to the changes in different network layers in the neural network model, and thus cannot perform a good dimension conversion process within the neural network model, affecting the operation of the neural network model.
[0035] In an embodiment of the present application, a dimension conversion chip is provided that can flexibly set dimension conversion parameters based on the requirements for the arrangement format of tensors when processing data at different neural network layers, based on at least one of the differences in the number of dimensions and the order of dimensions of the tensors between different neural network layers, thereby achieving targeted adjustment of the tensor format. The dimension conversion chip can be deployed in various computing devices to assist neural network models in performing more efficient data analysis and processing. The dimension conversion method performed by the dimension conversion chip can be applied to various data processing fields such as image processing and signal processing, and the embodiments of the present application are not limited to this.
[0036] Some terms in the embodiments of this application are explained below.
[0037] Tensor: A tensor is a multidimensional array that can be a scalar, vector, matrix, etc. A tensor is a data structure that can be used to store data and perform various mathematical operations.
[0038] Feature Map: In deep learning, a feature map is the intermediate output obtained by the forward propagation of a specific layer of a neural network. A feature map captures the important features of the input data. In a convolutional neural network, each convolutional layer generates a new feature map by applying a convolution kernel to the input image or the feature map of the previous layer. The feature map represents the abstract information processed by the network layer. In deep learning models, feature maps are stored and transmitted as tensors. These tensors can be stored in CPU or GPU memory to facilitate fast mathematical operations.
[0039] Tensor dimensions: The dimension of a tensor refers to the number of axes of the tensor. For example, a scalar is a 0-dimensional tensor, a vector is a 1-dimensional tensor, a matrix is a 2-dimensional tensor, and so on.
[0040] The shape of a tensor is one of the basic properties of a tensor, which refers to the size of the tensor in each dimension. The shape of an n-dimensional tensor can be expressed as (D0, D1, ..., D n-1 ) indicates that, among them, D0 to D n-1 are all positive integers, representing the size of each dimension. For example, a tensor with a shape of (3, 4) represents a matrix with 3 rows and 4 columns.
[0041] Axis of a tensor: The axis is relative to the shape. The axis represents the subscript of the shape of the tensor. For example, if tensor A is a two-dimensional array with 5 rows and 6 columns, that is, shape = (5, 6), then axis = 0 represents the first dimension of the tensor, that is, the row; axis = 1 represents the second dimension of the tensor, that is, the column.
[0042] Layout format: This defines the order in which tensors are stored and read in memory. In different computational stages of neural networks (such as convolutional layers and fully connected layers), data often needs to be manipulated in different dimensions. To ensure efficient data reading and processing during these operations, deep learning frameworks often define specific layout formats. Common layout formats include NCHW, NHWC, NC1HWC0, and others.
[0043] Tensor dimensionality transformation: Also known as tensor dimension conversion, this is a tensor operation used to change the layout of a tensor to meet different computational requirements. Basic dimensionality transformation operations include reshaping, inserting new dimensions, deleting dimensions, and transposing dimensions.
[0044] Transpose: This operation, also known as transposition, changes the order of the dimensions of a tensor. For example, a 4D tensor with the NHWC format can be transformed into the NCHW format by transposing the dimensions. Here, N represents the batch size, C represents the number of channels, H represents the height, and W represents the width.
[0045] The Data Transfer Engine (DTE) is a tool used to automatically transfer data, primarily between different memories. Data movement between memories consists of two phases: moving in and moving out. For example, in the first phase, the DTE moves tensors from high-bandwidth memory into the vector processing engine cache for vector calculations. In the second phase, the DTE moves the calculation results from the vector processing engine cache to high-bandwidth memory for output.
[0046] High Bandwidth Memory (HBM): A type of memory with high bandwidth and low latency. In the AI field, high-bandwidth memory can meet the needs of storing and processing large-scale neural network models, improving the training and inference efficiency of AI models.
[0047] Vector process engine buffer (VB): refers to the storage area allocated for the vector processing engine, which can improve the processing efficiency of the vector processing engine.
[0048] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the tensors, dimensionality conversion information, dimensionality conversion parameters and other contents involved in this application are all obtained with full authorization.
[0049] Secondly, the implementation environment involved in the embodiments of the present application is described. The dimension conversion chip provided in the embodiments of the present application can be deployed in a terminal, so that the terminal can perform the dimension conversion method based on the dimension conversion chip alone; it can also be deployed in a server, so that the server can perform the dimension conversion method based on the dimension conversion chip alone; the dimension conversion chip can also be deployed in at least one device of the terminal and the server, and the terminal and the server can perform the dimension conversion method based on the dimension conversion chip through data interaction. The embodiments of the present application are not limited to this. In some embodiments, the terminal and the server interact to perform the dimension conversion method based on the dimension conversion chip as an example for description.
[0050] Schematically, please refer to FIG. 1 , the implementation environment involves a terminal 110 and a server 120 , and the terminal 110 and the server 120 are connected via a communication network 130 .
[0051] In some embodiments, the terminal 110 has a data acquisition function for acquiring data that needs to be analyzed by the neural network model. For example, the terminal acquires at least one of multiple types of data such as image data, signal data, text data, etc. as data that needs to be analyzed.
[0052] In some embodiments, the terminal 110 sends data to the server 120 through the communication network 130. The server 120 is deployed with a neural network model (for example, the neural network model can be run by an AI core) and a dimension conversion chip. The dimension conversion chip can determine the parameters of the dimension conversion instruction according to the different requirements of different neural network layers for the arrangement format of tensors during the process of data processing by the neural network model, and execute the dimension conversion instruction according to the parameters (for example, in the process of data transfer between high bandwidth memory (HDM) and vector process engine buffer (VB)), thereby realizing the dimension conversion of the tensor from a first arrangement format to a second arrangement format, wherein the first arrangement format and the second arrangement format differ in at least one of the number of dimensions and the order of dimensions.
[0053] In some embodiments, the dimension conversion chip includes an information acquisition circuit for acquiring dimension conversion information of a tensor.
[0054] The dimension conversion information is used to instruct the conversion of the tensor from the first arrangement format to the second arrangement format, that is, the dimension conversion information is used to instruct the tensor in the first arrangement format to be dimensionally converted to obtain the tensor in the second arrangement format. In addition, the information acquisition circuit also sends the dimension conversion information to the parameter determination circuit.
[0055] In some embodiments, the dimension conversion chip further includes a parameter determination circuit for receiving dimension conversion information and determining parameters of the dimension conversion instruction based on the dimension difference between the first arrangement format and the second arrangement format.
[0056] Among them, the dimensional difference includes at least one of the difference in the number of dimensions and the difference in the order of dimensions. The difference in the number of dimensions is used to express the change in the number of dimensions between the first arrangement format and the second arrangement format. The difference in the order of dimensions is used to express the change in the order of dimensions between the first arrangement format and the second arrangement format. By executing the dimension conversion instruction according to the determined parameters, the tensor can be converted from the first arrangement format to the second arrangement format.
[0057] In addition, the parameter determination circuit is further configured to send the parameters of the dimension conversion instruction to the dimension conversion circuit.
[0058] In some embodiments, the dimension conversion circuit is used to receive parameters sent by the parameter determination circuit, read tensors from the high bandwidth memory (HBM), and execute dimension conversion instructions based on the parameters (for example, in the process of moving tensors between the HBM and the vector processing engine cache (VB)), thereby performing dimension conversion on the tensor in the first arrangement format to obtain a tensor in the second arrangement format for reading and processing by the neural network layer in the neural network model.
[0059] Illustratively, server 120 performs dimension conversion on a tensor in a first arrangement format based on a deployed dimension conversion chip, thereby obtaining a tensor in a second arrangement format. The tensor in the second arrangement format can be input into a neural network layer that supports processing tensors in the second arrangement format, thereby avoiding the limitation that AI hardware only supports transposition operations of specified dimensions.
[0060] In some embodiments, server 120 analyzes data and obtains data analysis results through multiple neural network layers in the neural network model based on the coordination between the neural network model and the dimension conversion chip. In some embodiments, server 120 sends the data analysis results to terminal 110 via communication network 130.
[0061] It is worth noting that the above-mentioned terminals include but are not limited to mobile terminals such as mobile phones, tablets, portable laptops, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc., and can also be implemented as desktop computers, etc.; the above-mentioned servers can be independent physical servers, or they can be server clusters or distributed systems composed of multiple physical servers, or they can be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0062] In some embodiments, the above-mentioned server can also be implemented as a node in a blockchain system.
[0063] In combination with the above-mentioned terminology introduction and application scenarios, the dimension conversion chip provided in this application is explained, and the chip deployment server is taken as an example. As shown in Figure 2, the dimension conversion chip includes an information acquisition circuit 210, a parameter determination circuit 220 and a dimension conversion circuit 230. The different circuit modules of the dimension conversion chip are explained separately below.
[0064] (1) Information Acquisition Circuit 210
[0065] The information acquisition circuit 210 is used to obtain dimension conversion information of the tensor.
[0066] Illustratively, the information acquisition circuit 210 is a module in the dimension conversion chip. When the server calls the dimension conversion chip to perform dimension conversion, the dimension conversion chip obtains dimension conversion information through the information acquisition circuit 210 .
[0067] Schematically, a neural network model refers to a computational model composed of neurons (or nodes) in a neural network. A neural network model can also be called a machine learning model. With the help of a neural network model, various types of data can be analyzed in a targeted manner to perform corresponding machine learning tasks.
[0068] In some embodiments, image data, text data, audio data and other types of data are analyzed through a neural network model; the neural network model is pre-set, which includes multiple neural network layers. The neural network layer is the basic component unit in the neural network model. A neural network layer contains at least one neuron. The multiple neural network layers that make up the neural network model include an input layer (Input Layer), a hidden layer (Hidden Layer), an output layer (Output Layer), other layers, etc., where the other layers can be implemented as a recurrent layer in a recurrent neural network (RNN), a convolutional layer in a convolutional neural network (CNN), etc.
[0069] That is, multiple neural network layers are pre-combined to form a neural network model. The multiple neural network layers in the neural network model used to process different machine learning tasks may be the same or different, and this is not limited here.
[0070] Schematically, a pre-configured neural network model is deployed in the server. Taking image data as an example, the image data is input into the input layer of the neural network model, and the image data is analyzed through multiple network layers in the neural network model to learn the mapping relationship from the input layer to the output layer to perform various machine learning tasks.
[0071] In some embodiments, taking the case where the data to be analyzed by the neural network model is image data, the image data as input data is usually represented as a multidimensional array, which can be regarded as a tensor; alternatively, the image data is implemented as an image, and the image data is converted into a tensor through the input layer of the neural network model.
[0072] Illustratively, an image is read from an image file, which may be a color image (RGB channels) or a grayscale image; a preprocessing operation is performed on the image (at least one of resizing, normalizing, enhancing contrast, or color space conversion); and the image is converted into a tensor. For example, for a color image, a three-dimensional tensor is usually used, whose dimensions are height (H), width (W), and number of channels (C). For a grayscale image, a two-dimensional tensor can be used.
[0073] In some embodiments, taking the transmission of tensors within a neural network model as an example, multiple neural network layers are connected sequentially, with each subsequent neural network layer receiving the tensor output by the previous neural network layer and performing corresponding processing on the tensor to obtain a tensor input to the next neural network layer. In other words, the tensor is adjusted accordingly based on the different neural network layers, and the data transmitted between different neural network layers in the neural network model can be regarded as tensors, which are the feature representations transmitted between different neural network layers.
[0074] Illustratively, the analysis is performed using a tensor in the first arrangement format received by any neural network layer in the neural network model. For example, if the tensor is data input to the input layer, the first arrangement format is the arrangement format of the image data corresponding to the tensor, such as (H, W, C); or if the tensor is data input to any intermediate layer, the first arrangement format is the arrangement format determined based on the output of the previous neural network layer, such as (W, H, C).
[0075] A layout format is typically used to describe the shape and layout of a tensor, representing at least one of the arrangement and size of each dimension in the tensor. A tensor's dimensionality sequence provides key information about the data structure. For example, image data is typically represented as a four-dimensional tensor with a shape of (batch_size, height, width, channels), abbreviated as NHWC. Batch_size represents the batch size, i.e., the number of samples processed simultaneously during training or inference; height represents the height of the image; width represents the width of the image; and channels represents the number of channels in the image, typically 3 for color (RGB) images. The layout format of a tensor expressed in NHWC can be referred to as the dimensionality sequence (0, 1, 2, 3), and dimensionality conversions can be performed based on this dimensionality sequence. Alternatively, the dimensionality sequence of a tensor expressed in NHWC can be referred to as the dimensionality sequence (0, 2, 1, 3), and dimensionality conversions can be performed based on this dimensionality sequence. In other words, the dimensionality sequence can be selected during dimensionality conversion, but the baseline dimensionality sequence must be fixed throughout the conversion.
[0076] The significance of dimension sequences lies in providing a clear understanding of the structure of tensors. This is also very useful when designing neural network architectures, adjusting the format of input data, and understanding model outputs. Tensor operations in deep learning frameworks (such as TensorFlow and PyTorch) often require knowledge of the dimensionality of the data, so understanding dimension sequences is very important.
[0077] The dimension conversion information is used to instruct to convert the first arrangement format into the second arrangement format, that is, to convert the first dimensional sequence into the second dimensional sequence. The first dimensional sequence and the second dimensional sequence are different.
[0078] In some embodiments, when it is determined that there is a dimensionality conversion requirement, the information acquisition circuit 210 acquires dimensionality conversion information.
[0079] The dimension transposition requirement is a requirement for performing dimension conversion on the first dimension sequence. Illustratively, during the operation of the neural network model, the dimension conversion chip detects the operation of the neural network model in real time. When the operation meets the dimension conversion conditions, it determines that a dimension conversion requirement exists, and obtains dimension conversion information through the information acquisition circuit 210. Alternatively, when the operation of the neural network model meets the dimension conversion conditions, the dimension conversion chip is called. When the dimension conversion chip is called, it determines that a dimension conversion requirement exists, causing the information acquisition circuit 210 in the dimension conversion chip to obtain the dimension conversion information.
[0080] Indicatively, the running status reflects the next neural network layer that processes the tensor during the operation of the current neural network model. If the next neural network layer does not support processing tensors in the first arrangement format, but supports processing tensors in the second arrangement format, then the running status is considered to meet the dimension conversion conditions and have dimension conversion requirements.
[0081] Alternatively, if the neural network layer currently processing the tensor is a neural network layer used to perform dimension conversion on the tensor, then the operation is deemed to meet the dimension conversion conditions and have dimension conversion requirements.
[0082] In some embodiments, the dimensionality conversion information is automatically generated information.
[0083] Illustratively, when the dimensionality conversion conditions are met, the neural network model automatically generates dimensionality conversion information that instructs the dimensionality conversion chip to perform dimensionality conversion processing; alternatively, when the dimensionality conversion conditions are met, the dimensionality conversion chip automatically generates dimensionality conversion information that prompts the chip to execute the dimensionality conversion process. The dimensionality conversion information includes information representing the first arrangement format (or first dimensional sequence) and information representing the next neural network layer; alternatively, the dimensionality conversion information is the conversion information used by the current neural network layer when performing dimensionality conversion processing.
[0084] In some embodiments, the dimensionality conversion information is manually entered information.
[0085] Illustratively, before or during the running of the neural network model, information requiring targeted dimension conversion processing of the tensor can be manually input as dimension conversion information, etc.
[0086] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0087] The information acquisition circuit 210 also sends the dimension conversion information to the parameter determination circuit 220 .
[0088] (2) Parameter determination circuit 220
[0089] Schematically, the parameter determination circuit 220 is another module in the dimension conversion chip. When the server calls the dimension conversion chip to perform dimension conversion processing, the dimension conversion chip determines the parameters used when executing the dimension conversion instruction on the tensor through the parameter determination circuit 220.
[0090] The parameter determination circuit 220 is configured to receive dimension conversion information, that is, receive the dimension conversion information sent by the information acquisition circuit 210 .
[0091] The dimension conversion information is information used to describe the conversion of the first dimension sequence into the second dimension sequence; therefore, the dimension conversion information contains the sequence changes between the first dimension sequence and the second dimension sequence.
[0092] The parameter determination circuit 220 is further configured to determine parameters of the dimension conversion instruction based on the dimension difference between the first dimension sequence and the second dimension sequence.
[0093] The dimension difference includes at least one of the dimension quantity difference and the dimension order difference; the dimension quantity difference is used to express the change in the quantity of the dimension, and the dimension order difference is used to express the change in the order of the dimension.
[0094] Schematically, if the first dimensional sequence is (0, 1, 2, 3) and the second dimensional sequence is (0, 1, 2, 3, 4), then the dimensional difference between the first dimensional sequence and the second dimensional sequence is the difference in the number of dimensions, which means that the number of dimensions between the first dimensional sequence and the second dimensional sequence has changed from four dimensions (four dimensions; 4 dimensions, 4D) to five dimensions (five dimensions; 5 dimensions, 5D); if the first dimensional sequence is (0, 1, 2, 3) and the second dimensional sequence is (0, 2, 1, 3), then the first dimensional sequence and the second dimensional sequence have different dimensions. The dimensionality difference between degree sequences is the dimensionality order difference, which means that there is a change in the dimensionality order between the first dimensional sequence and the second dimensional sequence, where in the second dimensional sequence, the order between the second dimension "1" and the third dimension "2" is reversed; if the first dimensional sequence is (0, 1, 2, 3) and the second dimensional sequence is (0, 1, 3, 2, 4), then the dimensionality difference between the first dimensional sequence and the second dimensional sequence is the difference in the number of dimensions and the difference in the order of dimensions, which means that there is not only a change in the order of dimensions but also a change in the number of dimensions between the first dimensional sequence and the second dimensional sequence.
[0095] The change in the number of dimensions and / or the change in the order of dimensions is usually caused by data processing, conversion, or adjustment of the network architecture. Hereinafter, a brief example of the difference in the number of dimensions and the difference in the order of dimensions will be given.
[0096] 1. Difference in the number of dimensions
[0097] 1.1 Feature Extraction: In a neural network model, the number of dimensions of a feature map (tensor) typically changes after passing through a convolutional layer (i.e., a neural network layer) such as a convolutional neural network (CNN). This is because each convolutional layer can change the size of the feature map.
[0098] 1.2. Pooling layer operation: The pooling layer is usually used to reduce the spatial dimension of the feature map by taking the maximum or average value of the local area to reduce the dimension. This operation will change the dimension of the tensor.
[0099] 1.3. Fully connected layer: At the end of the neural network, the fully connected layer may convert the high-dimensional feature map into a one-dimensional vector, thereby reducing the dimension.
[0100] 2. Differences in dimension order
[0101] 2.1. Network Architecture Design: Different neural network architectures may require different input data shapes. For example, many convolutional neural networks (CNNs) expect the input data to have a dimension sequence of (batch_size, height, width, channels).
[0102] 2.2. Cross-framework compatibility: A neural network model usually uses a deep learning framework, and multiple deep learning frameworks can also be integrated into one neural network model. Different deep learning frameworks may have different requirements for the dimension sequence of tensors. If you need to apply a neural network model with multiple deep learning frameworks or combine multiple different deep learning frameworks, the dimension sequence of tensors needs to be adjusted to meet the requirements of different deep learning frameworks.
[0103] 2.3 Task Requirements: Different deep learning tasks may require tensors of different shapes. For example, image classification and object detection tasks often require tensors of different shapes. The different shapes represent the differences in the row and column distribution of the matrix, representing the arrangement of sequences of different dimensions.
[0104] 2.4. Model fusion: In some cases, it may be necessary to fuse or connect multiple neural network models, which may involve adjusting the dimensional sequence of the input data to match the expected input corresponding to multiple neural network models.
[0105] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0106] The dimension conversion instruction is used to convert the arrangement format of the tensor from a first dimensional sequence to a second dimensional sequence according to determined parameters.
[0107] The parameter determination circuit 220 is further configured to send the parameters of the dimension conversion instruction to the dimension conversion circuit 230 .
[0108] (3) Dimension Conversion Circuit 230
[0109] Illustratively, the dimension conversion circuit 230 is another module in the dimension conversion chip. When the server calls the dimension conversion chip to perform dimension conversion processing, the dimension conversion chip executes the dimension conversion instruction through the dimension conversion circuit 230 to perform dimension conversion on the tensor.
[0110] The dimension conversion circuit 230 is configured to receive the parameters sent by the parameter determination circuit 220 .
[0111] The dimension conversion circuit 230 is further configured to read the tensor from a memory, such as a high bandwidth memory (HBM).
[0112] The dimension conversion circuit 230 is also used to execute a dimension conversion instruction based on the parameters during data transfer between the memory and the vector processing engine cache, thereby performing dimension conversion on the tensor in the first arrangement format to obtain a tensor in the second arrangement format.
[0113] The tensor with the second arrangement format is used to input the neural network layer in the neural network model for analysis based on the tensor with the second arrangement format.
[0114] Schematically, the multiple neural network layers in the neural network model include a neural network layer for analyzing tensors based on the second arrangement format. After obtaining the tensor with the second arrangement format based on dimensionality conversion, the tensor with the second arrangement format is input into the neural network layer so that the neural network layer can perform analysis.
[0115] In some embodiments, the dimension conversion instruction includes a first dimension transpose instruction and a second dimension transpose instruction, and the parameters include a first parameter corresponding to the first dimension transpose instruction and a second parameter corresponding to the second dimension transpose instruction;
[0116] The dimension conversion circuit is further used to: in the process of moving the tensor in the first arrangement format from the high-bandwidth memory to the vector processing engine cache, execute the first dimension transpose instruction on the tensor in the first arrangement format according to the first parameter to obtain a tensor with an intermediate arrangement format stored in the vector processing engine cache; in the process of moving the tensor with the intermediate arrangement format from the vector processing engine cache to the high-bandwidth memory, execute the second dimension transpose instruction on the tensor with the intermediate arrangement format according to the second parameter to obtain the tensor with the second arrangement format stored in the high-bandwidth memory.
[0117] In some embodiments, it further includes: a data transfer engine DTE, used to realize the transfer of the tensor between the high-bandwidth storage and the vector processing engine cache, wherein the first dimension transposition instruction and the second dimension transposition instruction are dimension transposition instructions under the specified dimension supported by the data transfer engine DTE.
[0118] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0119] In summary, the dimension conversion information of a tensor is acquired by means of an information acquisition circuit in a dimension conversion chip. A parameter determination circuit determines the parameters of a dimension conversion instruction based on the dimensional difference between the first and second arrangement formats represented by the dimension conversion information. The dimension conversion circuit then executes the dimension conversion instruction to perform dimension conversion on the tensor, obtaining a tensor having the second arrangement format for input into a neural network layer in a neural network model that analyzes the tensor based on the second arrangement format. The dimension conversion process is integrated with the dimension conversion chip to perform targeted adjustments to the arrangement format of the tensor based on multiple neural network layers in the neural network model that support tensors of different formats, thereby obtaining tensors of different formats that are more accurately input into the corresponding neural network layers for data processing. Furthermore, the chip fully considers dimensional differences, including at least one of differences in the number of dimensions and differences in the order of dimensions, thereby enabling flexible determination of dimensional conversion parameters based on the dimensional difference acquisition, avoiding the limitations of pre-setting dimensional conversion parameters.
[0120] In some embodiments, when determining the parameters of the dimension conversion instruction by the parameter determination circuit, differential processing can be performed according to the type of dimension conversion information. Schematically, when the dimension conversion information is implemented as transposition axis information (transposition axis index tuple), the number of elements contained in the transposition axis information can be determined, thereby determining the parameters of the dimension conversion instruction based on the difference in the number of dimensions between the number and the specified number of dimensions. Among them, the transposition axis information can reflect the difference in the number of dimensions and may also reflect the difference in the order of dimensions. Schematically, the parameter determination circuit is also used to perform the following.
[0121] In some embodiments, the parameter determination circuit analyzes the dimensionality conversion information.
[0122] Illustratively, after receiving the dimension conversion information, the parameter determination circuit analyzes the dimension conversion information to determine the information content included in the dimension conversion information.
[0123] In some embodiments, in response to the dimensionality conversion information being transposed axis information, a target number of dimensions represented by the transposed axis information is determined.
[0124] The transposed axis information is used to convert a tensor in a first arrangement format into a tensor in a second arrangement format.
[0125] Schematically, transpose axes information usually refers to exchanging axes of a tensor or matrix. In neural network models, operations that apply transpose axes information (axis transpose operations) are often used to adjust the shape of tensors to adapt to different neural network models or task requirements.
[0126] For example, a tensor is a two-dimensional matrix with a shape of (m, n), where m represents the number of rows and n represents the number of columns. Transposing the tensor by the axis means swapping the rows and columns of the tensor to generate a new matrix with a shape of (n, m).
[0127] In deep learning, high-dimensional tensors can be transposed in a similar manner. For example, a tensor with a shape of (0, 1, 2, 3) has a transposed axis of (0, 1, 3, 2). (0, 1, 2, 3) can be considered the first dimension sequence of the tensor, expressing four dimensions, meaning the original number of dimensions is 4; (0, 1, 3, 2) can be considered the transposed axis, also indicating four dimensions, meaning the target number of dimensions is also 4.
[0128] In some embodiments, parameters of the dimension conversion instruction are determined based on a difference in the number of dimensions between the target number of dimensions and the specified number of dimensions.
[0129] Illustratively, after determining the target number of dimensions and the specified number of dimensions, the numerical values of the target number of dimensions and the specified number of dimensions are compared to determine the difference in the number of dimensions between the target number of dimensions and the specified number of dimensions. For example, it is determined that the target number of dimensions is less than the specified number of dimensions; or that the target number of dimensions is equal to the specified number of dimensions; or that the target number of dimensions is greater than the specified number of dimensions.
[0130] In some embodiments, the parameter determination circuit includes an information analysis unit and a processing unit.
[0131] Illustratively, the dimension conversion chip shown in FIG. 2 may also be as shown in FIG. 3 , wherein the parameter determination circuit 310 includes an information analysis unit 311 and a processing unit 312 .
[0132] In some embodiments, the information analysis unit 311 is configured to generate a processing request under the condition that the dimension number difference indicates that the target dimension number is different from the specified dimension number.
[0133] Illustratively, the parameter determination circuit 310 performs a numerical comparison on the target number of dimensions and the specified number of dimensions through the information analysis unit 311 , and generates a processing request when it is determined that the target number of dimensions is different from the specified number of dimensions.
[0134] The processing request is used to adjust the transposition axis information, such as merging or splitting elements in the transposition axis information (ie, merging or splitting the transposition axis).
[0135] In some embodiments, AI hardware generally supports transposition operations of a specified number of dimensions (such as one of 3D, 4D, 5D, etc.), such as supporting conversion from one 4D (0, 1, 2, 3) to another 4D (0, 3, 1, 2), etc. If AI hardware is used to perform a transposition operation, it is necessary to process the transposition axis information when the target number of dimensions is different from the specified number of dimensions, so that when the dimension transposition is performed based on the transposition axis information, the transposition process under the specified number of dimensions can be performed. For example: if the specified number of dimensions is 4 and the target number of dimensions is 5, the transposition axis information needs to be adjusted to perform the transposition operation in 4D form. The following is an example of a transposition instruction supported by AI hardware with a specified number of dimensions of 4.
[0136] In some embodiments, the information analysis unit 311 is further configured to use the transposition axis information as a parameter of the dimension conversion instruction in response to the target dimension number being equal to the specified dimension number.
[0137] Illustratively, the information analysis unit 311 can perform a transposition operation on the tensor in the first arrangement format by using the transposition axis information under the condition that the target number of dimensions is determined to be the same as the specified number of dimensions, that is, using the transposition axis information as a parameter of the dimension conversion instruction.
[0138] For example: the first dimensional sequence is (0, 1, 2, 3), the transposed axis information is (0, 2, 3, 1), and based on the specified number of dimensions and the target number of dimensions being 4, the transposed axis information can be used as a parameter of the dimension conversion instruction to process the tensor of the first dimensional sequence, and the resulting second dimensional sequence is (0, 2, 3, 1).
[0139] Alternatively, the first dimensional sequence is (0, 1, 3, 2), and the transposed axis information is (0, 2, 1, 3). Based on the fact that the specified number of dimensions and the target number of dimensions are both 4, the transposed axis information can be used as a parameter of the dimension conversion instruction to process the tensor of the first dimensional sequence, and the resulting second dimensional sequence is (0, 3, 1, 2), where the 0th element of the transposed axis information is 0, indicating that the 0th element of the second dimensional sequence should be the 0th element in the first dimensional sequence, that is, 0; the 1st element of the transposed axis information is 2, indicating that the 1st element of the second dimensional sequence should be the 2nd element in the first dimensional sequence, that is, 3; the 2nd element of the transposed axis information is 1, indicating that the 2nd element of the second dimensional sequence should be the 1st element in the first dimensional sequence, that is, 1; the 3rd element of the transposed axis information is 3, indicating that the 3rd element of the second dimensional sequence should be the 3rd element in the first dimensional sequence, that is, 2.
[0140] Among them, the first dimensional sequence is used to characterize the distribution of each dimension (such as NHWC) in the tensor. The first dimensional sequence of the NHWC-arranged tensor can be expressed as (0, 1, 2, 3), and the first dimensional sequence of the NCWH-arranged tensor can be expressed as (0, 1, 2, 3), etc. Therefore, the first dimensional sequence can be regarded as a description of the initial arrangement and is not limited here.
[0141] The above transposition axis information and the specified number of dimensions being the same are only illustrative examples. The transposition axis information and the specified number of dimensions may also be different. For example, if the specified number of dimensions is 4, the transposition axis information is 5D. This is not limited here.
[0142] The information analysis unit 311 is further configured to send a processing request to the processing unit 312 .
[0143] Illustratively, the processing unit 312 is instructed to perform processing based on the sending of the processing request.
[0144] The processing unit 312 is configured to receive a processing request, that is, receive a processing request sent by the information analysis unit 311 to process the transposed axis information.
[0145] The processing unit 312 is further configured to process the transposed axis information based on a quantitative relationship between the target dimension quantity and the specified dimension quantity in the processing request to obtain parameters of the dimension conversion instruction.
[0146] Illustratively, based on the fact that the processing request is sent when the target dimension number is different from the specified dimension number, the processing request may include the quantitative relationship between the target dimension number and the specified dimension number, such as: when the target dimension number is less than the specified dimension number, the processing request includes the identifier 0; or, when the target dimension number is greater than the specified dimension number, the processing request includes the identifier 1, etc.
[0147] In some embodiments, the transposition axis information is processed differently according to the different quantitative relationships between the target dimension number and the specified dimension number to obtain the processed transposition axis information as a parameter of the dimension transposition instruction.
[0148] In some embodiments, the processing unit 312 is further configured to perform axis-binning processing on the transposed axis information in response to a processing request indicating that the target dimension number is greater than the specified dimension number, and obtain the processed transposed axis information as a parameter of the dimension transposition instruction.
[0149] The axis merging process is used to merge at least two dimensions in the first dimension sequence (transposed axis information).
[0150] In some embodiments, the processing request represents the comparison of the number of dimensions through the identifier carried therein. If the identifier carried therein is 0, it means that the processing request indicates that the target number of dimensions is less than the specified number of dimensions; or, if the identifier carried therein is 1, it means that the processing request indicates that the target number of dimensions is greater than the specified number of dimensions, etc.
[0151] Indicatively, the number of dimensions is specified as the number of dimensions that can be processed by AI hardware instructions; if the target number of dimensions is greater than the specified number of dimensions, the tensor of the first dimensional sequence cannot be directly processed based on the transposed axis information. In this case, the transposed axis information needs to be processed to select at least two dimensions from the elements representing each dimension in the transposed axis information for merging, and obtain the parameters of the specified number of dimensions based on the transposed axis information.
[0152] For example, the transposed axis information is (0, 1, 3, 2, 4), and the number of dimensions is 5. If the number of dimensions supported by the AI hardware instruction is 4, then the transposed axis information (0, 1, 3, 2, 4) needs to be paralleled based on the number of dimensions 4.
[0153] Among them, it can be assumed that the first dimensional sequence before dimensional transposition is (0, 1, 2, 3, 4) - the number of dimensions is 5, and the second dimensional sequence after transposition is (0, 1, 3, 2, 4) - the target number of dimensions is 5. The above-mentioned axis processing of the transposed axis information to obtain the parameters for dimensional transposition with a dimension number of 4 can be regarded as a process of reversely deducing from (0, 1, 3, 2, 4) to obtain (0, 1, 2, 3, 4), wherein (0, 1, 2, 3, 4) is obtained by dimensional conversion with a dimension number of 4 to obtain (0, 1, 3, 2, 4).
[0154] Figure 4 shows a schematic diagram of performing a tandem operation on transposed axis information to obtain dimensionality conversion parameters. The transposed axis information is (0, 1, 3, 2, 4) with a dimension of 5. This process can be achieved through two 4D permute operations, that is, from (0, 1, 2, 3, 4) to (0, 1, 3, 2, 4) through two 4D permute operations. Each permute operation uses a set of dimensionality conversion parameters, and two sets of dimensionality conversion parameters are derived based on (0, 1, 3, 2, 4).
[0155] Step 1: First 4D transformation 410 (obtaining the first set of dimensional transformation parameters)
[0156] In the first step, the two dimensions in the transposed axis information are first parallelized, such as dimension 0 and dimension 1 are parallelized to obtain the four dimensions supported by the AI hardware instruction, thereby obtaining (new0, 3, 2, 4); since the number of dimensions is 4, (new0, new2, new1, new3) can be used to represent the content with continuous numerical dimension representation when the number of dimensions is correct; based on the fact that this process is a reverse process, if it is necessary to restore (0, 1, 2, 3, 4), or to restore (0, 1, 2, 3), it means that during the first four-dimensional transformation 410, dimension 1 and dimension 2 need to be swapped in order to obtain (0, 1, 2, 3) from (0, 2, 1, 3), which is written in the form of a dimension conversion parameter with a dimension number of 4, that is, the first dimension conversion parameter obtained by the first four-dimensional transformation 410 is (0, 2, 1, 3), or permute (0, 2, 1, 3).
[0157] Step 2: Second four-dimensional transformation 420
[0158] In the second step, the two previously parallelized dimensions must first be restored. For example, the original dimensions 0 and 1 must be restored to obtain 5 dimensions, represented as (0, 1, new1, new2, 4). The second 4D permute should further adjust the dimensions to match the final result of the original 5D permute. Since the first 4D transform 410 already produces (0, 2, 1, 3), the second 4D transform 420 only requires a single identity operation. That is, the second dimensional conversion parameters obtained by the second 4D transform 420 are (0, 1, 2, 3), or permute (0, 1, 2, 3).
[0159] It is worth noting that the above parallel axis processing of dimension 0 and dimension 1 is only an illustrative example, and the embodiments of the present application are not limited to this.
[0160] In some embodiments, the processing unit 312 is further configured to, in response to the processing request indicating that the target dimension number is smaller than the specified dimension number, perform axis splitting processing on the transposed axis information to obtain dimension conversion parameters.
[0161] The axis splitting process is used to split at least one dimension in the first dimension sequence (transposed axis information).
[0162] Indicatively, the number of dimensions is specified as the number of dimensions that can be processed by AI hardware instructions. If the target number of dimensions is less than the specified number of dimensions, the tensor of the first dimension sequence cannot be directly processed based on the transposed axis information. In this case, the transposed axis information needs to be split to select at least one dimension from the elements representing each dimension in the transposed axis information for splitting, and obtain the dimension conversion parameters of the specified number of dimensions based on the transposed axis information.
[0163] For example, the transposed axis information is (2, 1, 0), and its dimension number is 3. If the AI hardware instruction supports processing of 4 dimensions, it is considered necessary to separate the transposed axis information (2, 1, 0) based on the dimension number 4.
[0164] Among them, it can be assumed that the dimension sequence before dimensional transposition is (0, 1, 2) - the number of dimensions is 3, and the dimension sequence after transposition is (2, 1, 0) - the number of dimensions is 3. The above-mentioned axis decomposition processing of the transposed axis information to obtain the dimension conversion parameters for dimensional transposition of the dimension number 4 can be regarded as a process of reversely deducing from (2, 1, 0) to obtain (0, 1, 2), wherein (0, 1, 2) is obtained through a dimension conversion instruction with a dimension number of 4 to obtain (2, 1, 0).
[0165] Schematically, a transformation from (0, 1, 2) to (2, 1, 0) means that dimension 0 moves to dimension 2; dimension 1 moves to dimension 1; and dimension 2 moves to dimension 0.
[0166] Furthermore, to convert a 3D array to a 4D array, a dimension needs to be added. This extra dimension (D3) can be a singleton dimension and does not affect the actual content of the data. For example, by adding it to the end as a new dimension, the data structure becomes 4D—(0, 1, 2, 3).
[0167] In some embodiments, two 4D permute operations can be performed to achieve the same effect as the original 3D permute (0, 1, 2).
[0168] (1) The first 4D Permute
[0169] In the first 4D permute, you can simulate some of the effects of the original 3D permute. For example, you can choose to move D2 (dimension 2) to D0 (dimension 0) while keeping the other dimensions relatively unchanged. Thus, the transpose axis (dimensionality conversion parameter) of the first 4D permute may be:
[0170] D2 (the original D2 is moved here); D0 (the original D0); D1 (the original D1); D3 (the new dimension remains unchanged); that is, the transposed axis of the first 4D permute is (2, 0, 1, 3).
[0171] (2) The second 4D Permute
[0172] The second 4D permute needs to further adjust the order to ensure that the final permutation order is the same as the original 3D permute. Based on the above, the first and third dimensions need to be swapped (now D0 and D1), while the second dimension (now D2) and the fourth dimension (D3) remain unchanged. Therefore, the transposed axis (dimensional conversion parameter) of the second 4D permute is:
[0173] D1 (moved from the third position in the first 4D permute to here); D2 (remains unchanged); D0 (moved from the first position in the first 4D permute to here); D3 (remains unchanged); that is, the transposed axis of the second 4D permute is (1, 2, 0, 3).
[0174] That is: based on the above analysis, the two dimensional conversion parameters are (2, 0, 1, 3) and (1, 2, 0, 3); that is: after processing (0, 1, 2) by (2, 0, 1, 3), the processing result of (0, 1, 2) is transposed by (1, 2, 0, 3), and (2, 1, 0) can be obtained.
[0175] It is worth noting that the above process of adding dimension 3 can be regarded as axis decomposition processing of axis 2, and the axis decomposition processing needs to not affect the expression of the original transposed axis information; the above is only an illustrative example, and the embodiments of the present application are not limited to this.
[0176] To summarize, the dimension transposition process is performed with the help of the dimension conversion chip integration, so that the arrangement format of the tensor can be adjusted in a targeted manner based on multiple neural network layers supporting different dimensions in the neural network model, thereby obtaining tensors with different arrangement formats that are convenient for more accurate input into the corresponding neural network layer for data processing; in addition, the dimensional differences including at least one of the difference in the number of dimensions and the difference in the order of dimensions are fully considered, so that the dimension conversion parameters can be flexibly set based on the dimensional difference, avoiding the limitations of pre-setting the dimensional conversion parameters.
[0177] In an embodiment of the present application, it is introduced that under the condition that the parameter of the dimension conversion instruction is the transposition axis information, the transposition axis information can be selectively processed according to the relationship between the target dimension number corresponding to the transposition axis information and the specified dimension number, so that the dimension conversion parameters with the same number of specified dimensions can be obtained based on the transposition axis information, which is convenient for performing a more standardized dimension transposition process on the tensor based on the first arrangement format and the dimension conversion parameters under the AI hardware instruction that supports a single dimension condition, thereby improving the acquisition efficiency and stability of the tensor in the second arrangement format.
[0178] In some embodiments, when obtaining the dimension conversion parameters through the parameter determination circuit, differential processing can be performed according to the content type of the dimension conversion information. Schematically, when the dimension conversion information is implemented as a first arrangement format and a second arrangement format, the dimension conversion parameters can be obtained according to the first arrangement format and the second arrangement format. The first arrangement format corresponds to the tensor before conversion, and the second arrangement format corresponds to the tensor after conversion; the first arrangement format and the second arrangement format can at least reflect the difference in dimension order, and may also reflect the difference in the number of dimensions. Schematically, the parameter determination circuit is also used to perform the following content.
[0179] In some embodiments, the parameter determination circuit analyzes the dimensionality conversion information.
[0180] Illustratively, after receiving the dimension conversion information, the parameter determination circuit analyzes the dimension conversion information to determine the type of information included in the dimension conversion information.
[0181] In some embodiments, in response to the dimension conversion information including the first arrangement format and the second arrangement format, the dimension conversion parameter is determined based on the first arrangement format and the second arrangement format.
[0182] Schematically, a layout format is used to describe the arrangement of multidimensional arrays (such as tensors in a neural network model) in memory. It can be considered the format of data supported by the neural network layers in a neural network model. Common layout formats include NCHW, NHWC, KhKwCiCo, and B0B1HW.
[0183] The four dimensions involved in NCHW or NHWC include the batch size N, the number of channels C, the feature map height H, and the feature map width W used to describe the feature map. In other words, the arrangement format of the four dimensions N, H, W, and C is based on the expression of the feature map.
[0184] KhKwCiCo involves four arrangements, including those used to describe the convolution kernel height (Kh), convolution kernel width (Kw), input channels (Ci), and output channels (Co). Kh represents the vertical size of the convolution kernel used for sliding in the convolution operation; Kw represents the horizontal size of the convolution kernel used for sliding in the convolution operation; Ci represents the number of channels of the input data. For example, the number of input channels of a color image is the number of color channels in the image, such as 3 channels for an RGB image; Co represents the number of output channels of the convolution operation, that is, the number of convolution kernels used in the convolution layer, and each convolution kernel corresponds to one output channel. In other words, the arrangement format including Kh, Kw, Ci, and Co is based on the expression of the convolution layer.
[0185] The four dimensions involved in B0B1HW include a batch dimension (Batch Dimension, B0), another batch dimension (Batch Dimension, Part 2, B1), a height dimension (Height Dimension, H), and a width dimension (Width Dimension, W). B0 represents the number of sample batches used in training, which corresponds to the number of batches in the input data; B1 indicates that two batch dimensions may appear in the representation, which may be related to certain models or operations. If there are two batch dimensions, it may indicate batch processing corresponding to two different aspects, or to support certain special operations; H represents the size of the data or feature in the vertical direction. In convolutional neural networks, it is usually the height of the image; W represents the size of the data or feature in the horizontal direction. In convolutional neural networks, it is usually the width of the image. In other words, this arrangement format is common in deep learning frameworks, especially in convolutional neural networks.
[0186] Schematically, the first arrangement format and the second arrangement format are two different arrangement formats. The first arrangement format is used to represent the arrangement format before dimensionality conversion, and the second arrangement format is used to represent the arrangement format after dimensionality conversion. Before transposing the tensor of the first arrangement format, it is first determined that the first arrangement format needs to be transposed based on the difference between the first arrangement format and the second arrangement format.
[0187] For example: in a neural network model, neural network layer 1 and neural network layer 2 are connected front and back, and the first arrangement format supported by neural network layer 1 is different from the second arrangement format supported by neural network layer 2. Then the tensor in the first arrangement format output by neural network layer 1 needs to be transposed to obtain a tensor in the second arrangement format, and then processed by neural network layer 2 based on the tensor in the second arrangement format, etc.
[0188] In some embodiments, the parameter determination circuit includes an information analysis unit and a parameter calling unit.
[0189] Illustratively, the dimension conversion chip shown in FIG. 2 may also be as shown in FIG. 5 , wherein the parameter determination circuit 510 includes an information analysis unit 511 and a parameter calling unit 512 .
[0190] The information analysis unit 511 is used to analyze the dimension conversion information.
[0191] The information analyzing unit 511 is further configured to send a parameter determination request to the parameter calling unit 512 in response to the dimension conversion information including the first arrangement format and the second arrangement format.
[0192] In some embodiments, the information analysis unit 511 analyzes the tensor output by the previous neural network layer when data is interacted between two neural network layers in the neural network model, and determines the arrangement format corresponding to the tensor as the first arrangement format; in addition, the information analysis unit 511 also determines the second arrangement format based on the arrangement format supported by the neural network layer when data processing is performed later.
[0193] Illustratively, if the dimension conversion information acquired by the information analysis unit 511 is implemented in the first arrangement format and the second arrangement format, a parameter determination request may be sent to the parameter calling unit 512 .
[0194] The parameter determination request is used to request to obtain dimension conversion parameters for performing dimension transposition on the tensor in the first arrangement format.
[0195] The parameter calling unit 512 is used to store the format transposition parameter table.
[0196] The format transposition parameter table stores transposition parameters corresponding to a plurality of arrangement format pairs, where the arrangement format pairs include a source arrangement format and a target arrangement format, and the transposition parameters are used to transpose from the source arrangement format to the target arrangement format.
[0197] In some embodiments, the format transposition parameter table is a table obtained based on a pre-analysis of the transposition conditions between the arrangement formats, including the transposition information for dimensional transposition between any two arrangement formats. By querying the format transposition parameter table, one arrangement format can be converted into another arrangement format. This process is achieved by performing dimensional transposition on the dimensional sequence corresponding to the arrangement format.
[0198] Schematically, an arrangement format pair A is implemented as arrangement format 1 - arrangement format 2, and the arrangement format pair A corresponds to the transposition parameter a; if the first dimensional sequence of arrangement format 1 needs to be transposed, the first dimensional sequence can be transposed through the transposition parameter a to obtain the dimensional sequence corresponding to arrangement format 2.
[0199] The parameter calling unit 512 is further configured to receive a parameter determination request, that is, the parameter calling unit 512 receives the parameter determination request sent by the information analyzing unit 511 .
[0200] The parameter calling unit 512 is further configured to query the format transposition parameter table based on the first arrangement format and the second arrangement format to obtain dimension conversion parameters.
[0201] Illustratively, based on the format transposition parameter, transposition parameters corresponding to a plurality of arrangement format pairs are stored. Therefore, after determining the first arrangement format and the second arrangement format, the dimension conversion parameter can be determined by querying the format transposition parameter table.
[0202] In some embodiments, the parameter calling unit 512 is further configured to query the source arrangement format in the format transposition parameter table in the first arrangement format to determine a plurality of candidate parameters corresponding to the first arrangement format.
[0203] The multiple candidate parameters are parameters corresponding to the first arrangement format. The source arrangement format is used to represent the arrangement format before dimensional transposition. Correspondingly, the target arrangement format corresponds to the arrangement format after dimensional transposition. The source arrangement format and the target arrangement format together constitute an arrangement format pair. The multiple candidate parameters correspond one-to-one to the multiple target arrangement formats that form the arrangement format pair with the first arrangement format.
[0204] FIG6 shows a schematic example of a format transposition parameter table, which includes nine arrangement formats.
[0205] The format transposition parameters consist of parameter table rows corresponding to the source arrangement format 610 and parameter table columns corresponding to the target arrangement format 620. The source arrangement format 610 includes nine arrangement formats, and the target arrangement format 620 also includes nine arrangement formats. Among the nine arrangement formats, NHWC, NCHW, and NC1HWC0 represent arrangement formats for sequences of different dimensions composed of the same dimension; KhKwCiCo, KwCiKhCo, Co1Ci1KhKwCi0Co0, and Co1KwCiKhCo0 represent arrangement formats for sequences of different dimensions composed of the same dimension, Co1 and Co0 are the contents obtained by splitting Co into two dimensions, and Ci1 and Ci0 are the contents obtained by splitting Ci into two dimensions; B0B1HW and B0B1W1HW0 represent arrangement formats for sequences of different dimensions composed of the same dimension, and W1 and W0 are the contents obtained by splitting W into two dimensions, etc. The nine arrangement formats here are only illustrative examples.
[0206] Under the condition that transposition is performed within different dimension sequences composed of the same dimension, the format transposition parameter table shown in Figure 6 is obtained. Therefore, based on any two arrangement formats, the dimension conversion parameters can be queried from the format transposition parameter table shown in Figure 6.
[0207] Schematically, by querying the format transposition parameter table with the first arrangement format, other source arrangement formats can be filtered out based on the first arrangement format before dimensional transposition to avoid the problem of wrong dimensional transposition direction, thereby obtaining multiple candidate parameters corresponding to the first arrangement format, and the multiple candidate parameters correspond one-to-one to other target arrangement formats that form arrangement format pairs with the source arrangement format.
[0208] In some embodiments, the parameter calling unit 512 is further configured to query multiple candidate parameters in a second arrangement format to obtain dimension conversion parameters.
[0209] The dimension conversion parameter is a transposition parameter of the arrangement format pair consisting of the first arrangement format and the second arrangement format.
[0210] Indicatively, the target arrangement formats corresponding to multiple candidate parameters are screened using the second arrangement format to obtain dimension conversion parameters corresponding to both the first arrangement format and the second arrangement format, i.e., the dimension conversion parameters correspond to the arrangement format pair of "first arrangement format-second arrangement format".
[0211] In some embodiments, the parameter calling unit 512 is further used to determine the number of first dimensions of the first dimensional sequence corresponding to the first arrangement format, and determine the number of second dimensions of the second dimensional sequence corresponding to the second arrangement format, in response to not finding the arrangement format pair consisting of the first arrangement format and the second arrangement format in the format transposition parameter table; in response to the first dimension number being different from the second dimension number, the second dimension sequence is processed based on the first dimension sequence to obtain a processed second dimension sequence having the first dimension number.
[0212] Illustratively, the first dimension number is the number of dimensions in the first dimension sequence, and the second dimension number is the number of dimensions in the second dimension sequence. If the first dimension number and the second dimension number are the same, the first dimension sequence can be used as a benchmark to perform dimension transposition on the tensor of the first dimension sequence through the second dimension sequence, that is, the second dimension sequence is used as a parameter of the dimension conversion instruction.
[0213] If the number of the first dimension is different from the number of the second dimension, the second dimension sequence may be adjusted based on the first dimension sequence.
[0214] Illustratively, taking the first dimensional sequence as a reference, that is, when adjusting the second dimensional sequence, the first dimensional sequence is obtained from the second dimensional sequence by reverse deduction, so as to determine the parameters used when converting the first dimensional sequence into the second dimensional sequence during the reverse deduction process.
[0215] In some embodiments, the parameter calling unit 512 is further configured to perform axis paralleling processing on the second dimension sequence in response to the number of the second dimension being greater than the number of the first dimension.
[0216] The axis merging process is used to merge at least two dimensions in the second dimension sequence.
[0217] In some embodiments, the parameter calling unit 512 is further configured to perform axis splitting processing on the second dimension sequence in response to the number of the second dimension being smaller than the number of the first dimension.
[0218] The axis splitting process is used to split at least one dimension in the second dimension sequence.
[0219] In some embodiments, after querying the format transposition parameter table, the dimension sequences corresponding to different arrangement formats can be more comprehensively transformed in combination with the above-mentioned axis splitting or axis merging processing; or, the dimension sequences corresponding to the arrangement formats can be first split or axis merging processed, and then the format transposition parameter table can be queried to cover more dimensional transposition situations.
[0220] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0221] To summarize, the dimension transposition process is performed with the help of the dimension conversion chip integration, so that the arrangement format of the tensor can be adjusted in a targeted manner for multiple neural network layers that support different arrangement formats in the neural network model, thereby obtaining tensors with different arrangement formats that are convenient for more accurate input into the corresponding neural network layer for data processing; in addition, the dimensional differences including at least one of the difference in the number of dimensions and the difference in the order of dimensions are fully considered, so that the dimension conversion parameters can be flexibly obtained based on the dimensional difference, avoiding the limitations of pre-setting the dimensional conversion parameters.
[0222] In an embodiment of the present application, it is introduced that under the condition that the dimensional conversion information is the first arrangement format and the second arrangement format, the format transposition parameter table can be queried according to the first arrangement format and the second arrangement format, so as to find the dimensional conversion parameters corresponding to the first arrangement format and the second arrangement format, so as to use the pre-calculated format transposition parameter table to more efficiently query the dimensional conversion parameters, facilitate a more efficient dimensional transposition process of the tensor, and improve the efficiency of obtaining the tensor in the second arrangement format.
[0223] In some embodiments, in addition to transposing the tensor in the first arrangement format to obtain the tensor in the second arrangement format through the dimension conversion circuit, it is also possible to determine whether to perform data padding operations on a specified dimension of the tensor in the second arrangement format based on the location of the target address for storing the tensor in the second arrangement format. Schematically, as shown in FIG7 , the dimension conversion circuit 710 further includes a dimension transposition unit 711, an address determination unit 711, and a dimension padding unit 713.
[0224] The dimension conversion circuit 710 is further configured to determine a target address.
[0225] Illustratively, the address determining unit 711 is used to determine a target address, where the target address is an address for storing a tensor having the second arrangement format.
[0226] Schematically, the address corresponding to the target address is the source address, and the source address is used to store the tensor having the first arrangement format; the target address is used to store the tensor having the second arrangement format obtained after the tensor is dimensionally converted.
[0227] The dimension conversion circuit 710 is further configured to perform data padding adjustment on a specified dimension of the tensor in the second arrangement format based on the position of the target address to obtain an adjusted tensor.
[0228] Illustratively, the dimension padding unit 713 in the dimension conversion circuit 710 performs data padding adjustment on a specified dimension in the tensor of the second arrangement format.
[0229] In some embodiments, during data processing using a neural network model, there may be an operation to pad dimensions, which is usually related to batch processing and hardware optimization.
[0230] For example, in neural network models, batch processing is often used to process multiple samples simultaneously. To efficiently use hardware (such as graphics processing units (GPUs)), it is necessary to ensure that all samples in each batch have the same dimensions. If the samples in the input data have different dimensions, in order to perform batch processing, the samples need to be padded to make them have the same dimensions.
[0231] Schematically, components such as GPUs are generally more efficient at performing matrix operations, which require input data to have the same dimensions. By padding the data, hardware can be more effectively utilized for parallel computing, improving training and inference speeds.
[0232] For example, some hardware may require data to be aligned in memory according to certain rules, which can improve memory read and write efficiency. By padding the dimensions, we can ensure that the layout of data stored in memory complies with hardware alignment rules.
[0233] In some embodiments, whether to perform padding adjustment on a specified dimension of the tensor in the second arrangement format to obtain an adjusted tensor is determined based on the location of the target address. The specified dimension is a pre-set dimension, such as a channel dimension.
[0234] Illustratively, if a specified dimension in the tensor of the second arrangement format needs to be padded and adjusted, the specified dimension is determined from the tensor, and then the data of the specified dimension in the tensor is padded to obtain an adjusted tensor.
[0235] In some embodiments, the dimension conversion circuit 710 is further configured to, in response to the target address being located in the in-core memory, perform padding adjustment on a specified dimension of the tensor in the second arrangement format to obtain an adjusted tensor.
[0236] Among them, in-core memory is the memory deployed in the AI processor (AI core) used to run the neural network model or the memory of the computing device that performs the dimensionality conversion process.
[0237] Illustratively, both the source address and the target address are located in the in-core memory; if it is determined that the target address is located in the in-core memory, the tensor is padded and adjusted.
[0238] In some embodiments, the dimension conversion circuit 710 is further configured to store the tensor having the second arrangement format at the target address in response to the first arrangement format and the second arrangement format satisfying a preset condition.
[0239] The preset condition refers to a mapping relationship between a first preset dimension in the first arrangement format and a second preset dimension in the second arrangement format.
[0240] Illustratively, when the first arrangement format and the second arrangement format meet a predetermined preset condition, there is no need to perform data padding operations on the specified dimension of the tensor, and the tensor with the second arrangement format can be directly stored in the target address.
[0241] In some embodiments, the preset condition can be implemented as at least one of the following situations.
[0242] Precondition 1:
[0243] (1) The source dimension 2 is mapped to the target dimension 3; that is, the second dimension in the first dimension sequence is mapped to the third dimension in the second dimension sequence;
[0244] (2) Source DTE_SRC_TILE_DIM2_SIZE == 3, 5, 7; that is, this content is used to compare whether the size of the second dimension of the tensor of the first dimension sequence is 3, 5, or 7;
[0245] (3)DTE_ENHANCE_PERMUTE_MODE=1.
[0246] Precondition 2:
[0247] (1) Source dimension 2 is mapped to target dimension 2; that is, the second dimension in the first dimension sequence is mapped to the second dimension in the second dimension sequence, and the second dimension remains unchanged;
[0248] (2) The source dimension 3 is mapped to the target dimension 3; that is, the third dimension in the first dimension sequence is mapped to the third dimension in the second dimension sequence, and the third dimension remains unchanged;
[0249] (3)DTE_ENHANCE_PERMUTE_MODE=1.
[0250] Among them, DTE_ENHANCE_PERMUTE_MODE=1 means starting the enhancement mode, that is, not padding the specified dimension in the tensor.
[0251] The dimension conversion circuit 710 is further configured to send the adjusted tensor to a target address for storage.
[0252] Illustratively, the dimension conversion circuit 710 sends the adjusted tensor to the target address, thereby storing the adjusted tensor through the target address; in addition, if there is no need to pad the specified dimension in the tensor, the dimension conversion circuit 710 sends the tensor in the second arrangement format to the target address, thereby storing the tensor in the second arrangement format through the target address, etc.
[0253] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0254] In summary, the dimension transposition process is integrated with the help of the dimension conversion chip so that the arrangement format of the tensor can be adjusted specifically for multiple neural network layers that support different arrangement formats in the neural network model, thereby obtaining tensors with different arrangement formats that are convenient for more accurate input into the corresponding neural network layer for data processing; in addition, the dimensional differences including at least one of the difference in the number of dimensions and the difference in the order of dimensions are fully considered, so that the dimension conversion parameters can be flexibly obtained based on the dimensional difference, avoiding the limitations of pre-setting the dimensional conversion parameters.
[0255] In an embodiment of the present application, the content of determining whether to perform data padding operation on the tensor based on the dimension of the target address is introduced. Through the relationship between the target address and the processor currently performing dimension transposition, it is possible to flexibly choose whether to continue processing the tensor in the second arrangement format. When the preset conditions are met, there is no need to pad the specified dimension of the tensor according to the preset data padding operation, thereby reducing the data processing amount to a certain extent and improving the efficiency of data processing of the neural network model while ensuring the dimension transposition effect.
[0256] In some embodiments, a dimension conversion chip can be deployed in a computing device such as a terminal or server, and the dimension conversion method can be executed by invoking the dimension conversion chip. In some embodiments, the dimension conversion method can be executed by a data handling engine (DTE). As shown in Figure 8, the dimension conversion method can be implemented as follows: Steps 810 to 840.
[0257] Step 810, obtaining dimension conversion information; the dimension conversion information is used to indicate that the tensor generated during the operation of the neural network model is converted from a first arrangement format to a second arrangement format.
[0258] Step 820: Determine parameters of a dimension conversion instruction based on the dimensionality difference between the first arrangement format and the second arrangement format; the dimensionality difference includes at least one of a difference in the number of dimensions and a difference in the order of dimensions.
[0259] Step 830: Read the tensor from the high bandwidth memory.
[0260] Step 840: Execute the dimension conversion instruction according to the parameters to perform dimension conversion on the tensor in the first arrangement format to obtain a tensor in the second arrangement format, and store the tensor in the second arrangement format in the high-bandwidth memory for input into the neural network layer in the neural network model.
[0261] In some embodiments, step 840 includes executing the dimension conversion instruction according to the parameter during data transfer between the high bandwidth memory and the vector processing engine cache.
[0262] In some embodiments, the dimension conversion instruction includes a first dimension transpose instruction and a second dimension transpose instruction, and the parameters include a first parameter corresponding to the first dimension transpose instruction and a second parameter corresponding to the second dimension transpose instruction;
[0263] Step 840 includes: in the process of moving the tensor in the first arrangement format from the high-bandwidth memory to the vector processing engine cache, executing the first dimension transpose instruction on the tensor in the first arrangement format according to the first parameter, so as to obtain a tensor with an intermediate arrangement format stored in the vector processing engine cache; in the process of moving the tensor with the intermediate arrangement format from the vector processing engine cache to the high-bandwidth memory, executing the second dimension transpose instruction on the tensor with the intermediate arrangement format according to the second parameter, so as to obtain the tensor with the second arrangement format stored in the high-bandwidth memory.
[0264] In some embodiments, the first dimension transposition instruction and the second dimension transposition instruction are dimension transposition instructions under specified dimensions supported by the data handling engine DTE.
[0265] In some embodiments, step 820 also includes: analyzing the dimension conversion information; in response to the dimension conversion information being transposition axis information, determining the target number of dimensions represented by the transposition axis information, the transposition axis information being used to transpose the tensor in the first arrangement format into a tensor in the second arrangement format; processing the transposition axis information based on the difference in the number of dimensions between the target number of dimensions and the specified number of dimensions, and determining the first parameter and the second parameter corresponding to the first dimension transposition instruction and the second dimension transposition instruction respectively according to the processed transposition axis information.
[0266] In some embodiments, processing the transposed axis information includes: merging or splitting the transposed axis information based on a quantitative relationship between the target dimension number and the specified dimension number.
[0267] In some embodiments, in response to the target number of dimensions being greater than the specified number of dimensions, performing a merging process on the transposed axis information, wherein the merging process is used to merge at least two dimensions in the transposed axis information;
[0268] In response to the target number of dimensions being smaller than the specified number of dimensions, an axis splitting process is performed on the transposed axis information, where the axis splitting process is used to split at least one dimension in the transposed axis information.
[0269] In some embodiments, in response to the target number of dimensions being equal to the specified number of dimensions, the first parameter and the second parameter corresponding to the first dimension transpose instruction and the second dimension transpose instruction, respectively, are determined according to the transpose axis information.
[0270] In some embodiments, step 820 further includes:
[0271] Analyze the dimension conversion information; in response to the dimension conversion information including the first arrangement format and the second arrangement format, determine the first parameter and the second parameter corresponding to the first dimension transposition instruction and the second dimension transposition instruction respectively based on the first arrangement format and the second arrangement format.
[0272] In some embodiments, the method further comprises:
[0273] A format transposition parameter table is stored, and the format transposition parameter table is queried based on the first arrangement format and the second arrangement format to obtain the first parameter and the second parameter; the format transposition parameter table stores first parameters and second parameters corresponding to multiple arrangement format pairs, and the arrangement format pairs include a source arrangement format and a target arrangement format.
[0274] In some embodiments, the method further comprises:
[0275] In response to not finding the arrangement format pair consisting of the first arrangement format and the second arrangement format in the format transposition parameter table, determining the first dimension number corresponding to the first arrangement format, and determining the second dimension number corresponding to the second arrangement format; in response to the first dimension number being different from the second dimension number, processing the second arrangement format based on the first dimension number so that the processed second arrangement format has the first dimension number.
[0276] In which, in response to the number of the second dimensions being greater than the number of the first dimensions, the dimension sequence corresponding to the second arrangement format is subjected to axis merging processing, and the axis merging processing is used to merge at least two dimensions in the dimension sequence; in response to the number of the second dimensions being less than the number of the first dimensions, the dimension sequence corresponding to the second arrangement format is subjected to axis splitting processing, and the axis splitting processing is used to split at least one dimension in the dimension sequence.
[0277] In some embodiments, the method further comprises:
[0278] Determine a target address, where the target address is used to store the tensor in the second arrangement format; based on the position of the target address, perform data padding adjustment on a specified dimension in the tensor in the second arrangement format to obtain an adjusted tensor; and send the adjusted tensor to the target address for storage.
[0279] In which, in response to the target address being located in the in-core memory, data padding adjustment is performed on the specified dimension of the tensor in the second arrangement format to obtain the adjusted tensor; wherein, the in-core memory is the memory of the AI processor that runs the neural network model.
[0280] In which, in response to the first arrangement format and the second arrangement format meeting a preset condition, the tensor having the second arrangement format is stored in the target address, and the preset condition is used to define the mapping relationship between the first preset dimension in the first arrangement format and the second preset dimension in the second arrangement format.
[0281] To summarize, the dimension transposition process is performed with the help of the dimension conversion chip integration, so that the arrangement format of the tensor can be adjusted in a targeted manner based on multiple neural network layers that support different arrangement formats in the neural network model, thereby obtaining tensors with different arrangement formats that are convenient for more accurate input into the corresponding neural network layer for data processing; in addition, the dimensional differences including at least one of the difference in the number of dimensions and the difference in the order of dimensions are fully considered, so that the parameters of the dimension conversion instruction can be flexibly determined based on the dimensional difference, avoiding the limitation problem of pre-setting the parameters of the dimension conversion instruction.
[0282] In some embodiments, the above-mentioned dimension conversion method can also be called "an accelerated design method for multi-layout conversion and tensor transposition". This method can be executed by a data transfer engine (DTE), or integrated into a dimension conversion chip and deployed in the data transfer engine for execution, etc.
[0283] Illustratively, the parameter for dimension transposition within the data handling engine is called "DTE Permute." DTE Permute typically supports four dimensions, of which the third dimension, dimension 3, is data-contiguous, indicating continuous data storage. If the four dimensions are represented by (0, 1, 2, 3), then source dimensions 0 / 1 / 2 / 3 can be mapped to any target dimension. The source dimension is the dimension before the dimension transposition (which can be considered the first dimension sequence mentioned above), and the target dimension is the dimension after the dimension transposition (which can be considered the second dimension sequence mentioned above). Source dimension 0 is the first dimension within the source dimension, source dimension 1 is the second dimension within the source dimension, and so on.
[0284] Figure 9 shows an example of mapping the source dimensions corresponding to the layout format NCHW910 to the target dimensions corresponding to the layout format NHWC920, where N represents the batch size, C represents the number of channels, H represents the image height, and W represents the image width. This can be thought of as mapping the source dimensions (0, 1, 2, 3) to the target dimensions (0, 2, 3, 1). Taking the dimension represented by dim as an example, this process is implemented using the following four-dimensional dimension conversion parameter - permute(4D):
[0285] DTE_TILE_PERMUTE_DIM3_MAP=1, target dim3 comes from source dim1——C;
[0286] DTE_TILE_PERMUTE_DIM2_MAP=3, target dim2 comes from source dim3——W;
[0287] DTE_TILE_PERMUTE_DIM1_MAP=2, target dim1 comes from source dim2——H;
[0288] DTE_TILE_PERMUTE_DIM0_MAP=0, target dim1 comes from source dim0——N.
[0289] In some embodiments, when the target address is in in-core memory, the target dimension 3 will be padded to bandwidth alignment. However, when DTE_ENHANCE_PERMUTE_MODE (enhancement mode) is enabled, the target dimension 3 will not be padded. Enhancement mode can only be enabled if the following pre-set conditions are met:
[0290] Precondition 1:
[0291] (1) Source dim2 is mapped to target dim3. For example (map[2]=3, map[3]=2);
[0292] (2) Source DTE_SRC_TILE_DIM2_SIZE==3, 5, 7;
[0293] (3)DTE_ENHANCE_PERMUTE_MODE=1.
[0294] Precondition 2:
[0295] (1) Source dim2 is mapped to target dim2;
[0296] (2) Source dim3 is mapped to target dim3;
[0297] (3)DTE_ENHANCE_PERMUTE_MODE=1.
[0298] The product of the target dim2 size and dim3 size will be automatically padded to align with the bandwidth.
[0299] In linear algebra, when a matrix is transposed, the rows of the matrix are actually converted to columns, and the columns to rows. This means that the positions of the elements in the original matrix have changed in the transposed matrix. Suppose there is a matrix A with m rows and n columns, and its transposed matrix is denoted as A^T. Then the dimension of A^T is n rows and m columns, that is, the number of columns of the original matrix becomes the number of rows of the transposed matrix, and the number of rows of the original matrix becomes the number of columns of the transposed matrix. The principle of the transposition operation is: for each element A[i][j] in the matrix A, place it at position A^T[j][i] in the transposed matrix A^T. In other words, the element in the i-th row and j-th column of the matrix A becomes the element in the j-th row and i-th column of the transposed matrix A^T. In a traditional central processing unit (CPU), the transposition operation can be implemented by exchanging the row index and column index of the matrix.
[0300] During the dimension transposition process, the accelerator's DTE Permute instruction performs a basic 4D transposition. The first on-path transposition occurs while data is being moved from the source storage location, the High Bandwidth Memory (HBM), to the target storage space, the Vector Process Engine Buffer (VB). On-path transposition refers to the simultaneous dimension transposition of the data.
[0301] Schematically, Figure 10 shows a schematic diagram of a single on-path transposition. The source dim (axis 1) and dest dim (axis 2) data of the matrix data block are swapped, A1, A4, and A7 are transformed from columns to rows, and A1, A2, and A3 are transformed from rows to columns. In VB1010, ping-pong parallelism is a parallel computing mode typically used to describe interactive communication and computation between two or more processes or threads, and to describe processes or threads taking turns sending messages and performing computations. Ping and pong are each considered components of VB1010.
[0302] The data in the ping is transported to VB 1010, and the data in the pong is calculated. The ping-pong switching is pipelined and parallelized to improve the transport and calculation efficiency of the transposition transformation. If only one transposition transformation is required, after the data is calculated, it is moved out of VB1010 and returned to HBM1020 to obtain the calculation result of a successful dimension transposition.
[0303] Schematically, as shown in FIG11 , it is a schematic diagram of performing two road-side transpositions.
[0304] For a 4D tensor that needs to be transposed twice, for example, if the dimension sequence is (0, 1, 2, 3), axis 1 (dim0) and axis 2 (dim1) need to be transposed, and axis 2 (dim2) and axis 4 (dim3) need to be transposed. The first on-path transposition 1110 transposes axis 1 and axis 2, that is, the rows and columns of each data block A / B / C / D are transposed from HBM to the ping / pong buffer in VB; the second on-path transposition 1120 needs to transpose axis 3 and axis 4. The second on-path transposition 1120 can be performed during the process of moving out from the source storage location ping / pong VB to HBM, and the data sorted in rows AB are transposed into columns AB, and the data in columns AC are transposed into rows AC, to obtain the final result data of the two transpositions, so that the rows and columns of data blocks A, B, C, and D are transposed, and the row and column data in each data block are also transposed.
[0305] In some embodiments, since the permute instruction of AI hardware generally supports permutation of 4D tensors, in actual neural network application scenarios, the need for transposing tensors of 5D, 6D, or even larger dimensions often arises. In this case, it is necessary to decompose the permute of a 5D or 6D tensor into two or more permutations of a 4D tensor, and at the same time, automatically adjust its permute parameters. For example, for a 5D transpose, its permute parameters are (0, 1, 3, 2, 4), and the axes are decomposed into (0, 2, 1, 3) and (0, 1, 2, 3); wherein the 0-axis and 1-axis of the original permute are first merged into one axis (axis processing), and then its permute parameters are adjusted to 4D (0, 2, 1, 3). After performing a 4D permute, the transposition effect is actually achieved. No additional operation is required for the second transposition, so the permute parameters are kept at (0, 1, 2, 3), thus completing the 5D transpose.
[0306] In some embodiments, common arrangement formats in AI chips include: feature map formats: NHWC, NCHW, NC1HWC0 (5D arrangement format, based on the c dimension splitting of NHWC into c1 and c0, c0 is often 32 or 64), weight arrangement formats: KhKwCiCo, KwCiKhCo, Co1Ci1KhKwCi0Co0 (6D weight format, based on the ci and co dimension splitting of KhKwCiCo), Co1KwCiKhCo0 (5D weight format, kw ci kh co, co split, used for scenes with small input channels 5D weight format, kw ci kh co, co split, used for scenarios with small input channels) and Ci1KhKwCi0NCo (5D weight format, ci split, co expansion, used for scenarios with small output channels), matmul and batchmatmul layout formats: B0B1HW, B0B1W1HW0; in specific network scenarios, only the data layout format of the source src tensor and the data layout format of the known target dst tensor are often declared. There is no specific transpose node in the network and no permute parameter is provided for such layout format conversion. Therefore, a general multi-layout conversion table (that is, the above-mentioned format transposition parameter table) can be established. The input is the data layout format of the source src tensor and the data layout format of the target dst tensor, and the output is the two permute parameters for implementing the src to dst conversion, thereby realizing the automatic layout conversion operations between feature maps, weights, batch matmul and matmul, so that the network operation and calculation meet the layout requirements, and there is no layout format error and the whole network operation error.
[0307] Figure 12 shows the architecture for layout conversion and multi-dimensional (N Dimension, ND) transposition. Throughout the design, the required inputs are the source dimension format 1210 (src tensor format) and the target dimension format 1220 (dst tensor format), or transpose axes information 1230. Source dimension format 1210 and target dimension format 1220 can be considered one form of dimensional conversion information: source dimension format 1210 is the first layout format corresponding to the first dimensional sequence, and target dimension format is the second layout format corresponding to the second dimensional sequence. Transpose axes information can be considered another form of dimensional conversion information.
[0308] The entire design transposition uses the input source and destination formats to query the layout conversion table (Figure 6) to obtain the permute axes information twice. For example, when converting from NCHW to NHWC, the permute axes information obtained is (0, 1, 2, 3) for the first and (0, 2, 3, 1) for the second. The transposition operation is performed based on the obtained permute axes information, and the source tensor in the HBM is transposed to obtain the dst tensor and stored back in the HBM.
[0309] In addition, you can also input arbitrary ND transpose axes information, for example: (0, 1, 3, 2, 4) after splitting and merging the axes to obtain two transpose axes information (i.e., dimension conversion parameters). Similarly, according to the obtained Permute_1_axes and Permute_2_axes as information input to transpose_6d (here 6D is used as an example, it can also be input to 4D, 5D, etc.), and perform two transpose_4d operations to efficiently complete the ND transpose operation.
[0310] In some embodiments, when performing dimensional transposition on tensors with channel numbers of 3, 5, and 7 (channel 3, 5, 7), a dedicated hardware instruction enhancement mode can be used. When transposing along the way, no additional padding is required to align to 128 bytes, that is, when the data type is fp32, the channel is padded to 32, and when the data type is fp16, the channel is padded to 64; when the channel is not 3, 5, or 7 but is also a small channel, the corresponding padding is padded to 3, 5, or 7 instead of 32 or 64, which greatly reduces the amount of non-valid data that needs to be transported and improves permute performance.
[0311] When DTE_ENHANCE_PERMUTE_MODE is enabled, target dimension 3 will not be filled. Enhancement mode can only be enabled when the following pre-conditions are met:
[0312] Precondition 1:
[0313] Source dim2 is mapped to target dim3;
[0314] Source DTE_SRC_TILE_DIM2_SIZE==3, 5, 7;
[0315] DTE_ENHANCE_PERMUTE_MODE=1.
[0316] For b32, the target dim2 will be automatically filled with 64, and for b16, the target dim2 will be automatically filled with 128, where b is the abbreviation of batch.
[0317] Precondition 2:
[0318] source dim2 should be mapped to destination dim2;
[0319] Source dim3 should be mapped to destination dim3;
[0320] DTE_ENHANCE_PERMUTE_MODE=1.
[0321] And the multipliers of the target dim2 dimension and dim3 dimension will be automatically padded to align with the bandwidth.
[0322] In some embodiments, the method can be applied to the application scenario of a neural network model. The transpose operation is one of the most common operations used to change the shape and arrangement of Tensors to adapt to the arrangement format requirements between different layers.
[0323] Schematically, in models such as convolutional neural networks, recurrent neural networks, and Transformer, since the dimensional order and arrangement required for the input data and the weights of each layer may be different, operations such as transpose need to be used to adjust the input data.
[0324] For example: for convolutional neural networks, the input data is usually in NCHW format, and the weights of the convolutional layer are in OC3 format, which includes out_channels, in_channels, kernel_height, and kernel_width parameters; for recurrent neural networks, the input data is usually in the format of (batch_size, sequence_length, input_size), and the output data of the recurrent layer is in the format of (batch_size, hidden_size).
[0325] In addition to the above scenarios, there are other application scenarios that may involve transpose operations. For example, in image segmentation tasks, we can use transpose to transpose the feature map so that it can be aligned with the original image. In natural language processing, we can use transpose to transpose the word vector matrix so that it can be calculated with the context matrix. In the attention mechanism, we can use transpose to transpose the attention matrix so that it can be multiplied with the query matrix. The transpose operation is widely used in deep learning neural networks, and the data layout transformation requirements of AI processors are also quite common. An efficient transpose operation is very important.
[0326] As shown in Figure 13, in a dot product (Dot) operation 1310, feature representation can be changed based on processes such as transposition and resizing. In an element-wise (Eltwise) operation 1320, feature representation can be changed based on processes such as transposition and activation transformation. In a convolution (Conv) operation 1330, feature representation can be changed based on processes such as resizing and transposition. The transposition operation can change the dimensional sequence corresponding to the tensor.
[0327] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0328] To summarize, the dimension transposition process is performed with the help of the dimension conversion chip integration, so that the arrangement format of the tensor can be adjusted in a targeted manner based on multiple neural network layers that support different arrangement formats in the neural network model, thereby obtaining tensors with different arrangement formats that are convenient for more accurate input into the corresponding neural network layer for data processing; in addition, the dimensional differences including at least one of the difference in the number of dimensions and the difference in the order of dimensions are fully considered, so that the parameters of the dimension conversion instruction can be flexibly obtained based on the dimensional difference, avoiding the limitation problem of pre-setting the parameters of the dimension conversion instruction.
[0329] In the embodiments of the present application, by using at least one dimension conversion information in the layout format and transposed axis information, that is, by querying at least one of the format transposition parameter table and splitting / merging axis processing, the parameters of the dimension conversion instruction can be obtained more efficiently and selectively, thereby effectively improving the efficiency and versatility of layout conversion and optimizing the transposition performance when performing dimension transposition on tensors of different layout formats. This meets the needs of various deep learning networks and effectively improves the competitiveness of data processing.
[0330] Figure 14 shows a schematic diagram of the structure of a server provided by some exemplary embodiments of the present application. The server 1400 includes a central processing unit (CPU) 1401, a system memory 1404 including a random access memory (RAM) 1402 and a read-only memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the CPU 1401. The server 1400 also includes a mass storage device 1406 for storing an operating system 1413, application programs 1414, and other program modules 1415.
[0331] The mass storage device 1406 is connected to the central processing unit 1401 through a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1406 and its associated computer-readable media provide non-volatile storage for the server 1400. In other words, the mass storage device 1406 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0332] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. The system memory 1404 and mass storage device 1406 described above may be collectively referred to as memory.
[0333] According to various embodiments of the present application, the server 1400 may also be connected to a remote computer on a network such as the Internet for operation. That is, the server 1400 may be connected to the network 1412 via the network interface unit 1411 connected to the system bus 1405, or the network interface unit 1411 may be used to connect to other types of networks or remote computer systems (not shown).
[0334] The memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.
[0335] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the dimensional conversion method provided by the above-mentioned method embodiments.
[0336] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the dimensional conversion method provided by the above-mentioned method embodiments.
[0337] Embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the dimensionality conversion method described in any of the above embodiments.
[0338] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A dimension conversion chip, comprising: An information acquisition circuit, configured to acquire dimension conversion information and send the dimension conversion information to a parameter determination circuit; The dimension conversion information is used to indicate that a tensor generated during the operation of a neural network model is converted from a first arrangement format to a second arrangement format; The parameter determination circuit is configured to receive the dimension conversion information, determine parameters of a dimension conversion instruction based on a dimension difference between the first arrangement format and the second arrangement format, and send the parameters to a dimension conversion circuit; the dimension difference includes at least one of a dimension quantity difference and a dimension order difference; The dimension conversion circuit is configured to receive the parameters, read the tensor from a high-bandwidth memory, execute the dimension conversion instruction according to the parameters, so as to perform dimension conversion on the tensor in the first arrangement format to obtain a tensor in the second arrangement format; the tensor in the second arrangement format is stored in the high-bandwidth memory and is used for input to a neural network layer in the neural network model.
2. The chip according to claim 1, wherein: The dimension conversion circuit is further configured to execute the dimension conversion instruction according to the parameters during the process of data transfer between the high-bandwidth memory and a vector processing engine cache.
3. The chip according to claim 2, wherein, The dimension conversion instruction includes a first dimension transpose instruction and a second dimension transpose instruction, and the parameters include a first parameter corresponding to the first dimension transpose instruction and a second parameter corresponding to the second dimension transpose instruction; The dimension conversion circuit is further configured to: during the process of moving the tensor in the first arrangement format from the high-bandwidth memory to the vector processing engine cache, execute the first dimension transpose instruction on the tensor in the first arrangement format according to the first parameter to obtain a tensor in an intermediate arrangement format stored in the vector processing engine cache; during the process of moving the tensor in the intermediate arrangement format from the vector processing engine cache to the high-bandwidth memory, execute the second dimension transpose instruction on the tensor in the intermediate arrangement format according to the second parameter to obtain the tensor in the second arrangement format stored in the high-bandwidth memory.
4. The chip according to claim 3, further comprising: A data transfer engine DTE, configured to implement the transfer of the tensor between the high-bandwidth memory and the vector processing engine cache, wherein the first dimension transpose instruction and the second dimension transpose instruction are dimension transpose instructions under a specified dimension supported by the data transfer engine DTE.
5. The chip according to claim 3 or 4, wherein: The parameter determination circuit is further configured to analyze the dimension conversion information; in response to the dimension conversion information being transpose axis information, determine a target dimension quantity represented by the transpose axis information, and the transpose axis information is used to transpose the tensor in the first arrangement format into the tensor in the second arrangement format; Based on a dimension quantity difference between the target dimension quantity and a specified dimension quantity, process the transpose axis information, and determine the first parameter and the second parameter respectively corresponding to the first dimension transpose instruction and the second dimension transpose instruction according to the processed transpose axis information.
6. The chip according to claim 5, wherein: The parameter determination circuit is further configured to; Based on the quantitative relationship between the number of target dimensions and the number of specified dimensions, perform axis merging or axis splitting on the transposed axis information.
7. The chip according to claim 5 or 6, wherein, The parameter determination circuit is further configured to; In response to the number of target dimensions being greater than the number of specified dimensions, perform axis merging on the transposed axis information, where the axis merging is used to merge at least two dimensions in the transposed axis information; In response to the number of target dimensions being less than the number of specified dimensions, perform axis splitting on the transposed axis information, where the axis splitting is used to split at least one dimension in the transposed axis information.
8. The chip according to any one of claims 5-7, wherein, The parameter determination circuit is further configured to; In response to the number of target dimensions being equal to the number of specified dimensions, determine the first parameter and the second parameter corresponding to the first dimension transposition instruction and the second dimension transposition instruction respectively according to the transposed axis information.
9. The chip according to claim 3 or 4, wherein, The parameter determination circuit is further configured to: Analyze the dimension conversion information; in response to the dimension conversion information including the first layout format and the second layout format, determine the first parameter and the second parameter corresponding to the first dimension transposition instruction and the second dimension transposition instruction respectively based on the first layout format and the second layout format.
10. The chip according to claim 9, wherein, The parameter determination circuit is further configured to; Store a format transposition parameter table, query the format transposition parameter table based on the first layout format and the second layout format, and obtain the first parameter and the second parameter; The format transposition parameter table stores the first parameter and the second parameter corresponding to multiple layout format pairs respectively, and the layout format pair includes a source layout format and a target layout format.
11. The chip according to claim 10, wherein, The parameter determination circuit is further configured to; In response to not querying the layout format pair composed of the first layout format and the second layout format in the format transposition parameter table, determine the number of first dimensions corresponding to the first layout format, and determine the number of second dimensions corresponding to the second layout format; in response to the number of first dimensions being different from the number of second dimensions, based on the number of first dimensions, process the second layout format so that the processed second layout format has the number of first dimensions.
12. The chip according to claim 11, wherein: The parameter determination circuit is further configured to; In response to the number of second dimensions being greater than the number of first dimensions, perform axis merging on the dimension sequence corresponding to the second layout format, where the axis merging is used to merge at least two dimensions in the dimension sequence; in response to the number of second dimensions being less than the number of first dimensions, perform axis splitting on the dimension sequence corresponding to the second layout format, where the axis splitting is used to split at least one dimension in the dimension sequence.
13. The chip according to any one of claims 1 to 12, wherein The dimension conversion circuit is further configured to determine a target address for storing the tensor having the second layout format; based on the position of the target address, perform data padding adjustment on the specified dimension in the tensor of the second layout format to obtain an adjusted tensor; and send the adjusted tensor to the target address for storage.
14. The chip according to claim 13, wherein: The dimension conversion circuit is also used to perform data padding adjustment on the specified dimension of the tensor in the second arrangement format in response to the target address being located in the in-core memory to obtain the adjusted tensor; wherein the in-core memory is the memory of the AI processor running the neural network model.
15. The chip according to claim 13 or 14, wherein: The dimension conversion circuit is also used to store the tensor having the second arrangement format in the target address in response to the first arrangement format and the second arrangement format meeting a preset condition, and the preset condition is used to define a mapping relationship between the first preset dimension in the first arrangement format and the second preset dimension in the second arrangement format.
16. A dimension conversion method, executed by a data handling engine, comprising: Get dimension conversion information; The dimension conversion information is used to instruct the conversion of the tensor generated during the operation of the neural network model from the first arrangement format to the second arrangement format; determining a parameter of a dimension conversion instruction based on a dimensionality difference between the first arrangement format and the second arrangement format; The dimensional difference includes at least one of a difference in the number of dimensions and a difference in the order of dimensions; The tensor is read from the high-bandwidth memory; the dimension conversion instruction is executed according to the parameter to perform dimension conversion on the tensor in the first arrangement format to obtain a tensor in a second arrangement format, and the tensor in the second arrangement format is stored in the high-bandwidth memory for input into the neural network layer in the neural network model.
17. A computer device comprising a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the dimensionality conversion method according to claim 16.
18. A computer-readable storage medium storing at least one program, wherein the at least one program is loaded and executed by a processor to implement the dimension conversion method according to claim 16.
19. A computer program product comprising computer instructions, wherein when the computer instructions are executed by a processor, the dimensionality conversion method according to claim 16 is implemented.
Citation Information
Patent Citations
Data processing method and system and related equipment
CN114968612A
Data processing method, computer equipment and chip
CN116821019A
Method for permuting dimensions of a multi-dimensional tensor
US20220129744A1
Cited By
Processor, chip product, computer equipment and tensor processing method
CN121785664A