Dimension transpose chip, method, device, storage medium and program product

Through the coordinated work of the information acquisition, parameter acquisition and transposition module of the dimension transpose chip, the problem that transposed parameters in the neural network structure cannot adapt to changes in different network layers is solved, and flexible dimension adjustment and efficient data processing of tensor data are realized.

CN120354897APending Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410086998.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing neural network structure, the fixedly set transpose parameters cannot flexibly adapt to the changes of different network layers, resulting in poor dimensional transformation process and affecting the operation of the neural network structure.

Method used

It provides a dimension transposition chip, which obtains the dimension transposition information of tensor data through the information acquisition module, and the parameter acquisition module obtains the dimension transposition parameters according to the dimension difference, and the dimension transposition module performs the dimension transposition of tensor data to adapt to the needs of different neural network layers.

Benefits of technology

It realizes flexible dimension adjustment of tensor data, adapts to the needs of multiple neural network layers, improves the accuracy and efficiency of data processing, and avoids pre-set limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354897A_ABST
    Figure CN120354897A_ABST
Patent Text Reader

Abstract

The invention discloses a dimension transposition chip, method and device, a storage medium and a program product, and relates to the technical field of computers. The chip comprises an information acquisition module used for acquiring dimension transposition information corresponding to tensor data, and the tensor data is a feature representation which is generated in the operation process of a neural network structure and is expressed by a first dimension sequence; the parameter acquisition module is used for acquiring a dimension transposition parameter based on the dimension difference between the first dimension sequence and the second dimension sequence; and the dimension transposition module is used for carrying out dimension transposition on the tensor data of the first dimension sequence according to the dimension transposition parameters to obtain tensor data with a second dimension sequence. Through the above mode, the dimension transposition parameter can be flexibly acquired by considering at least one dimension difference in the dimension number difference and the dimension sequence difference according to the dimension sequence condition when data processing is performed on different neural network layers. The method can be applied to various scenes such as cloud technology, artificial intelligence and intelligent transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and particularly to a dimension transposition chip, method, device, storage medium, and program product. Background Art

[0002] A neural network is a computational model that mimics the structure and function of a biological neural network and is widely used in various fields. Neural networks usually store data in the form of tensors, and the layout of the data in the tensor may have different sequence expressions between different network layers of the neural network structure.

[0003] In related technologies, usually a neural network structure supports the dimension conversion process within one dimension (such as four dimensions or three dimensions, etc.). After obtaining the tensor data, the dimension conversion of the tensor data is performed through the transposition parameters preset for different network layers respectively, so as to achieve the dimension conversion process under the dimension corresponding to the tensor data.

[0004] However, the above transposition parameters, as preset content, have great rigidity. And when a neural network structure needs to add other network layers based on the improvement of the network structure, if the other network layers need to support the conversion between different dimensions, the fixedly set transposition parameters cannot flexibly adapt to the changes of different network layers in the neural network structure, thus unable to perform a good dimension conversion process within the neural network structure, affecting the operation of the neural network structure. Summary of the Invention

[0005] Embodiments of the present application provide a dimension transposition chip, method, device, storage medium, and program product, which can specifically adjust the dimension sequence of tensor data according to the dimension sequence situation when processing data by different neural network layers, and flexibly obtain dimension transposition parameters by considering at least one of the dimension quantity difference and the dimension sequence difference. The technical solution is as follows.

[0006] On the one hand, a dimension transposition chip is provided, and the chip includes:

[0007] An information acquisition module, configured to acquire dimension transposition information corresponding to tensor data, where the tensor data is a feature representation expressed in a first dimension sequence generated during the operation of a neural network structure, and the dimension transposition information is used to indicate transposing the first dimension sequence into a second dimension sequence; and sending the dimension transposition information to a parameter acquisition module;

[0008] The parameter acquisition module is configured to receive dimension transposition information; obtain dimension transposition parameters based on the dimension difference between the first dimension sequence and the second dimension sequence, where the dimension difference includes at least one of a dimension quantity difference and a dimension sequence difference, the dimension quantity difference is used to represent the change in the number of dimensions, the dimension sequence difference is used to represent the change in the order of dimensions, and the dimension transposition parameters are used to transpose the first dimension sequence to the second dimension sequence; and send the dimension transposition parameters to the dimension transposition module.

[0009] The dimension transposition module is configured to receive the dimension transposition parameters; obtain the tensor data; perform dimension transposition on the tensor data of the first dimension sequence with the dimension transposition parameters to obtain tensor data with a second dimension sequence, and the tensor data with the second dimension sequence is used to input a neural network layer that analyzes based on the second dimension sequence in the neural network structure.

[0010] On the other hand, a dimension transposition method is provided, and the method includes:

[0011] Obtain tensor data, where the tensor data is a feature representation generated during the operation of the neural network structure and expressed in the arrangement of the first dimension sequence.

[0012] Obtain the dimension transposition information corresponding to the tensor data, where the dimension transposition information is used to indicate transposing the first dimension sequence corresponding to the tensor data to a second dimension sequence.

[0013] Based on the dimension difference between the first dimension sequence and the second dimension sequence, obtain dimension transposition parameters, where the dimension difference includes at least one of a dimension quantity difference and a dimension sequence difference, the dimension quantity difference is used to represent the change in the number of dimensions, the dimension sequence difference is used to represent the change in the order of dimensions, and the dimension transposition parameters are used to transpose the first dimension sequence to the second dimension sequence.

[0014] Perform dimension adjustment on the tensor data of the first dimension sequence with the dimension transposition parameters to obtain the tensor data with the second dimension sequence, and the tensor data with the second dimension sequence is used to input a network layer that analyzes based on the second dimension sequence in the neural network.

[0015] On the other hand, a computer device is provided, and the computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the dimension transposition method as described in any one of the above embodiments of the present application.

[0016] On the other hand, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the dimension transposition method as described in any one of the embodiments of the present application above.

[0017] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the dimension transposition method as described in any one of the above embodiments.

[0018] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0019] By means of the information acquisition module in the dimension transposition chip, the dimension transposition information of the tensor data is acquired, and the dimension transposition parameters are acquired by the parameter acquisition module according to the dimension difference between the first dimension sequence and the second dimension sequence represented by the dimension transposition information. Furthermore, under the dimension transposition module, the tensor data is dimension-transposed with the dimension transposition parameters to obtain tensor data with the second dimension sequence, so as to be input into a neural network layer for analysis based on the second dimension sequence in the neural network structure. By integrating the dimension transposition chip to execute the dimension transposition process, the dimension sequence of the tensor data can be targeted adjusted based on multiple neural network layers supporting different dimensions in the neural network structure, so as to obtain tensor data with different dimensions that is more convenient to accurately input into the corresponding neural network layer for data processing. In addition, the dimension difference situation including at least one of the dimension quantity difference and the dimension sequence difference is fully considered, so that the dimension transposition parameters can be flexibly acquired based on the dimension difference, avoiding the limitation problem of presetting dimension transposition parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;

[0022] Figure 2 is a schematic diagram of the structure of a dimension transposition chip provided by an exemplary embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of a dimension transposed chip provided by another exemplary embodiment of the present application;

[0024] Figure 4 It is a schematic diagram of parallel-axis processing provided by an exemplary embodiment of the present application;

[0025] Figure 5 It is a schematic structural diagram of a dimension transposed chip provided by still another exemplary embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of a format transposed parameter table provided by an exemplary embodiment of the present application;

[0027] Figure 7 It is a schematic structural diagram of a dimension transposed chip provided by yet another exemplary embodiment of the present application;

[0028] Figure 8 It is a flowchart of a dimension transposition method provided by an exemplary embodiment of the present application;

[0029] Figure 9 It is a schematic diagram of dimension transposition provided by an exemplary embodiment of the present application;

[0030] Figure 10 It is a schematic diagram of single-pass in-path transposition provided by an exemplary embodiment of the present application;

[0031] Figure 11 It is a schematic diagram of multiple-pass in-path transposition provided by an exemplary embodiment of the present application;

[0032] Figure 12 It is a schematic diagram of a dimension transposed chip performing dimension transposition processing provided by an exemplary embodiment of the present application;

[0033] Figure 13 It is a schematic diagram of an application scenario of dimension transposition processing provided by an exemplary embodiment of the present application;

[0034] Figure 14 It is a structural block diagram of a server provided by an exemplary embodiment of the present application. Detailed implementation manners

[0035] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0036] In the related art, generally, a neural network structure supports the dimension conversion process within a certain dimension (such as four - dimensional or three - dimensional, etc.). After obtaining tensor data, the dimension of the tensor data is converted through the transpose parameters preset for different network layers respectively, so as to realize the dimension conversion process under the dimension corresponding to the tensor data. However, as the preset content, the above - mentioned transpose parameters have great rigidity. And when a neural network structure needs to add other network layers based on the improvement of the network structure, if the other network layers need to support the conversion between different dimensions, the fixedly set transpose parameters cannot flexibly adapt to the changes of different network layers in the neural network structure, thus unable to perform a good dimension conversion process within the neural network structure, affecting the operation of the neural network structure.

[0037] In the embodiments of the present application, a dimension transpose chip is introduced, which can specifically adjust the dimension sequence of tensor data according to the dimension sequence situation when processing data by different neural network layers. By considering at least one of the dimension quantity difference and the dimension sequence difference, the dimension transpose parameters are flexibly obtained. The dimension transpose chip can be deployed in various computing devices, so as to assist the neural network structure in performing a more efficient data analysis and processing process. The dimension transpose method executed by the dimension transpose chip can be applied to various data processing fields such as the image processing field and the signal processing field. The embodiments of the present application do not limit this.

[0038] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the tensor data, dimension transpose information, dimension transpose parameters, etc. involved in the present application are all obtained under full authorization.

[0039] Secondly, the implementation environment involved in the embodiments of the present application is described. The dimension transpose chip provided by the embodiments of the present application can be deployed in the terminal, so that the terminal independently executes the dimension transpose method based on the dimension device chip; it can also be deployed in the server, so that the server independently executes the dimension transpose method based on the dimension device chip; it can also deploy the dimension transpose chip inside at least one of the terminal and the server, and the terminal and the server execute the dimension transpose method based on the dimension transpose chip through data interaction. The embodiments of the present application do not limit this. Optionally, taking the interaction between the terminal and the server based on the dimension transpose chip to execute the dimension transpose method as an example for description.

[0040] Schematically, please refer to Figure 1, in this implementation environment, the terminal 110 and the server 120 are involved, and the terminal 110 and the server 120 are connected through the communication network 130.

[0041] In some embodiments, the terminal 110 has a data acquisition function for acquiring data that needs to be analyzed by the neural network structure. For example, the terminal acquires at least one of various types of data such as image data, signal data, and text data as the data to be analyzed.

[0042] Optionally, the terminal 110 sends the data to the server 120 through the communication network. The neural network structure and the dimension transpose chip are deployed in the server 120. The dimension transpose chip can perform a dimension transpose process on the tensor data obtained by converting the data during the process of processing the data through the neural network structure. The tensor data is a feature representation expressed in the first dimension sequence generated during the operation of the neural network structure.

[0043] In some embodiments, the dimension transpose chip includes an information acquisition module for acquiring dimension transpose information corresponding to the tensor data.

[0044] Among them, the dimension transpose information is used to indicate transposing the first dimension sequence to the second dimension sequence, that is, the dimension transpose information is used to indicate the dimension transpose process of the tensor data of the first dimension sequence through the dimension transpose information. In addition, the information acquisition module will also send the dimension transpose information to the parameter acquisition module.

[0045] In some embodiments, the dimension transpose chip also includes a parameter acquisition module for receiving the dimension transpose information. In addition, the parameter acquisition module is also used to obtain dimension transpose parameters based on the dimension difference between the first dimension sequence and the second dimension sequence.

[0046] Among them, the dimension difference includes at least one of the dimension quantity difference and the dimension sequence difference. The dimension quantity difference is used to express the change in the number of dimensions, and the dimension sequence difference is used to express the change in the dimension order. The dimension transpose parameters are used to transpose the first dimension sequence to the second dimension sequence.

[0047] In addition, the parameter acquisition module is also used to send the dimension transpose parameters to the dimension transpose module.

[0048] In some embodiments, the dimension transpose module is used to receive the dimension transpose parameters sent by the parameter acquisition module and is also used to acquire the tensor data.

[0049] In addition, the dimension transpose module is also used to transpose the tensor data of the first dimension sequence with the dimension transpose parameters to obtain the tensor data with the second dimension sequence.

[0050] Among them, the tensor data with the second - dimension sequence is used as the input for the neural - network layer in the neural - network structure for analysis based on the second - dimension sequence.

[0051] Schematically, the server 120 performs dimension transposition on the tensor data of the first - dimension sequence based on the deployed dimension - transposition chip, so as to obtain the tensor data of the second - dimension sequence. This tensor data of the second - dimension sequence can be input into the neural - network layer that supports processing the second - dimension sequence, so that the neural - network layer can perform targeted analysis on the tensor data of the second - dimension sequence, avoiding the limitation that different neural - network layers cannot analyze tensor data from different dimension sequences.

[0052] In some embodiments, the server 120, based on the coordination between the neural - network structure and the dimension - transposition chip, analyzes the data through multiple neural - network layers in the neural - network structure and obtains the data - analysis result. Optionally, the server 120 sends the data - analysis result to the terminal 110 through the communication network 130.

[0053] It should be noted that the above - mentioned terminal includes, but is not limited to, mobile terminals such as mobile phones, tablet computers, portable laptop computers, intelligent voice - interaction devices, intelligent home appliances, in - vehicle terminals, etc., and can also be implemented as a desktop computer, etc.; the above - mentioned server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud - computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain - name services, security services, Content Delivery Network (CDN), and big - data and artificial - intelligence platforms.

[0054] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, application programs, and networks within a wide - area network or local - area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management - platform technology, application technology, etc. based on the cloud - computing business model, which can form a resource pool, be used on - demand, and be flexible and convenient.

[0055] In some embodiments, the above - mentioned server can also be implemented as a node in a blockchain system.

[0056] Combined with the above - mentioned noun introduction and application scenarios, the dimension - transposition chip provided by this application is described. Taking the server where the chip is deployed as an example, as Figure 2 shown, the dimension - transposition chip includes an information - acquisition module 210, a parameter - acquisition module 220, and a dimension - transposition module 230. The different modules of the dimension - transposition chip are described separately below.

[0057] (1) Information acquisition module 210

[0058] The information acquisition module 210 is used to obtain the dimension transposition information corresponding to the tensor data.

[0059] Illustratively, the information acquisition module 210 is a module within the dimension transposition chip. When the server invokes the dimension transposition chip for dimension transposition processing, the dimension transposition chip obtains the dimension transposition information corresponding to the tensor data through the information acquisition module 210.

[0060] Among them, the tensor data is a feature representation expressed in the first dimension sequence generated during the operation of the neural network structure.

[0061] Illustratively, the neural network structure refers to the organization and connection manner between neurons (or nodes) in the neural network. As a computational model composed of neurons, the neural network structure can also be called a neural network model, a machine learning model, etc. With the help of the neural network structure, various types of data can be analyzed specifically to perform corresponding machine learning tasks.

[0062] Optionally, data such as image data, text data, and audio data are analyzed through the neural network structure; the neural network structure is a pre-set structure, which includes multiple neural network layers. The neural network layer is the basic component unit in the neural network structure. At least one neuron is included in one neural network layer. The multiple neural network layers that make up the neural network structure include an input layer (Input Layer), a hidden layer (Hidden Layer), an output layer (Output Layer), and other layers. Among them, the other layers can be implemented as a recurrent layer in a recurrent neural network (RNN), a convolutional layer in a convolutional neural network (CNN), etc.

[0063] That is: multiple neural network layers are pre-combined to obtain the neural network structure. The multiple neural network layers in the neural network structure for processing different machine learning tasks may be the same or different, which is not limited here.

[0064] Illustratively, a pre-configured neural network structure is deployed in the server. Taking the image data analyzed by the neural network structure as an example, the image data is input into the input layer of the neural network structure, and the image data is analyzed through multiple network layers within the neural network structure to learn the mapping relationship from the input layer to the output layer so as to perform various machine learning tasks.

[0065] In some embodiments, taking the data to be analyzed by a neural network structure as image data as an example, the image data, as input data, is usually represented as a multi-dimensional array, which can be regarded as tensor data; or, the image data is implemented as an image, and the image data is converted into tensor data through the input layer of the neural network structure, etc.

[0066] Schematically, an image is read from an image file, and the image may be a color image (RGB channels) or a grayscale image; preprocessing operations are performed on the image (at least one of resizing, normalizing, enhancing contrast, or color space conversion, etc.); the image is converted into tensor data. For example, for a color image, it is usually represented by three-dimensional tensor data, and its dimensions are height (Height, H), width (Weight, W), and number of channels (Channel, C). For a grayscale image, it can be represented by two-dimensional tensor data.

[0067] In some embodiments, taking the transfer of tensor data in a neural network structure as an example, multiple neural network layers are connected in succession. The subsequent neural network layer receives the tensor data output by the previous neural network layer and performs corresponding processing on the tensor data to obtain the tensor data input to the next neural network layer. That is: the tensor data is adjusted accordingly based on different neural network layers, and the data transferred between different neural network layers in the neural network structure can all be regarded as tensor data, and the tensor data is the feature representation transferred between different neural network layers.

[0068] Schematically, analyzing the tensor data received by any neural network layer in the neural network structure, the tensor data is a feature representation expressed in a first-dimensional sequence. For example: the tensor data is the data input to the input layer, and the first-dimensional sequence is the dimensional expression form of the image data corresponding to the tensor data, such as expressed as (H, W, C); or, the tensor data is the data input to any intermediate layer, and the first-dimensional sequence is the dimensional expression form determined based on the output of the previous neural network layer, such as expressed as (W, H, C), etc.

[0069] The dimension sequence is usually used to describe the shape of tensor data, representing at least one of the arrangement and size of each dimension in the tensor data. The dimension sequence of the tensor data provides key information about the data structure. For example, taking the tensor data as image data, image data is usually expressed in the form of a four-dimensional tensor, and its shape is (batch_size, height, width, channels), abbreviated as NHWC. batch_size represents the batch size, that is, the number of samples processed simultaneously during training or inference; height represents the height of the image; width represents the width of the image; channels represents the number of channels of the image, which is usually 3 for color images (RGB). The dimension sequence of the tensor data expressed in NHWC can be called (0, 1, 2, 3), so that the dimension transposition can be performed based on this dimension sequence as a benchmark; of course, the dimension sequence of the tensor data expressed in NHWC can also be called (0, 2, 1, 3), so that the dimension transposition can be performed based on this dimension sequence as a benchmark. That is, the setting of the dimension sequence during dimension transposition can be selected, but the benchmark dimension sequence needs to be fixed during one transposition process.

[0070] The meaning of the dimension sequence lies in providing a clear understanding of the data tensor structure, and it is also very useful when designing neural network architectures, adjusting the format of input data, and understanding the output of the model. Tensor operations in deep learning frameworks (such as TensorFlow, PyTorch) usually require knowledge of the dimension information of the data, so understanding the dimension sequence is very important.

[0071] Among them, the dimension transposition information is used to indicate transposing the first dimension sequence into the second dimension sequence. Since the dimension transposition process is used to adjust the dimension sequence, the first dimension sequence and the second dimension sequence are different.

[0072] Optionally, when there is a dimension transposition requirement, the information acquisition module 210 will obtain the dimension transposition information.

[0073] Among them, the dimension transposition requirement is the requirement to perform dimension transposition on the first dimension sequence. Schematically, during the operation of the neural network structure, the dimension transposition chip continuously detects the operation situation in the neural network structure. When the operation situation meets the dimension transposition condition, it is regarded as having a dimension transposition requirement, and the dimension transposition information is obtained through the information acquisition module 210; or, when the operation situation of the neural network structure meets the dimension transposition condition, the dimension transposition chip is called, and this process is regarded as having a dimension transposition requirement, so that the information acquisition module 210 in the dimension transposition chip can obtain the dimension transposition information.

[0074] Schematically, the operating condition reflects the neural network layer to be processed next during the operation of the current neural network structure for the tensor data. If the next neural network layer to be processed does not support processing the tensor data of the first - dimension sequence but supports processing the tensor data of the second - dimension sequence, it is considered that the operating condition meets the dimension transposition condition and there is a dimension transposition requirement.

[0075] Alternatively, if the current neural network layer for processing the tensor data is a neural network layer for performing dimension transposition on the tensor data, it is considered that the operating condition meets the dimension transposition condition and there is a dimension transposition requirement, etc.

[0076] In some embodiments, the dimension transposition information is automatically generated information.

[0077] Schematically, when the dimension transposition condition is met, the neural network structure automatically generates dimension transposition information for instructing the dimension transposition chip to perform dimension transposition processing; or, the dimension transposition chip automatically generates dimension transposition information to prompt its operation when the dimension transposition condition is met, etc. The dimension transposition information includes information representing the first - dimension sequence and information representing the next neural network layer; or, the dimension transposition information is the transposition information used by the current neural network layer for dimension transposition processing, etc.

[0078] In some embodiments, the dimension transposition information is manually input information.

[0079] Schematically, before or during the operation of the neural network structure, information for performing targeted tensor transposition processing on the tensor data can be manually input as the dimension transposition information, etc.

[0080] It should be noted that the above are only schematic examples, and the embodiments of the present application are not limited thereto.

[0081] Among them, the information acquisition module 210 will also send the dimension transposition information to the parameter acquisition module 220.

[0082] (2) Parameter acquisition module 220

[0083] Schematically, the parameter acquisition module 220 is another module within the dimension transposition chip. When the server invokes the dimension transposition chip to perform dimension transposition processing, the dimension transposition chip determines the dimension transposition parameters used for dimension transposing the tensor data through the parameter acquisition module 220.

[0084] The parameter acquisition module 220 is used to receive the dimension transposition information, that is, to receive the dimension transposition information sent by the information acquisition module 210.

[0085] Among them, the dimension transposition information is used to describe the information for converting the first - dimension sequence into the second - dimension sequence; therefore, the sequence change situation between the first - dimension sequence and the second - dimension sequence is contained in the dimension transposition information.

[0086] The parameter acquisition module 220 is further configured to acquire dimension transposition parameters based on the dimension difference between the first - dimension sequence and the second - dimension sequence.

[0087] Among them, the dimension difference includes at least one of the dimension - number difference and the dimension - sequence difference; the dimension - number difference is used to express the change in the number of dimensions, and the dimension - sequence difference is used to express the change in the order of dimensions.

[0088] Illustratively, if the first - dimension sequence is (0, 1, 2, 3) and the second - dimension sequence is (0, 1, 2, 3, 4), then the dimension difference between the first - dimension sequence and the second - dimension sequence is realized as a dimension - number difference, representing that there is a change in the number of dimensions between the first - dimension sequence and the second - dimension sequence, changing from four - dimensional (four dimensions; 4 dimensions, 4D) to five - dimensional (five dimensions; 5 dimensions, 5D); if the first - dimension sequence is (0, 1, 2, 3) and the second - dimension sequence is (0, 2, 1, 3), then the dimension difference between the first - dimension sequence and the second - dimension sequence is realized as a dimension - sequence difference, representing that there is a change in the order of dimensions between the first - dimension sequence and the second - dimension sequence, where the order between the second dimension "1" and the third dimension "2" is reversed; if the first - dimension sequence is (0, 1, 2, 3) and the second - dimension sequence is (0, 1, 3, 2, 4), then the dimension difference between the first - dimension sequence and the second - dimension sequence is realized as a dimension - number difference and a dimension - sequence difference, representing that there are not only changes in the order of dimensions but also changes in the number of dimensions between the first - dimension sequence and the second - dimension sequence, etc.

[0089] Among them, the change in the number of dimensions and / or the change in the order of dimensions are usually caused by data processing, conversion, or network - architecture adjustment. Illustratively, a brief example of the dimension - number difference and the dimension - sequence difference is given.

[0090] 1. Dimension - number difference

[0091] 1.1 Feature extraction: In a neural - network structure, if after passing through a convolutional layer (i.e., a neural - network layer) such as a Convolutional Neural Network (CNN), the number of dimensions of the feature map (tensor data) usually changes. This is because each convolutional layer can change the size of the feature map.

[0092] 1.2 Pooling layer operation: The pooling layer is usually used to reduce the spatial dimension of the feature map by taking the maximum or average value of the local area to reduce the dimension, and this operation will change the dimension of the tensor data.

[0093] 1.3 Fully connected layer: At the end of the neural network, the fully connected layer may convert the high-dimensional feature map into a one-dimensional vector to reduce the dimension.

[0094] 2 Dimension sequence differences

[0095] 2.1 Network architecture design: Different neural network architectures may require different input data shapes. For example, many Convolutional Neural Networks (CNNs) expect the dimension sequence of the input data to be (batch_size, height, width, channels).

[0096] 2.2 Cross-frame compatibility: A neural network structure usually adopts one deep learning framework, and multiple deep learning frameworks can also be integrated into one neural network structure; different deep learning frameworks may have different requirements for the dimension sequence of tensor data. If it is necessary to apply a neural network structure with multiple deep learning frameworks or a combination of multiple different deep learning frameworks, then the dimension sequence of the tensor data needs to be adjusted to meet the requirements of different deep learning frameworks.

[0097] 2.3 Task requirements: Different deep learning tasks may require tensor data of different shapes. For example, image classification tasks and object detection tasks usually require tensor data of different shapes, where different shapes refer to the difference in the row and column distribution of the matrix, representing the arrangement of different dimension sequences.

[0098] 2.4 Model fusion: In some cases, it may be necessary to fuse or connect multiple neural network structures, which may involve adjusting the dimension sequence of the input data to match the expected inputs corresponding to the multiple neural network structures respectively.

[0099] It should be noted that the above are only illustrative examples, and the embodiments of the present application are not limited thereto.

[0100] Among them, the dimension transposition parameter is used to transpose the first dimension sequence to the second dimension sequence.

[0101] Schematically, the dimension transposition parameter is the parameter for performing dimension transposition. By applying the dimension transposition parameter, the first dimension sequence can be transposed into the second dimension sequence, so as to adapt to different neural network layers in the neural network structure.

[0102] The parameter acquisition module 220 is further configured to send the dimension transposition parameter to the dimension transposition module 230.

[0103] (3) Dimension Transposition Module 230

[0104] Schematically, the dimension transposition module 230 is another module within the dimension transposition chip. When the server invokes the dimension transposition chip for dimension transposition processing, the dimension transposition chip performs the dimension transposition process on the tensor data with the dimension transposition parameters through the dimension transposition module 230.

[0105] The dimension transposition module 230 is used to receive the dimension transposition parameters, that is, the dimension transposition parameters sent by the parameter acquisition module 220.

[0106] Among them, the dimension transposition information is used to convert the first dimension sequence into the second dimension sequence; therefore, the dimension transposition parameters contain the transposition information for transposing the first dimension sequence into the second dimension sequence.

[0107] The dimension transposition module 230 is also used to obtain the tensor data.

[0108] Among them, the tensor data is the data used to perform the dimension transposition process. Optionally, the process of obtaining the tensor data is performed through the dimension transposition module 230 so that the dimension transposition module 230 can perform the dimension transposition process on the tensor data based on the dimension transposition parameters.

[0109] The dimension transposition module 230 is also used to perform dimension transposition on the tensor data of the first dimension sequence with the dimension transposition parameters to obtain the tensor data with the second dimension sequence.

[0110] Schematically, matrix transformation is performed on the tensor data of the first dimension sequence with the dimension transposition parameters to implement the dimension transposition process and obtain the tensor data with the second dimension sequence.

[0111] Among them, the tensor data with the second dimension sequence is used to input into the neural network layer in the neural network structure for analysis based on the second dimension sequence.

[0112] Schematically, among the multiple neural network layers in the neural network structure, there is a neural network layer for analysis based on the second dimension sequence. After obtaining the tensor data with the second dimension sequence based on dimension transposition, the tensor data with the second dimension sequence is input into this neural network layer for analysis.

[0113] It should be noted that the above is only a schematic example, and the embodiments of the present application are not limited thereto.

[0114] In summary, the dimension transposition information of the tensor data is obtained by means of the information acquisition module in the dimension transposition chip, and the dimension transposition parameters are obtained by the parameter acquisition module according to the dimension difference between the first dimension sequence and the second dimension sequence represented by the dimension transposition information. Then, the tensor data is dimension-transposed with the dimension transposition parameters under the dimension transposition module to obtain the tensor data with the second dimension sequence, so as to be input into the neural network layer in the neural network structure for analysis based on the second dimension sequence. By integrating the dimension transposition chip to execute the dimension transposition process, the dimension sequence of the tensor data can be specifically adjusted based on multiple neural network layers in the neural network structure that support different dimensions, so as to obtain tensor data with different dimensions that is more convenient to accurately input into the corresponding neural network layer for data processing. In addition, the dimension difference situation including at least one of the dimension number difference and the dimension sequence difference is fully considered, so that the dimension transposition parameters can be flexibly obtained based on the dimension difference, avoiding the limitation problem of presetting the dimension transposition parameters.

[0115] In an optional embodiment, when obtaining the dimension transposition parameters through the parameter acquisition module, the dimension transposition information can be differentially processed according to the content of the dimension transposition parameters. Illustratively, when the dimension transposition parameter is implemented as a sequence transformation value, the sequence dimension number represented by the sequence transformation value can be determined, and then the dimension conversion parameter can be obtained based on the dimension number difference between the sequence dimension number and the first dimension number corresponding to the first dimension sequence. Among them, the sequence transformation value can reflect the dimension number difference and may also reflect the dimension sequence difference. Illustratively, the parameter acquisition module is also used to perform the following content.

[0116] In some embodiments, the parameter acquisition module analyzes the dimension transposition information.

[0117] Illustratively, after receiving the dimension transposition information, the parameter acquisition module analyzes the dimension transposition information to determine the information content included in the dimension transposition information.

[0118] In some embodiments, in response to the dimension transposition information being a sequence transformation value, the sequence dimension number represented by the sequence transformation value is determined.

[0119] Among them, the sequence transformation value is used to transpose the tensor data of the first dimension sequence into the tensor data of the second dimension sequence.

[0120] Illustratively, the sequence transformation value can also be called the transpose axes information, which generally refers to the exchange of axes for tensor data or matrices. In the neural network structure, the operation of applying the transpose axes information (axis transposition operation) is often used to adjust the shape of the tensor data to adapt to different neural network structures or task requirements.

[0121] For example, the tensor data is a two-dimensional matrix with a shape of (m, n), where m represents the number of rows and n represents the number of columns; performing an axis transpose operation on this tensor data means swapping the number of rows and columns of the tensor data to generate a new matrix with a shape of (n, m).

[0122] In deep learning, the axis transpose operation can be performed on high-dimensional tensor data in a similar way. For example, there is a tensor data with a shape of (0, 1, 2, 3) and the transpose axis information is (0, 1, 3, 2). Here, (0, 1, 2, 3) can be regarded as the first-dimensional sequence of the tensor data, expressing four dimensions, that is, the number of the first dimension is 4; (0, 1, 3, 2) can be regarded as the sequence transformation value and also expresses four dimensions, that is, the number of sequence dimensions is also 4.

[0123] In some embodiments, based on the dimension number difference between the number of sequence dimensions and the number of the first dimension corresponding to the first-dimensional sequence, dimension transpose parameters are obtained.

[0124] Schematically, after determining the number of sequence dimensions and the number of the first dimension, compare the numerical sizes between the number of sequence dimensions and the number of the first dimension, that is, determine the dimension number difference between the number of sequence dimensions and the number of the first dimension. For example, it is determined that the number of sequence dimensions is less than the number of the first dimension; or, the number of sequence dimensions is equal to the number of the first dimension, or, the number of sequence dimensions is greater than the number of the first dimension.

[0125] In an alternative embodiment, the parameter acquisition module includes an information analysis unit and a sequence processing unit.

[0126] Schematically, the above Figure 2 shown dimension transpose chip can also be as Figure 3 shown, where the parameter acquisition module 310 includes an information analysis unit 311 and a sequence processing unit 312.

[0127] In some embodiments, the information analysis unit 311 is configured to generate a sequence processing request under the condition that the dimension number difference represents that the number of sequence dimensions is different from the number of the first dimension.

[0128] Schematically, the parameter acquisition module 310 numerically compares the number of sequence dimensions and the number of the first dimension through the information analysis unit 311. When it is determined that the number of sequence dimensions is different from the number of the first dimension, a sequence processing request will be generated.

[0129] Among them, the sequence processing request is used to perform sequence adjustment on the sequence transformation value.

[0130] Optionally, within a neural network structure, a transpose operation for a specific number of dimensions (such as one of 3D, 4D, 5D, etc.) is typically supported. For example, it supports conversion from one 4D (0, 1, 2, 3) to another 4D (0, 3, 1, 2), etc. When performing a transpose operation under this neural network structure, when the number of sequence dimensions is different from the number of the first dimension, the sequence transformation value needs to be adjusted to enable the transpose process under the number of the first dimension when performing dimension transposition based on the sequence transformation value. For example, when the number of the first dimension is 4 and the number of sequence dimensions is 5, the sequence transformation value needs to be adjusted to perform the transpose operation in 4D form.

[0131] In some embodiments, the information analysis unit 311 is further configured to, in response to the number of sequence dimensions being equal to the number of the first dimension, use the sequence transformation value as the dimension transpose parameter.

[0132] Illustratively, when the information analysis unit 311 determines through comparison that the number of sequence dimensions is the same as the number of the first dimension, it can perform a transpose operation on the tensor data corresponding to the first dimension sequence using the sequence transformation value, that is, use the sequence transformation value as the dimension transpose parameter.

[0133] For example: If the first dimension sequence is (0, 1, 2, 3) and the sequence transformation value is (0, 2, 3, 1), based on both the number of the first dimension and the number of sequence dimensions being 4, the sequence transformation value can be used as the dimension transpose parameter to process the tensor data of the first dimension sequence, and the resulting second dimension sequence is (0, 2, 3, 1).

[0134] Or, if the first dimension sequence is (0, 1, 3, 2) and the sequence transformation value is (0, 2, 3, 1), based on both the number of the first dimension and the number of sequence dimensions being 4, the sequence transformation value can be used as the dimension transpose parameter to process the tensor data of the first dimension sequence, and the resulting second dimension sequence is (0, 3, 1, 2). Here, the first number of the sequence transformation value is 0, indicating that the first number in the final arrangement should be the 0th number in the first dimension sequence, that is, 0; the second number of the sequence transformation value is 2, indicating that the second number in the final arrangement should be the 2nd number in the first dimension sequence, that is, 3; the third number of the sequence transformation value is 1, indicating that the third number in the final arrangement should be the 1st number in the first dimension sequence, that is, 1; the fourth number of the sequence transformation value is 3, indicating that the fourth number in the final arrangement should be the 3rd number in the first dimension sequence, that is, 2.

[0135] Among them, as the first - dimension sequence is used to characterize the distribution of each parameter (such as NHWC) in the tensor data, the first - dimension sequence of the tensor data arranged in NHWC can be regarded as (0, 1, 2, 3), and the first - dimension sequence of the tensor data arranged in NCWH can also be regarded as (0, 1, 2, 3), etc. Therefore, the first - dimension sequence can be regarded as a description of the initial arrangement, which is not limited here.

[0136] The above - mentioned sequence transformation value having the same number of dimensions as the tensor data is only a schematic example. The transposed axis information and the tensor data may also have different numbers of dimensions. For example, the tensor data is 4D and the transposed axis information is 5D, which is not limited here.

[0137] The information analysis unit 311 is also used to send a sequence processing request to the sequence processing unit 312.

[0138] Schematically, based on the sending of the sequence processing request to instruct the sequence processing unit 312 to execute the sequence processing process.

[0139] The sequence processing unit 312 is used to receive the sequence processing request, that is, to receive the sequence processing request sent by the information analysis unit 311, so as to perform sequence processing on the sequence transformation value.

[0140] The sequence processing unit 312 is also used to perform sequence processing on the sequence transformation value based on the quantitative relationship between the number of dimensions of the sequence in the sequence processing request and the number of the first dimensions, so as to obtain the dimension transposition parameter.

[0141] Schematically, based on the sequence processing request being sent when the number of dimensions of the sequence is different from the number of the first dimensions, the sequence processing request may include the quantitative relationship between the number of dimensions of the sequence and the number of the first dimensions. For example, when the number of dimensions of the sequence is less than the number of the first dimensions, the sequence processing request includes the identifier 0; or when the number of dimensions of the sequence is greater than the number of the first dimensions, the sequence processing request includes the identifier 1, etc.

[0142] Optionally, according to different quantitative relationships between the number of dimensions of the sequence and the number of the first dimensions, perform differential sequence processing on the sequence transformation value to obtain the dimension transposition parameter for transposing the first - dimension sequence.

[0143] In some embodiments, the sequence processing unit 312 is also used to perform an axis - combining process on the sequence transformation value in response to the sequence processing request indicating that the number of dimensions of the sequence is greater than the number of the first dimensions, so as to obtain the dimension transposition parameter.

[0144] Among them, the axis - combining process is used to combine at least two dimensions in the first - dimension sequence.

[0145] Optionally, the sequence processing request characterizes the comparison situation of the number of dimensions through the sequence identifier carried therein. If the sequence identifier carried therein is 0, it indicates that the sequence processing request indicates that the number of sequence dimensions is greater than the first number of dimensions; or, if the sequence identifier carried therein is 1, it indicates that the sequence processing request indicates that the number of sequence dimensions is greater than the first number of dimensions, etc.

[0146] Schematically, the first number of dimensions is the number of dimensions corresponding to the first-dimensional sequence that can be processed by the neural network structure, representing the number of dimensions supported by the neural network structure for processing; if the number of sequence dimensions is greater than the first number of dimensions, it is impossible to directly process the tensor data of the first-dimensional sequence based on the sequence transformation value. At this time, sequence processing needs to be performed on the sequence transformation value to select at least two dimensions for merging from the values representing each dimension in the sequence transformation value, and obtain the dimension transpose parameter of the first number of dimensions based on the sequence transformation value.

[0147] For example: the sequence transformation value is (0, 1, 3, 2, 4), and its number of sequence dimensions is 5; if the number of dimensions supported by the neural network structure for processing is 4, it is regarded as necessary to perform axis merging processing on the sequence transformation value (0, 1, 3, 2, 4) based on the first number of dimensions 4.

[0148] Among them, it can be assumed that the dimension sequence before dimension transposition is (0, 1, 2, 3, 4) - the number of dimensions is 5, and the sequence transformation value is the transposed sequence dimension (0, 1, 3, 2, 4) - the number of dimensions is 5. The above-mentioned axis merging processing on the sequence transformation value to obtain the dimension transpose parameter for dimension transposition of the first number of dimensions 4 can be regarded as the process of inversely deducing (0, 1, 2, 3, 4) from (0, 1, 3, 2, 4). Among them, (0, 1, 2, 3, 4) is used to obtain (0, 1, 3, 2, 4) through the dimension transpose parameter with the number of dimensions 4.

[0149] As Figure 4 shown, it is a schematic diagram of performing an axis merging operation on the sequence transformation value to obtain the dimension transpose parameter. Among them, the sequence transformation value is (0, 1, 3, 2, 4) with the number of dimensions 5. This process can be implemented through two four-dimensional transformation (4D permute) operations, that is, converting from (0, 1, 2, 3, 4) to (0, 1, 3, 2, 4) through two four-dimensional transformation operations. Each transformation operation uses a dimension transpose parameter, and two dimension transpose parameters will be deduced based on (0, 1, 3, 2, 4).

[0150] The first step: the first four-dimensional transformation 410 (obtaining the first dimension transpose parameter)

[0151] In the first step, first perform an axis merging operation on two dimensions in the sequence transformation value. For example, perform an axis merging operation on dimension 0 and dimension 1 to obtain 4 dimensions supported by the neural network structure, thereby obtaining (new0, 3, 2, 4); based on the number of dimensions being 4, it can be (new0, new2, new1, new3), so as to represent this content in the form of dimensions with continuous numerical values when the number of dimensions is correct; since this process is an inverse process, if it is necessary to restore to (0, 1, 2, 3, 4), or rather, restore to (0, 1, 2, 3), it means that when performing the first four-dimensional transformation 410, it is necessary to swap the positions of dimension 1 and dimension 2 to obtain (0, 2, 1, 3) from (0, 1, 2, 3). Written in the form of the dimension transposition parameter with the number of dimensions being 4, that is: the first dimension transposition parameter obtained by the first four-dimensional transformation 410 is (0, 1, 3, 2), or permute(0, 1, 3, 2).

[0152] Step 2: The second four-dimensional transformation 420

[0153] In the second step, first, it is necessary to perform a restoration operation on the two dimensions that were axis-merged previously. For example, perform a restoration operation on the original dimension 0 and the original dimension 1 to obtain 5 dimensions, expressed as (0, 1, new1, new2, 4); the second 4D permute should further adjust the dimensions to match the final result of the original 5D permute. Since the first four-dimensional transformation 410 can already restore to (0, 1, 2, 3), the second four-dimensional transformation 420 only requires an identity operation, that is: the second dimension transposition parameter obtained by the second four-dimensional transformation 420 is (0, 1, 2, 3), or permute(0, 1, 2, 3).

[0154] It should be noted that the above example of performing an axis merging operation on dimension 0 and dimension 1 is only illustrative, and the embodiments of the present application are not limited thereto.

[0155] In some embodiments, the sequence processing unit 312 is further configured to, in response to a sequence processing request indicating that the sequence dimension number is less than the first dimension number, perform an axis splitting operation on the sequence transformation value to obtain a dimension transposition parameter.

[0156] Among them, the axis splitting operation is used to split at least one dimension in the first dimension sequence.

[0157] Optionally, the sequence processing request characterizes the comparison of the dimension number through the sequence identifier carried therein. If the sequence identifier carried therein is 0, it means that the sequence processing request indicates that the sequence dimension number is less than the first dimension number; or, if the sequence identifier carried therein is 1, it means that the sequence processing request indicates that the sequence dimension number is less than the first dimension number, etc.

[0158] Schematically, the first - dimension quantity is the dimension quantity corresponding to the first - dimension sequence that can be processed by the neural - network structure, representing the dimension quantity supported by the neural - network structure for processing; if the sequence dimension quantity is less than the first - dimension quantity, it is impossible to directly process the tensor data of the first - dimension sequence based on the sequence transformation value. At this time, sequence processing needs to be performed on the sequence transformation value to select at least one dimension for splitting from the values representing each dimension in the sequence transformation value, and obtain the dimension transpose parameter of the first - dimension quantity based on the sequence transformation value.

[0159] For example: The sequence transformation value is (2, 1, 0), and its sequence dimension quantity is 3; if the dimension quantity supported by the neural - network structure for processing is 4, it is regarded as necessary to perform axis - splitting processing on the sequence transformation value (2, 1, 0) based on the first - dimension quantity 4.

[0160] Among them, it can be assumed that the dimension sequence before dimension transpose is (0, 1, 2) - - the dimension quantity is 3, and the sequence transformation value is the transposed sequence dimension (2, 1, 0) - - the dimension quantity is 3. The above - mentioned axis - splitting processing of the sequence transformation value to obtain the dimension transpose parameter for dimension transpose of the first - dimension quantity 4 can be regarded as the process of inversely deducing (0, 1, 2) from (2, 1, 0). Among them, (0, 1, 2) is used to obtain (2, 1, 0) through the dimension transpose parameter with a dimension quantity of 4.

[0161] Schematically, transforming from (0, 1, 2) to (2, 1, 0) means that dimension 0 moves to dimension 2; dimension 1 moves to dimension 1; dimension 2 moves to dimension 0.

[0162] In addition, in order to convert the dimension quantity 3 (3D) into the dimension quantity 4 (4D), an additional dimension needs to be added. This additional dimension (D3) can be a single - element dimension, which does not affect the actual content of the data. For example: Add it to the end as a new dimension, and the data structure becomes 4D - - (0, 1, 2, 3).

[0163] Optionally, two 4D permute operations can be used to achieve the same effect as the original 3D permute (0, 1, 2).

[0164] (1) The first 4D Permute

[0165] In the first 4D permute, part of the effect of the original 3D permute can be simulated. For example, it can be selected to move D2 (dimension 2) to dimension 0 while keeping other dimensions relatively unchanged. In this way, the transpose axis (dimension transpose parameter) of the first 4D permute may be:

[0166] D2 (The original D2 is moved here); D0 (the original D0); D1 (the original D1); D3 (new dimension, remains unchanged); that is, the transposed axes of the first 4D permute are (2, 0, 1, 3).

[0167] (2) The second 4D Permute

[0168] The second 4D permute needs to further adjust the order to ensure that the final permutation order is the same as the original 3D permute. Based on the above, it is necessary to swap the first and third dimensions (now D0 and D1), while keeping the second dimension (now D2) and the fourth dimension (D3) unchanged. Therefore, the transposed axes (dimension transpose parameters) of the second 4D permute are:

[0169] D1 (moved from the third position in the first 4D permute to here); D2 (remains unchanged); D0 (moved from the first position in the first 4D permute to here); D3 (remains unchanged); that is, the transposed axes of the second 4D permute are (1, 2, 0, 3).

[0170] That is: Based on the above analysis, the two dimension transpose parameters are (2, 0, 1, 3) and (1, 2, 0, 3) respectively; that is: after processing (0, 1, 2) through (2, 0, 1, 3), and then transposing the processing result of (0, 1, 2) through (1, 2, 0, 3), (2, 1, 0) can be obtained.

[0171] It should be noted that the above process of adding dimension 3 can be regarded as splitting axis 2, and the splitting process needs to not affect the expression of the original sequence transformation value; the above is only an illustrative example, and the embodiments of the present application are not limited thereto.

[0172] In summary, by means of a dimension transpose chip to integrally execute the dimension transpose process, so as to targetedly adjust the dimension sequence of tensor data based on multiple neural network layers supporting different dimensions in a neural network structure, thereby obtaining tensor data of different dimensions that is more convenient for accurate input into the corresponding neural network layer for data processing; in addition, fully considering the dimension difference situation including at least one of the dimension number difference and the dimension sequence difference, so that the dimension transpose parameters can be flexibly obtained based on the dimension difference, avoiding the limitation problem of presetting dimension transpose parameters.

[0173] In the embodiments of the present application, it is introduced that under the condition that the dimension transposition information is a sequence transformation value, the sequence transformation value can be selectively processed according to the relationship between the number of sequence dimensions corresponding to the sequence transformation value and the number of the first dimensions, so that based on the sequence transformation value, a dimension transposition parameter with the same number of dimensions as the number of the first dimensions can be obtained, which is convenient for performing a more standardized dimension transposition process on the tensor data based on the first dimension sequence and the dimension transposition parameter under the neural network structure that supports the single dimension condition, and improving the acquisition efficiency and acquisition stability of the tensor data of the second dimension sequence.

[0174] In an alternative embodiment, when obtaining the dimension transposition parameter through the parameter acquisition module, the dimension transposition information can be differentially processed according to the content of the dimension transposition parameter. Illustratively, when the dimension transposition parameter is implemented as a first data format and a second data format, the dimension transposition parameter can be obtained according to the first data format and the second data format. Among them, the first data format corresponds to the first dimension sequence, and the second data format corresponds to the second dimension sequence; at least the dimension sequence difference can be reflected through the first data format and the second data format, and the dimension number difference may also be reflected. Illustratively, the parameter acquisition module is further configured to perform the following content.

[0175] In some embodiments, the parameter acquisition module analyzes the dimension transposition information.

[0176] Illustratively, after receiving the dimension transposition information, the parameter acquisition module analyzes the dimension transposition information to determine the information content included in the dimension transposition information.

[0177] In some embodiments, in response to the first data format corresponding to the first dimension sequence and the second data format corresponding to the second dimension sequence being included in the dimension transposition information, the dimension transposition parameter is obtained based on the first data format and the second data format.

[0178] Illustratively, the data format is a way to describe the arrangement of a multi-dimensional array (such as the tensor data flowing in the neural network structure) in memory, and can be regarded as the format of the data supported by the neural network layer in the neural network structure. Common data formats include NCHW, NHWC, KhKwCiCo, and B0B1HW, etc.

[0179] The 4 parameters involved in both NCHW and NHWC include the batch processing size N for describing the feature map, the number of channels C, the height H of the feature map, and the width W of the feature map. That is: the data format including these 4 parameters N, H, W, and C is an expression based on the feature map.

[0180] The four parameters involved in KhKwCiCo include the height of the convolutional kernel (Kernel Height, Kh), the width of the convolutional kernel (Kernel Width, Kw), the number of input channels (Input Channels, Ci), and the number of output channels (Output Channels, Co). Among them, Kh represents the vertical size of the convolutional kernel used for sliding in the convolutional operation; Kw represents the horizontal size of the convolutional kernel used for sliding in the convolutional operation; Ci represents the number of channels of the input data. For example, the number of input channels of a color image is the number of color channels in the image. For example, the number of channels of an RGB image is 3; Co represents the number of output channels of the result of the convolutional operation, that is, the number of convolutional kernels used in the convolutional layer. Each convolutional kernel corresponds to an output channel. That is: The data format including these four parameters Kh, Kw, Ci, and Co is an expression based on the convolutional layer, etc.

[0181] The four parameters involved in B0B1HW include a batch dimension (Batch Dimension, B0), another batch dimension (Batch Dimension, Part 2, B1), a height dimension (Height Dimension, H), and a width dimension (Width Dimension, W). Among them, B0 represents the number of sample batches used in training, and this dimension corresponds to the number of batches in the input data; B1 is that there may be two batch dimensions in the representation method, which may be related to certain models or operations. If there are two batch dimensions, it may represent batch processing corresponding to two different aspects, or it is to support certain special operations; H represents the size of the data or features in the vertical direction. In a convolutional neural network, it is usually the height of the image; W represents the size of the data or features in the horizontal direction. In a convolutional neural network, it is usually the width of the image. That is: This data format is common in deep learning frameworks, especially in convolutional neural networks.

[0182] Schematically, the first data format and the second data format are two different data formats. The first data format is used to represent the data format corresponding to the first-dimensional sequence, and the second data format is used to represent the data format corresponding to the second-dimensional sequence. Before transposing the first-dimensional sequence, it is first determined that the first-dimensional sequence needs to be transposed according to the difference between the first data format and the second data format.

[0183] For example: In a neural network structure, neural network layer 1 and neural network layer 2 are connected in sequence. If the first data format supported by neural network layer 1 is different from the first data format supported by neural network layer 2, then the first-dimensional sequence output by neural network layer 1 needs to be transposed to obtain a second-dimensional sequence, and then neural network layer 2 processes it based on the second-dimensional sequence, etc.

[0184] In an optional embodiment, the parameter acquisition module includes an information analysis unit and a parameter call unit.

[0185] Schematically, the above Figure 2 The dimension transpose chip shown can also be as Figure 5 shown, where the parameter acquisition module 510 includes an information analysis unit 511 and a parameter call unit 512.

[0186] The information analysis unit 511 is used to analyze the dimension transpose information.

[0187] The information analysis unit 511 is further configured to send a parameter acquisition request to the parameter call unit 512 in response to the dimension transpose information including a first data format and a second data format.

[0188] Optionally, when the information analysis unit 511 performs data interaction between two neural network layers in the neural network structure, it analyzes the tensor data output by the previous neural network layer, determines the data format corresponding to the tensor data as the first data format, and based on the tensor data corresponding to the first dimension sequence, so it is called the first dimension sequence corresponding to the first data format; in addition, the information analysis unit 511 also determines the second data format based on the data format supported by the neural network layer during subsequent data processing.

[0189] Schematically, if the dimension transpose information obtained by the information analysis unit 511 is implemented as a first data format and a second data format, a parameter acquisition request can be sent to the parameter call unit 512.

[0190] Among them, the parameter acquisition request is used to request to obtain the dimension transpose parameters for transposing the first dimension sequence corresponding to the tensor data.

[0191] The parameter call unit 512 is used to store a format transpose parameter table.

[0192] Among them, the format transpose parameter table stores transpose parameters corresponding to multiple data format pairs. The data format pairs include a start data format and an end data format. The transpose parameters are used to transpose from the start data format to the end data format, and the data format is used to express the layout of the dimension sequence.

[0193] Optionally, the format transpose parameter table is a table pre-analyzed based on the transpose situation between data formats, which includes transpose information for dimension transpose between any two data formats. By querying the format transpose parameter table, one data format can be converted into another data format, and this process is achieved by transposing the dimension sequence corresponding to the data format.

[0194] Schematically, a data format pair A is implemented as Data Format 1 - Data Format 2, and this data format pair A corresponds to a transpose parameter a. Then, if it is necessary to transpose the first - dimension sequence of Data Format 1, the first - dimension sequence can be transposed through the transpose parameter a to obtain the dimension sequence corresponding to Data Format 2.

[0195] The parameter call unit 512 is also used to receive a parameter acquisition request. That is: the parameter call unit 512 receives the parameter acquisition request sent by the information analysis unit 511.

[0196] The parameter call unit 512 is also used to query the format transpose parameter table based on the first data format and the second data format to obtain the dimension transpose parameter.

[0197] Schematically, based on the format transpose parameter, the transpose parameters corresponding to multiple data format pairs are stored. Therefore, after determining the first data format and the second data format, the dimension transpose parameter can be determined by querying the format transpose parameter table.

[0198] In some embodiments, the parameter call unit 512 is also used to query the starting data format in the format transpose parameter table in the first data format to determine multiple candidate parameters corresponding to the first data format.

[0199] Among them, the multiple candidate parameters are the parameters corresponding to the first data format. The starting data format is used to represent the data format before dimension transposition. Correspondingly, the starting data format corresponds to an ending data format, and the starting data format and the ending data format together form a data format pair. The multiple candidate parameters correspond one - to - one to the multiple ending data formats that form data format pairs with the first data format.

[0200] As Figure 6 shown, it is a schematic example of the format transpose parameter table, which includes 9 data formats.

[0201] The format transposition parameters are composed of the parameter table rows corresponding to the starting data format 610 and the parameter table columns corresponding to the ending data format 620. The starting data format 610 includes 9 data formats, and the ending data format 620 also includes 9 data formats. Among the 9 data formats, NHWC, NCHW, and NC1HWC0 represent data formats of different dimensional sequences composed of the same parameters; KhKwCiCo, KwCiKhCo, Co1Ci1KhKwCi0Co0, and Co1KwCiKhCo0 represent data formats of different dimensional sequences composed of the same parameters. Co1 and Co0 are the contents obtained by splitting Co into two dimensions, and Co1 and Co0 are the contents obtained by splitting Ci into two dimensions; B0B1HW and B0B1W1HW0 represent data formats of different dimensional sequences composed of the same parameters. W1 and W0 are the contents obtained by splitting W into two dimensions, etc. Here, the 9 data formats are only illustrative examples.

[0202] Under the condition of transposing within different dimensional sequences with the same parameters, the format transposition parameter table as shown in Figure 6 is obtained. Thus, based on any two data formats, the dimensional transposition parameters can be queried from the Figure 6 shown format transposition parameter table.

[0203] Illustratively, when querying the format transposition parameter table with the first data format, other starting data formats can be filtered out first based on the first data format before dimensional transposition to avoid the problem of incorrect dimensional transposition direction, thereby obtaining multiple candidate parameters corresponding to the first data format. The multiple candidate parameters correspond one by one to the ending data formats that form data format pairs with other starting data formats.

[0204] In some embodiments, the parameter calling unit 512 is further configured to query multiple candidate parameters with the second data format to obtain the dimensional transposition parameters.

[0205] Among them, the dimensional transposition parameter is the transposition parameter of the data format pair composed of the first data format and the second data format.

[0206] Illustratively, based on the need to transpose the first dimensional sequence of the first data format into the second dimensional sequence of the second data format, the ending data formats corresponding to the multiple candidate parameters are screened with the second data format, thereby obtaining the dimensional transposition parameter jointly corresponding to the first data format and the second data format, that is: the dimensional transposition parameter corresponds to the data format pair of "first data format - second data format".

[0207] In an alternative embodiment, the parameter calling unit 512 is further configured to, in response to the data format pair composed of the first data format and the second data format not being found in the format transposition parameter table, determine the number of dimensions of the first dimension sequence corresponding to the first data format, and determine the number of dimensions of the second dimension sequence corresponding to the second data format; in response to the number of dimensions of the first dimension sequence being different from the number of dimensions of the second dimension sequence, sequence process the second dimension sequence based on the first dimension sequence to obtain the dimension transposition parameter.

[0208] Illustratively, the number of dimensions of the first dimension sequence is the number of dimensions in the first dimension sequence, and the number of dimensions of the second dimension sequence is the number of dimensions in the second dimension sequence. If the number of dimensions of the first dimension sequence is the same as the number of dimensions of the second dimension sequence, the tensor data of the first dimension sequence can be dimension transposed by the second dimension sequence based on the first dimension sequence, that is, the second dimension sequence is used as the dimension transposition parameter.

[0209] If the number of dimensions of the first dimension sequence is different from the number of dimensions of the second dimension sequence, the second dimension sequence can be sequence processed based on the first dimension sequence.

[0210] Illustratively, based on the first dimension sequence, that is, when performing sequence processing on the second dimension sequence, reverse infer how to obtain the first dimension sequence from the second dimension sequence to determine the dimension transposition parameter used when converting the first dimension sequence into the second dimension sequence during the reverse inference process.

[0211] Optionally, the parameter calling unit 512 is further configured to, in response to the number of dimensions of the second dimension sequence being greater than the number of dimensions of the first dimension sequence, perform an axis merging process on the second dimension sequence to obtain the dimension transposition parameter.

[0212] Among them, the axis merging process is used to merge at least two dimensions in the second dimension sequence.

[0213] Optionally, the parameter calling unit 512 is further configured to, in response to the number of dimensions of the second dimension sequence being less than the number of dimensions of the first dimension sequence, perform an axis splitting process on the second dimension sequence to obtain the dimension transposition parameter.

[0214] Among them, the axis splitting process is used to split at least one dimension in the second dimension sequence.

[0215] In an alternative embodiment, after querying the format transposition parameter table, the above axis splitting process or axis merging process can be combined to more comprehensively transform the dimension sequences corresponding to different data formats; or, after performing the axis splitting process or axis merging process on the dimension sequence corresponding to the data format, query the format transposition parameter table, so as to cover more dimension transposition situations.

[0216] It should be noted that the above are only illustrative examples, and the embodiments of the present application are not limited thereto.

[0217] In summary, by means of the dimension transpose chip, the dimension transpose process is integrated to specifically adjust the dimension sequence of tensor data based on multiple neural network layers that support different dimensions in the neural network structure, so as to obtain tensor data with different dimensions that is more convenient for accurately inputting into the corresponding neural network layer for data processing. In addition, the dimension difference situation including at least one of the dimension quantity difference and the dimension sequence difference is fully considered, so that the dimension transpose parameters can be flexibly obtained based on the dimension difference, avoiding the limitation problem of presetting dimension transpose parameters.

[0218] In the embodiment of the present application, it is introduced that under the condition that the dimension transpose information is in the first data format and the second data format, the format transpose parameter table can be queried according to the first data format and the second data format, so as to find the dimension transpose parameters corresponding to the tensor data and the second data format, and use the pre-computed format transpose parameter table to more efficiently query the dimension transpose parameters, facilitating a more efficient dimension transpose process for the tensor data and improving the efficiency of obtaining tensor data with the second dimension sequence.

[0219] In an optional embodiment, in addition to transposing the tensor data with the first dimension sequence through the dimension transpose module to obtain the tensor data with the second dimension sequence, it is also possible to determine whether it is necessary to perform data padding operations on the specified dimension in the tensor data with the second dimension sequence according to the location of the target address for storing the tensor data with the second dimension sequence. Schematically, as Figure 7 shown, the dimension transpose module 710 further includes a dimension transpose unit 711, an address determination unit 711, and a dimension padding unit 713.

[0220] The dimension transpose module 710 is also used to determine the target address.

[0221] Schematically, the address determination unit 711 is used to determine the target address. Among them, the target address is for storing the tensor data with the second dimension sequence.

[0222] Schematically, the address corresponding to the target address is the source address, and the source address is used to store the tensor data with the first dimension sequence; the target address is used to store the tensor data with the second dimension sequence obtained after dimension conversion of the tensor data.

[0223] The dimension transpose module 710 is also used to perform data padding adjustment on the specified dimension in the tensor data with the second dimension sequence based on the location of the target address to obtain the adjusted tensor data.

[0224] Schematically, the dimension padding unit 713 in the dimension transpose module 710 performs data padding adjustment on the specified dimension in the tensor data with the second dimension sequence.

[0225] Optionally, during the process of data processing through a neural network structure, there may be an operation to pad dimensions, which is usually related to batch processing and hardware optimization.

[0226] Schematically, in a neural network structure, multiple samples are usually processed simultaneously using batch processing. To efficiently utilize hardware (such as a Graphics Processing Unit, GPU), it is necessary to ensure that all samples in each batch have the same dimensions. If the dimensions of the samples in the input data are different, in order to perform batch processing operations, it is necessary to pad the samples to make them have the same dimensions.

[0227] Schematically, components such as GPUs are usually more efficient in performing matrix operations, and matrix operations require the input data to have the same dimensions. By padding the data, hardware can be more effectively utilized for parallel computing, improving the speed of training and inference.

[0228] Schematically, there may be some hardware requirements for data to be aligned in memory according to certain rules, which can improve the efficiency of memory reading and writing. By padding the dimensions, it can be ensured that the layout of the data stored in memory conforms to the hardware alignment rules, etc.

[0229] In some embodiments, based on the location of the target address, it is determined whether to perform padding adjustment on a specified dimension in the second dimension sequence to obtain the adjusted tensor data. The specified dimension is a pre-set dimension, such as the channel dimension, etc.

[0230] Schematically, if it is necessary to perform padding adjustment on a specified dimension in the second dimension sequence, then the specified dimension is determined from the tensor data, and then the data of the specified dimension in the tensor data is padded to obtain the adjusted tensor data.

[0231] In some embodiments, the dimension transpose module 710 is further configured to, in response to the target address being located in the on-chip memory, perform padding adjustment on a specified dimension in the second dimension sequence to obtain the adjusted tensor data.

[0232] Among them, the on-chip memory is the memory deployed in the computing device used to perform the dimension transpose process.

[0233] Schematically, both the source address and the target address are located in the on-chip memory; if it is determined that the target address is located in the on-chip memory, then padding adjustment is performed on the tensor data.

[0234] In some embodiments, the dimension transpose module 710 is further configured to, in response to a preset condition being met between the first dimension sequence and the second dimension sequence, store the tensor data with the second dimension sequence in the target address.

[0235] Among them, the preset condition is a condition that limits the dimension conversion between the first preset dimension in the first dimension sequence and the second preset dimension in the second dimension sequence.

[0236] Schematically, when the preset condition preset between the first dimension sequence and the second dimension sequence is met, there is no need to perform a data padding operation on the specified dimension of the tensor data, and the tensor data with the second dimension sequence can be directly stored in the target address.

[0237] Optionally, the preset condition can be implemented as at least one of the following situations.

[0238] Preset condition 1:

[0239] (1) The source dimension 2 is mapped to the target dimension 3; that is: the second dimension in the first dimension sequence is mapped to the third dimension in the second dimension sequence;

[0240] (2) The source DTE_SRC_TILE_DIM2_SIZE == 3, 5, 7; that is: this content is used to compare whether the size of the tensor data in the first dimension sequence in the second dimension is 3 or 5 or 7;

[0241] (3) DTE_ENHANCE_PERMUTE_MODE = 1.

[0242] Preset condition 2:

[0243] (1) The source dimension 2 is mapped to the target dimension 2; that is: the second dimension in the first dimension sequence is mapped to the second dimension in the second dimension sequence, and the second dimension remains unchanged;

[0244] (2) The source dimension 3 is mapped to the target dimension 3; that is: the third dimension in the first dimension sequence is mapped to the third dimension in the second dimension sequence, and the third dimension remains unchanged;

[0245] (3) DTE_ENHANCE_PERMUTE_MODE = 1.

[0246] Among them, DTE_ENHANCE_PERMUTE_MODE = 1 means starting the enhancement mode, that is, not padding the specified dimension in the tensor data.

[0247] The dimension transpose module 710 is also used to send the adjusted tensor data to the target address for storage.

[0248] Schematically, the dimension transposition module 710 sends the adjusted tensor data to the target address, so as to store the adjusted tensor data through the target address; in addition, if it is not necessary to pad the specified dimension in the tensor data, the dimension transposition module 710 sends the tensor data of the second dimension sequence to the target address, so as to store the tensor data of the second dimension sequence through the target address, etc.

[0249] It should be noted that the above is only a schematic example, and the embodiments of the present application are not limited thereto.

[0250] In summary, by means of the dimension transposition chip, the dimension transposition process is integrally executed, so as to targetedly adjust the dimension sequence of the tensor data based on multiple neural network layers supporting different dimensions in the neural network structure, so as to obtain tensor data of different dimensions that is convenient for more accurate input into the corresponding neural network layer for data processing; in addition, the dimension difference situation including at least one of the dimension quantity difference and the dimension sequence difference is fully considered, so that the dimension transposition parameters can be flexibly obtained based on the dimension difference, avoiding the limitation problem of presetting the dimension transposition parameters.

[0251] In the embodiments of the present application, the content of determining whether to perform data padding operation on the tensor data according to the dimension of the target address is introduced. Through the relationship between the target address and the device currently performing dimension transposition, it is possible to flexibly select whether to continue processing the tensor data of the second dimension sequence. When the preset conditions are met, it is not necessary to pad the specified dimension of the tensor data according to the preset data padding operation, thereby reducing the data processing amount to a certain extent and improving the data processing efficiency of the neural network structure while ensuring the dimension transposition effect.

[0252] In an optional embodiment, the dimension transposition chip can be deployed in computing devices such as terminals and servers, and the execution process of the dimension transposition method is realized by calling the dimension transposition chip. As Figure 8 shown, the dimension transposition method can be implemented as the following steps 810 to 840.

[0253] Step 810, obtain tensor data.

[0254] Among them, the tensor data is a feature representation generated during the operation of the neural network structure and expressed in the arrangement manner of the first dimension sequence.

[0255] Step 820, obtain the dimension transposition information corresponding to the tensor data.

[0256] Among them, the dimension transposition information is used to indicate transposing the first dimension sequence corresponding to the tensor data into the second dimension sequence.

[0257] Step 820 can refer to the relevant content of the above information acquisition module, which will not be elaborated here.

[0258] Step 830: Obtain a dimension transposition parameter based on the dimension difference between the first dimension sequence and the second dimension sequence.

[0259] Among them, the dimension difference includes at least one of the dimension quantity difference and the dimension sequence difference. The dimension quantity difference is used to represent the change in the number of dimensions, the dimension sequence difference is used to represent the change in the order of dimensions, and the dimension transposition parameter is used to transpose the first dimension sequence to the second dimension sequence.

[0260] In an optional embodiment, in response to the dimension transposition information being a sequence transformation value, determine the number of sequence dimensions represented by the sequence transformation value.

[0261] Among them, the sequence transformation value is used to transpose the tensor data of the first dimension sequence into the tensor data of the second dimension sequence.

[0262] In an optional embodiment, obtain a dimension transposition parameter based on the dimension quantity difference between the number of sequence dimensions and the number of first dimensions corresponding to the first dimension sequence.

[0263] In some embodiments, under the condition that the dimension quantity difference represents that the number of sequence dimensions is different from the number of first dimensions, perform sequence processing on the sequence transformation value based on the quantitative relationship between the number of sequence dimensions and the number of first dimensions in the sequence processing request to obtain a dimension transposition parameter.

[0264] Optionally, in response to the number of sequence dimensions being greater than the number of first dimensions, perform axis merging processing on the sequence transformation value to obtain a dimension transposition parameter.

[0265] Optionally, in response to the number of sequence dimensions being less than the number of first dimensions, perform axis splitting processing on the sequence transformation value to obtain a dimension transposition parameter. The axis splitting processing is used to split at least one dimension in the first dimension sequence.

[0266] In some embodiments, in response to the number of sequence dimensions being equal to the number of first dimensions, use the sequence transformation value as the dimension transposition parameter.

[0267] In an optional embodiment, in response to the dimension transposition information including the first data format corresponding to the first dimension sequence and the second data format corresponding to the second dimension sequence, obtain a dimension transposition parameter based on the first data format and the second data format.

[0268] In some embodiments, in response to the dimension transposition information including the first data format and the second data format, query a format transposition parameter table based on the first data format and the second data format.

[0269] Among them, the format transposition parameter table stores transposition parameters corresponding to multiple data format pairs. The data format pairs include a starting data format and an ending data format. The transposition parameter is used to transpose from the starting data format to the ending data format, and the data format is used to express the layout of the dimension sequence.

[0270] In some embodiments, dimension transposition parameters are obtained based on query results.

[0271] Optionally, query the starting data format in the format transposition parameter table in the first data format to determine multiple candidate parameters corresponding to the first data format.

[0272] Among them, the multiple candidate parameters are parameters corresponding to the first data format, and the multiple candidate parameters correspond one-to-one to the multiple ending data formats that form data format pairs with the first data format.

[0273] Optionally, query the multiple candidate parameters in the second data format to obtain dimension transposition parameters. The dimension transposition parameter is the transposition parameter of the data format pair formed by the first data format and the second data format.

[0274] Step 830 may refer to the content related to the above parameter acquisition module and will not be elaborated here.

[0275] Step 840, perform dimension adjustment on the tensor data of the first dimension sequence with the dimension transposition parameter to obtain tensor data with a second dimension sequence.

[0276] Among them, the tensor data with the second dimension sequence is used to input a network layer in the neural network for analysis based on the second dimension sequence.

[0277] In an optional embodiment, a target address is determined.

[0278] Among them, the target address is used to store the tensor data with the second dimension sequence.

[0279] In an optional embodiment, based on the location of the target address, perform data padding adjustment on the specified dimension in the tensor data of the second dimension sequence to obtain adjusted tensor data.

[0280] Among them, the adjusted tensor data is used to be sent to the target address for storage.

[0281] In some embodiments, in response to the target address being in the on-chip memory, perform data padding adjustment on the specified dimension in the tensor data of the second dimension sequence to obtain adjusted tensor data.

[0282] Among them, the on-chip memory is the memory deployed in the computing device for performing the dimension transposition process.

[0283] In some embodiments, in response to a preset condition being met between the first - dimension sequence and the second - dimension sequence, tensor data having the second - dimension sequence is stored at a target address.

[0284] Wherein, the preset condition is a condition that defines the dimension conversion between a first preset dimension in the first - dimension sequence and a second preset dimension in the second - dimension sequence.

[0285] Step 840 can refer to the content related to the above - mentioned dimension transpose module, which will not be elaborated here.

[0286] In summary, by means of a dimension transpose chip to integrally execute the dimension transpose process, so as to perform targeted adjustment on the dimension sequence of tensor data based on multiple neural network layers that support different dimensions in a neural network structure, thereby obtaining tensor data of different dimensions that is convenient for more accurate input into the corresponding neural network layer for data processing; in addition, fully considering the dimension difference situation including at least one of the dimension number difference and the dimension sequence difference, so that the dimension transpose parameters can be flexibly obtained based on the dimension difference, avoiding the limitation problem of presetting dimension transpose parameters.

[0287] In an alternative embodiment, the above - mentioned dimension transpose method can also be called "an accelerated design method for multi - layout conversion and tensor transpose". This method can be executed by a Data Transfer Engine (DTE), or integrated in a dimension transpose chip and deployed within the Data Transfer Engine for execution, etc.

[0288] Schematically, the parameters for performing dimension transpose within the Data Transfer Engine can be called "DTE Permute". DTE Permute generally supports 4 dimensions, where the third dimension - dimension 3 is data - continuous, representing a continuous storage state between data; if (0, 1, 2, 3) represents 4 dimensions, the source dimensions 0 / 1 / 2 / 3 can be mapped to any target dimension. The source dimension is the dimension before dimension transpose (which can be regarded as the above - mentioned first - dimension sequence), and the target dimension is the dimension after dimension transpose (which can be regarded as the above - mentioned second - dimension sequence). Source dimension 0 is the first dimension within the source dimension, source dimension 1 is the second dimension within the source dimension, etc.

[0289] Such as Figure 9 The example shows a schematic diagram of mapping the source dimension corresponding to the data format NCHW910 to the target dimension corresponding to the data format NHWC920, where N represents the batch size, C represents the number of channels, H represents the image height, and W represents the image width. It can be regarded as the source dimension (0, 1, 2, 3) being mapped to the target dimension (0, 2, 3, 1). Taking the dimension as dim for example, this process is achieved through the following four - dimensional dimension transpose parameter - permute(4D) expression:

[0290] DTE_TILE_PERMUTE_DIM3_MAP = 1, the target dim3 comes from the source dim1 - C;

[0291] DTE_TILE_PERMUTE_DIM2_MAP = 3, the target dim2 comes from the source dim3 - W;

[0292] DTE_TILE_PERMUTE_DIM1_MAP = 2, the target dim1 comes from the source dim2 - H;

[0293] DTE_TILE_PERMUTE_DIM0_MAP = 0, the target dim1 comes from the source dim0 - N.

[0294] Optionally, when the target address is in the on - chip memory, the target dimension 3 will be padded to bandwidth alignment. However, when DTE_ENHANCE_PERMUTE_MODE (enhanced mode) is enabled, the target dimension 3 will not be padded. The enhanced mode can only be enabled when the following preset conditions are met:

[0295] Preset condition 1:

[0296] (1) The source dim2 is mapped to the target dim3. For example (map[2] = 3, map[3] = 2);

[0297] (2) The source DTE_SRC_TILE_DIM2_SIZE == 3, 5, 7;

[0298] (3) DTE_ENHANCE_PERMUTE_MODE = 1.

[0299] Preset condition 2:

[0300] (1) The source dim2 is mapped to the target dim2;

[0301] (2) The source dim3 is mapped to the target dim3;

[0302] (3) DTE_ENHANCE_PERMUTE_MODE = 1.

[0303] The product of the target dim2 size and the dim3 size will be automatically padded to bandwidth alignment.

[0304] In linear algebra, when performing a transpose operation on a matrix, the rows of the matrix are actually changed to columns, and the columns are changed to rows. This means that the positions of the elements in the original matrix change in the transposed matrix. Suppose there is a matrix A with m rows and n columns, and its transposed matrix is denoted as A^T. Then the dimension of A^T is n rows and m columns, that is, the number of columns of the original matrix becomes the number of rows of the transposed matrix, and the number of rows of the original matrix becomes the number of columns of the transposed matrix. The principle of the transpose operation is: for each element A[i][j] in matrix A, place it at the position A^T[j][i] in the transposed matrix A^T. In other words, the element in the i-th row and j-th column of matrix A becomes the element in the j-th row and i-th column of the transposed matrix A^T. In a traditional Central Processing Unit (CPU), the transpose operation can be achieved by swapping the row index and column index of the matrix.

[0305] During the dimension transpose process, the accelerator DTE Permute instruction is used to implement the basic 4D transpose. During the process of data being transferred from the source storage location - High Bandwidth Memory (HBM) to the target storage space - Vector Process Engine Buffer (VB), the first in-line transpose is performed. Here, the in-line transpose is used to represent the dimension transpose process while transferring data.

[0306] Schematically, as Figure 10 shown, is a schematic diagram for performing a single in-line transpose. Among them, the srcdim (axis 1) and dest dim (axis 2) data of the matrix data block are swapped with each other, A1, A4, A7 are changed from columns to rows, and A1, A2, A3 are changed from rows to column arrangements; ping-pong parallelism is carried out in VB1010. Ping-pong parallelism is a parallel computing mode, usually used to describe the interactive communication and computing between two or more processes or threads, used to describe the process or threads taking turns to send messages and execute calculations, and ping and pong are regarded as components within VB1010 respectively.

[0307] Among them, the data in ping is transferred into VB1010, the data in pong is calculated, and the pipelined parallel ping-pong is switched to improve the transfer and calculation efficiency of the transpose transformation. If only one transpose transformation needs to be performed, after the data is calculated, it is moved out of VB1010 and back to HBM1020, and the calculation result of successful dimension transpose is obtained.

[0308] Schematically, as Figure 11 shown, is a schematic diagram for performing two in-line transposes.

[0309] For the case where 4D tensor data needs to be transposed twice, for example: the dimension sequence is (0, 1, 2, 3), and it is required to transpose axis 1 (dim0) and axis 2 (dim1), and axis 2 (dim2) and axis 4 (dim3). The first in-line transpose 1110 transposes axis 1 and axis 2, that is, the rows and columns of each data block of A / B / C / D are transposed with each other from HBM and stored in the ping / pong buffer in VB; the second in-line transpose 1120 needs to transpose axis 3 and axis 4. It can be carried out when moving out for the second in-line transpose 1120. During the process of moving out from the source storage location ping / pong VB to HBM, the data with row AB sorted is transposed into column AB, and the data of column AC is transposed into row AC, obtaining the final data after two transposes, so that the rows and columns of data blocks A, B, C, and D are transposed, and the row and column data in each data block are also transposed.

[0310] In an optional embodiment, since the permute instruction of AI hardware usually supports the permute of 4D tensors, in actual neural network application scenarios, there are often transpose requirements for 5D, 6D or even higher-dimensional tensors. At this time, the permute of 5D or 6D tensors needs to be disassembled into two or more permutes of 4D tensors, and at the same time, its permute parameters are automatically adjusted. For example: for a 5D transpose, its permute parameter is (0, 1, 3, 2, 4), which is split into axes (0, 2, 1, 3) and (0, 1, 2, 3); where the 0th axis and 1st axis of the original permute are first merged into one axis (axis merging process), and then its permute parameter is adjusted to 4D (0, 2, 1, 3) for a 4D permute. After that, the transpose effect has actually been achieved, and no additional operation is required for the second transpose, so the permute parameter is kept as (0, 1, 2, 3), thus completing the 5D transpose.

[0311] In some embodiments, the AI chip data formats include: feature map formats: NHWC, NCHW, NC1HWC0 (5D data format, splitting the c dimension in NHWC into c1 and c0, where c0 is often 32 or 64), weight data formats: KhKwCiCo, KwCiKhCo, Co1Ci1KhKwCi0Co0 (6D weight format, splitting the ci and co dimensions based on KhKwCiCo), Co1KwCiKhCo0 (5D weight format, kw ci kh co, co split, for scenarios with small input channels), and Ci1KhKwCi0NCo (5D weight format, ci split, co inflated, for scenarios with small output channels), matmul and batchmatmul data formats: B0B1HW, B0B1W1HW0; in specific network scenarios, usually only the data layout format of the source src tensor and the known data layout format of the target dst tensor are declared, and there are no specific transpose nodes in the network and no permute parameters are provided for such data format conversions. Therefore, a general multi-layout conversion table (i.e., the above format transpose parameter table) can be established. The input is the data layout format of the source src tensor and the data layout format of the target dst tensor, and the output is the two permute parameters for realizing the conversion from src to dst, so as to realize the automatic layout conversion operations among feature map, weight, batch matmul, and matmul respectively, so that the network operation and calculation meet the requirements of the layout, and there are no data format errors resulting in incorrect operation of the entire network.

[0312] The architecture diagram of layout conversion and multi-dimensional (N Dimension, ND) transpose is as Figure 12 shown. In the whole design, what needs to be input is the source dimension format 1210 (src tensor format) and the target dimension format 1220 (dst tensor format), or the input sequence transformation value (transpose axes axis information 1230); among them, the source dimension format 1210 and the target dimension format 1220 can be regarded as a form of dimension transpose information. The source dimension format 1210 is the first data format corresponding to the first dimension sequence, the target dimension format is the second data format corresponding to the second dimension sequence, and the sequence transformation value can be regarded as another form of dimension transpose information.

[0313] The entire design transposition will query the layout conversion table as shown in Figure 6 to obtain the permute axes information for two transpositions based on the input src format and dst format information. For example, when converting from NCHW to NHWC, the permute axes information obtained for the two transpositions is: 1st: (0, 1, 2, 3); 2nd: (0, 2, 3, 1). According to the obtained permute axes information for the two transpositions, a transposition operation is performed to transpose the src tensor in HBM and store the resulting dst tensor back in HBM.

[0314] In addition, it is also possible to input arbitrary ND transpose axes information. For example, (0, 1, 3, 2, 4) is split and combined to obtain the permute axes information for two transpositions (i.e., the dimension transposition parameters). Similarly, based on the obtained Permute_1_axes and Permute_2_axes as input information to transpose_6d (here 6D is used as an example, and it can also be input to 4D, 5D, etc.), two transpose_4d operations are performed to efficiently complete the ND transpose operation.

[0315] In some embodiments, when performing dimension transposition on tensor data with channel numbers of 3, 5, 7 (channel 3, 5, 7), the dedicated hardware instruction enhancement mode (enhance mode) can be used. During in-line transposition, there is no need for additional padding to align to 128 bytes. That is, when the data type is fp32, the channel is padded to 32, and when the data type is fp16, the channel is padded to 64; when the channel is not 3, 5, 7 but is also a small channel, the corresponding nearby padding is padded to 3, 5, 7 instead of 32 or 64, greatly reducing the amount of non-effective data that needs to be transferred and improving the permute performance.

[0316] When DTE_ENHANCE_PERMUTE_MODE is enabled, the target dimension 3 will not be padded. The enhancement mode can only be enabled when the following preset conditions are met:

[0317] Preset condition 1:

[0318] The source dim2 is mapped to the target dim3;

[0319] The source DTE_SRC_TILE_DIM2_SIZE == 3, 5, 7;

[0320] DTE_ENHANCE_PERMUTE_MODE = 1.

[0321] For b32, the target dim2 will be automatically filled with 64, and for b16, the target dim2 will be automatically filled with 128, where b is the abbreviation of batch.

[0322] Preset condition 2:

[0323] The source dim2 should be mapped to the target dim2;

[0324] The source dim3 should be mapped to the target dim3;

[0325] DTE_ENHANCE_PERMUTE_MODE = 1.

[0326] And the multipliers of the target dim2 size and dim3 size will be automatically filled to bandwidth alignment.

[0327] In an optional embodiment, the method can be applied to the application scenario of a neural network structure. The transpose operation is one of the very common operations, used to change the shape and arrangement of the Tensor to adapt to the data format requirements between different layers.

[0328] Schematically, in models such as convolutional neural networks, recurrent neural networks, and Transformers, since the required dimension orders and arrangement ways of the input data and the weights of each layer may be different, operations such as transpose are needed to adjust the input data.

[0329] For example: For a convolutional neural network, the input data is usually in the NCHW format, while the weights of the convolutional layer are in the OC3 format, that is, including out_channels, in_channels, kernel_height, and kernel_width parameters; for a recurrent neural network, the input data is usually in the format of (batch_size, sequence_length, input_size), while the output data of the recurrent layer is in the format of (batch_size, hidden_size).

[0330] In addition to the above scenarios, there are some other application scenarios that may involve the transpose operation. For example, in an image segmentation task, we can use transpose to transpose the feature map so as to align it with the original image. In natural language processing, we can use transpose to transpose the word vector matrix so as to calculate it with the context matrix. In the attention mechanism, we can use transpose to transpose the attention matrix so as to perform multiplication calculation with the query matrix. The transpose operation is widely used in deep learning neural networks, and the need for data layout transformation in AI processors is also quite common. An efficient transpose operation is very important.

[0331] As Figure 13 shown, under the dot product (Dot) operation 1310, the feature representation can be changed based on processes such as transpose and dimension transformation; under the element-wise (Eltwise) operation 1320, the feature representation can be changed based on processes such as transpose and activation transformation; under the convolution (Conv) operation 1330, the feature representation can be changed based on processes such as dimension transformation and transpose, etc. Among them, the dimension sequence corresponding to the tensor data can be changed based on the transpose operation.

[0332] It should be noted that the above are only illustrative examples, and the embodiments of the present application are not limited thereto.

[0333] In summary, by means of a dimension transpose chip integrated to execute the dimension transpose process, the dimension sequence of the tensor data can be specifically adjusted based on multiple neural network layers supporting different dimensions in the neural network structure, so as to obtain tensor data of different dimensions that is more convenient for accurate input into the corresponding neural network layer for data processing; in addition, the dimension difference situation including at least one of the dimension quantity difference and the dimension sequence difference is fully considered, so that the dimension transpose parameters can be flexibly obtained based on the dimension difference, avoiding the limitation problem of presetting the dimension transpose parameters.

[0334] In the embodiments of the present application, by means of at least one of the dimension transpose information of the data format and the sequence transformation value, the dimension transpose parameters can be obtained more efficiently and selectively by querying at least one of the format transpose parameter table and the split / merge axis processing, so that the efficiency and generality of the layout conversion can be effectively improved, and the transpose performance when transposing tensor data of different data formats is optimized. It meets the requirements of various deep learning networks and effectively improves the competitiveness of data processing.

[0335] Figure 14The structural schematic diagram of a server provided by an exemplary embodiment of the present application is shown. The server 1400 includes a Central Processing Unit (CPU) 1401, a system memory 1404 including a Random Access Memory (RAM) 1402 and a Read Only Memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the central processing unit 1401. The server 1400 further includes a mass storage device 1406 for storing an operating system 1413, application programs 1414, and other program modules 1415.

[0336] The mass storage device 1406 is connected to the central processing unit 1401 through a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1406 and its associated computer-readable medium provide non-volatile storage for the server 1400. That is to say, the mass storage device 1406 may include computer-readable media (not shown) such as a hard disk or a Compact Disc Read Only Memory (CD-ROM) drive.

[0337] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The above-mentioned system memory 1404 and mass storage device 1406 may be collectively referred to as memory.

[0338] According to various embodiments of the present application, the server 1400 may also be run by connecting to a remote computer on the network through a network such as the Internet. That is, the server 1400 may be connected to the network 1412 through a network interface unit 1411 connected to the system bus 1405, or in other words, the network interface unit 1411 may also be used to connect to other types of networks or remote computer systems (not shown).

[0339] The above-mentioned memory further includes one or more programs, and one or more programs are stored in the memory and are configured to be executed by the CPU.

[0340] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and at least one instruction, at least one program, a code set, or an instruction set is loaded and executed by the processor to implement the dimension transposition method provided by the above-mentioned method embodiments.

[0341] An embodiment of the present application further provides a computer-readable storage medium, on which at least one instruction, at least one program segment, a code set or an instruction set is stored, and the at least one instruction, at least one program segment, the code set or the instruction set is loaded and executed by a processor to implement the dimension transposition method provided by each of the above method embodiments.

[0342] An embodiment of the present application further provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the dimension transposition method described in any one of the above embodiments.

[0343] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A dimension transposition chip, characterized in that, The chip includes: An information acquisition module, configured to acquire dimension transposition information corresponding to tensor data, where the tensor data is a feature representation expressed in a first dimension sequence generated during the operation of a neural network structure, and the dimension transposition information is used to indicate transposing the first dimension sequence into a second dimension sequence; and send the dimension transposition information to a parameter acquisition module; The parameter acquisition module, configured to receive the dimension transposition information; based on the dimension difference between the first dimension sequence and the second dimension sequence, acquire a dimension transposition parameter, where the dimension difference includes at least one of a dimension quantity difference and a dimension sequence difference, the dimension quantity difference is used to represent the change in the number of dimensions, the dimension sequence difference is used to represent the change in the order of dimensions, and the dimension transposition parameter is used to transpose the first dimension sequence to the second dimension sequence; and send the dimension transposition parameter to a dimension transposition module; The dimension transposition module, configured to receive the dimension transposition parameter; acquire the tensor data; transpose the tensor data of the first dimension sequence with the dimension transposition parameter to obtain tensor data with a second dimension sequence, and the tensor data with the second dimension sequence is used to be input into a neural network layer in the neural network structure for analysis based on the second dimension sequence.

2. The chip according to claim 1, wherein The parameter acquisition module is further configured to analyze the dimension transposition information; in response to the dimension transposition information being a sequence transformation value, determine the sequence dimension quantity characterized by the sequence transformation value, where the sequence transformation value is used to transpose the tensor data of the first dimension sequence into the tensor data of the second dimension sequence; and based on the dimension quantity difference between the sequence dimension quantity and the first dimension quantity corresponding to the first dimension sequence, acquire the dimension transposition parameter.

3. The chip according to claim 2, characterized in that, The parameter acquisition module includes an information analysis unit and a sequence processing unit; The information analysis unit is further configured to generate a sequence processing request for performing sequence adjustment on the sequence transformation value under the condition that the dimension quantity difference represents that the sequence dimension quantity is different from the first dimension quantity, and the sequence processing request is used to perform sequence adjustment on the sequence transformation value; Send the sequence processing request to the sequence processing unit; The sequence processing unit is configured to receive the sequence processing request; Based on the quantitative relationship between the sequence dimension quantity and the first dimension quantity in the sequence processing request, perform sequence processing on the sequence transformation value to obtain the dimension transposition parameter.

4. The chip according to claim 3, wherein The sequence processing unit is further configured to, in response to the sequence processing request indicating that the sequence dimension quantity is greater than the first dimension quantity, perform an axis merging process on the sequence transformation value to obtain the dimension transposition parameter, and the axis merging process is used to merge at least two dimensions in the sequence transformation value; The sequence processing unit is further configured to, in response to the sequence processing request indicating that the number of sequence dimensions is less than the number of the first dimensions, perform axis splitting processing on the sequence transformation value to obtain the dimension transposition parameter, where the axis splitting processing is used to split at least one dimension in the sequence transformation value.

5. The chip according to claim 3, wherein the information analysis unit is further configured to, in response to the number of sequence dimensions being equal to the number of the first dimensions, use the sequence transformation value as the dimension transposition parameter.

6. The chip according to claim 1, wherein the parameter acquisition module is further configured to analyze the dimension transposition information; in response to the dimension transposition information including a first data format corresponding to the first dimension sequence and a second data format corresponding to the second dimension sequence, obtain the dimension transposition parameter based on the first data format and the second data format.

7. The chip according to claim 6, characterized in that, The parameter acquisition module includes an information analysis unit and a parameter call unit; the information analysis unit is configured to analyze the dimension transposition information; in response to the dimension transposition information including the first data format and the second data format, send a parameter acquisition request to the parameter call unit; the parameter call unit is configured to store a format transposition parameter table, in which a plurality of transposition parameters respectively corresponding to data format pairs are stored, the data format pair includes a starting data format and an ending data format, the transposition parameter is used to transpose from the starting data format to the ending data format, and the data format is used to represent the layout of the dimension sequence; receive the parameter acquisition request; query the format transposition parameter table based on the first data format and the second data format to obtain the dimension transposition parameter.

8. The chip according to claim 7, wherein the parameter call unit is further configured to query the starting data format in the format transposition parameter table in the first data format, determine a plurality of candidate parameters corresponding to the first data format, the plurality of candidate parameters are parameters corresponding to the first data format, and the plurality of candidate parameters correspond one-to-one to a plurality of ending data formats that form the data format pair with the first data format; query the plurality of candidate parameters in the second data format to obtain the dimension transposition parameter, and the dimension transposition parameter is the transposition parameter of the data format pair formed by the first data format and the second data format.

9. The chip according to claim 6, wherein the parameter call unit is further configured to, in response to not querying the data format pair formed by the first data format and the second data format in the format transposition parameter table, determine the number of the first dimensions of the first dimension sequence corresponding to the first data format and determine the number of the second dimensions of the second dimension sequence corresponding to the second data format; in response to the number of the first dimensions being different from the number of the second dimensions, perform sequence processing on the second dimension sequence based on the first dimension sequence to obtain the dimension transposition parameter.

10. The chip according to claim 9, wherein the parameter calling unit is further configured to, in response to the number of the second dimensions being greater than the number of the first dimensions, perform an axis merging process on the second dimension sequence to obtain the dimension transposition parameter, where the axis merging process is used to merge at least two dimensions in the second dimension sequence; or, in response to the number of the second dimensions being less than the number of the first dimensions, perform an axis splitting process on the second dimension sequence to obtain the dimension transposition parameter, where the axis splitting process is used to split at least one dimension in the second dimension sequence.

11. The chip according to any one of claims 1 to 10, wherein the dimension transposition module is further configured to determine a target address for storing the tensor data having the second dimension sequence; based on the position of the target address, perform data padding adjustment on a specified dimension in the tensor data of the second dimension sequence to obtain adjusted tensor data; and send the adjusted tensor data to the target address for storage.

12. The chip according to claim 11, wherein the dimension transposition module is further configured to, in response to the target address being located in the on-chip memory, perform data padding adjustment on a specified dimension in the tensor data of the second dimension sequence to obtain the adjusted tensor data; wherein the on-chip memory is a memory deployed in a computing device for performing the dimension transposition process.

13. The chip according to claim 11, wherein the dimension transposition module is further configured to, in response to a preset condition being met between the first dimension sequence and the second dimension sequence, store the tensor data having the second dimension sequence in the target address, where the preset condition is a condition that defines the dimension conversion between a first preset dimension in the first dimension sequence and a second preset dimension in the second dimension sequence.

14. A dimension transposition method, characterized in that, The method includes: obtaining tensor data, which is a feature representation expressed in an arrangement manner of a first dimension sequence generated during the operation of a neural network structure; obtaining dimension transposition information corresponding to the tensor data, where the dimension transposition information is used to indicate transposing the first dimension sequence corresponding to the tensor data into a second dimension sequence; obtaining a dimension transposition parameter based on the dimension difference between the first dimension sequence and the second dimension sequence, where the dimension difference includes at least one of a dimension number difference and a dimension sequence difference, the dimension number difference is used to express the change in the number of dimensions, the dimension sequence difference is used to express the change in the order of dimensions, and the dimension transposition parameter is used to transpose the first dimension sequence to the second dimension sequence; performing dimension adjustment on the tensor data of the first dimension sequence with the dimension transposition parameter to obtain the tensor data having the second dimension sequence, where the tensor data having the second dimension sequence is used to input a network layer in the neural network for analysis based on the second dimension sequence.

15. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the dimension transposition method as described in claim 14.

16. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the dimension transposition method as described in claim 14.

17. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the dimension transposition method as described in claim 14 is implemented.