Method, device, system, and storage medium for generating an operator for tensor operations

By generating tensor operations, the input information is parsed and converted into abstract syntax tree information and intermediate expression information, which directly reflects the hardware characteristics, solving the problem that traditional methods cannot fully utilize processor performance, and achieving efficient generation of high-performance tensor operations operator library.

CN114564686BActive Publication Date: 2025-07-11SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210237709.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-07-11
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

Traditional tensor operation operator generation methods cannot fully utilize the hardware performance of each processor, especially the hardware characteristics of the GPU, and it is difficult to generate a high-performance tensor operation operator library.

Method used

By analyzing the input information, generating matrix sequences, symbol representation information is generated based on the matrix sequence and symbol vectors, and converting them into abstract syntax tree information and intermediate expression information, and finally generating assembly code, which directly reflects the hardware characteristics to optimize operator generation.

Benefits of technology

The development efficiency of tensor operation acceleration library is improved, and the hardware performance of each processor, especially the performance of GPU, is fully utilized to generate high-performance tensor operation operator library.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114564686B_ABST
    Figure CN114564686B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, a computing device, a computing system, and a storage medium for generating an operator for tensor operations. The method includes: parsing input information for generating a matrix sequence; generating symbolic representation information for operations based on the generated matrix sequence and an input symbolic vector; converting the generated symbolic representation information into abstract syntax tree information or intermediate representation information; and generating assembly code for an operator for tensor operations based on basic description information included in the generated abstract syntax tree information and intermediate representation information. The present disclosure can make full use of the hardware performance of each processor, and can efficiently generate an operator library for high-performance tensor operations for a GPU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of information processing, and more particularly to a method, a computing device, and a storage medium for generating an operator for tensor operations. Background Art

[0002] A tensor is a high-dimensional array. Tensor operations are the basic operations in the field of digital signal processing, especially in deep learning. Taking the Fourier Transform as an example, it is a very important mathematical transformation method in the field of digital signal processing, used to implement the transformation process of a signal from the time domain to the frequency domain. The Discrete Fourier Transform (DFT) is the representation form of the continuous Fourier Transform in a discrete system. The Fast Fourier Transform (FFT) is a general term for efficient and fast calculation methods for calculating the DFT using a computer, which can greatly reduce the number of multiplications required for a computer to calculate the discrete Fourier transform.

[0003] Traditional methods for generating an operator for tensor operations, such as but not limited to CPU-based FFT operator generation techniques (e.g., which can be implemented by calling the FFTW library), generally first generate code in a high-level language, such as C or C++; and then generate the final target file for a specific hardware through compilation by a compiler. In order to achieve generality, the compilation process of the compiler will ignore the specific hardware characteristics. Therefore, although the above methods have high flexibility, they cannot make full use of the hardware performance of each processor. If you want to further improve the performance of the program, special optimization means are still required.

[0004] In addition, currently, technologies (such as technologies for automatically generating a high-performance FFT operator library) and products for efficiently generating an operator for tensor operations on a GPU are relatively rare. The reasons are as follows. On the one hand, due to different hardware characteristics and software abstractions, the method for generating an operator for tensor operations on a CPU (e.g., an FFT operator) cannot be directly applied to a GPU. On the other hand, even if it is possible to generate a high-level language implementation for a GPU using a code generation technology similar to that on a CPU and then compile it into GPU code, the performance of the code is difficult to guarantee, and there is still a certain gap from the performance of a tensor operation algorithm (e.g., an FFT algorithm) designed directly using assembly. In addition, in the traditional tensor operation operator (e.g., an FFT operator) generation scheme, small code snippet templates still need to be manually written for specialization and assembly, which is not conducive to improving the development efficiency of a tensor operation (e.g., an FFT) acceleration library.

[0005] In summary, the disadvantages of the traditional solution for generating operators for tensor operations are as follows: it cannot make full use of the hardware performance of each processor, and it is difficult to effectively generate a high-performance operator library for tensor operations for GPUs. Summary of the Invention

[0006] The present disclosure provides a method, a computing device, and a computer-readable storage medium for generating operators for tensor operations, which can make full use of the hardware performance of each processor and can efficiently generate a high-performance operator library for tensor operations for GPUs.

[0007] According to a first aspect of the present disclosure, a method for generating an operator for tensor operations is provided. The method includes: parsing input information to generate a matrix sequence; generating symbolic representation information for operations based on the generated matrix sequence and an input symbol vector; converting the generated symbolic representation information into abstract syntax tree information and intermediate representation information; and generating assembly code for an operator for tensor operations based on basic description information included in the generated abstract syntax tree information and intermediate representation information.

[0008] According to a second aspect of the present invention, a computing device is further provided. The computing device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the computing device can perform the method according to the first aspect of the present disclosure.

[0009] According to a third aspect of the present disclosure, a computer-readable storage medium is further provided. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a machine, it performs the method according to the first aspect of the present disclosure.

[0010] In some embodiments, the tensor operation is a fast Fourier transform, and the basic description information at least includes: the size of the matrix, the operation associated with the matrix, and description parameters of the hardware resources required for the operation associated with the matrix.

[0011] In some embodiments, the method for generating an operator for tensor operations further includes: storing information about the matrix sequence in a tree-shaped data structure, and the information about the matrix sequence at least includes: a plurality of sub-matrices included in the matrix sequence, operators associated with the sub-matrices, and operands.

[0012] In some embodiments, converting the generated symbol representation information into abstract syntax tree information and intermediate representation information includes: determining whether the symbol representation information is an algebraic expression based on the identification information of the symbol representation information; in response to determining that the symbol representation information is an algebraic expression, converting the symbol representation information into abstract syntax tree information via an abstract syntax tree parsing algorithm; and in response to determining that the symbol representation information is not an algebraic expression, generating intermediate representation information based on the generated symbol representation information.

[0013] In some embodiments, generating a matrix sequence includes: obtaining the size and transformation type of the input data to be transformed; performing prime factorization on the size of the input data so as to decompose it into multiple prime factors; and performing matrix factorization based on the multiple prime factors and the transformation type; so as to generate a matrix sequence.

[0014] In some embodiments, generating symbol representation information for operations based on the generated matrix sequence and the input symbol vector includes: multiplying the current sub-matrix among the multiple sub-matrices included in the matrix sequence by the input symbol vector corresponding to the current sub-matrix so as to generate symbol representation information for the current-level operation; and multiplying the next sub-matrix among the multiple sub-matrices included in the matrix sequence by the input symbol vector corresponding to the next sub-matrix so as to generate symbol representation information for the next-level operation, where the input symbol vector corresponding to the next sub-matrix is generated by the product of the current sub-matrix and the input symbol vector corresponding to the current sub-matrix. In some embodiments, the operators and operands associated with the sub-matrices are generated based on the parsing of a specific type of sub-matrix, and the specific type of sub-matrix is determined based on the size of the sub-matrix, the computational behavior related to the sub-matrix, and the constraint data of the hardware resources.

[0015] In some embodiments, generating assembly code for an operator of tensor operations includes: performing optimization of machine-independent code; performing register allocation; converting the operation operations indicated in the intermediate representation information into corresponding assembly instructions; parsing the abstract syntax tree information to generate intermediate representation information; and performing scheduling optimization on the generated intermediate representation information for generating assembly code for an operator of fast Fourier transform.

[0016] In some embodiments, generating assembly code for an operator of tensor operations includes: performing optimization of machine-independent code; performing register allocation; converting the operation operations indicated in the intermediate representation information into corresponding assembly instructions; converting the abstract syntax tree information into assembly code for scheduling; and generating assembly code for an operator of fast Fourier transform.

[0017] In some embodiments, generating assembly code for an operator of a tensor operation includes: generating assembly code for an operator of a fast Fourier transform based on basic description information and the type of the basic description information, where the type of the basic description information includes: permutation, discrete Fourier transform, and twiddle factor.

[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements.

[0020] Figure 1 A schematic diagram of a computing device for a method of generating an operator of a tensor operation according to an embodiment of the present disclosure is shown.

[0021] Figure 2 A flowchart of a method of generating an operator of a tensor operation according to an embodiment of the present disclosure is shown.

[0022] Figure 3 A schematic diagram of a tree-shaped data structure according to an embodiment of the present disclosure is shown.

[0023] Figure 4 A schematic diagram of a method of generating symbolic representation information for an operation according to an embodiment of the present disclosure is shown.

[0024] Figure 5 A flowchart of a method of generating assembly code for an operator of a tensor operation according to an embodiment of the present disclosure is shown.

[0025] Figure 6 A flowchart of a method of converting the generated symbolic representation information into abstract syntax tree information according to an embodiment of the present disclosure is shown.

[0026] Figure 7 A flowchart of a method of generating a matrix sequence according to an embodiment of the present disclosure is shown.

[0027] Figure 8 A block diagram of an electronic device suitable for implementing the embodiments of the present disclosure is schematically shown.

[0028] In each of the drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure will be more thorough and complete, and can fully convey the scope of the present disclosure to those skilled in the art.

[0030] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.

[0031] As described above, the disadvantage of the traditional scheme for generating operators for tensor operations is that it cannot fully utilize the hardware performance of each processor, and it is difficult to effectively generate an operator library for high-performance tensor operations for GPUs.

[0032] To at least partially solve one or more of the above problems and other potential problems, example embodiments of the present disclosure propose a method, a computing device, a computing system, and a computer-readable storage medium for generating operators for tensor operations. In the solution of the present disclosure: a matrix sequence is generated by parsing input information, and symbolic representation information for a series of operations is generated based on the matrix sequence and an input symbol vector; then, symbolic representation information for the operations is generated based on the generated matrix sequence and input symbol vector; the generated symbolic representation information is then converted into abstract syntax tree information and intermediate representation information; and assembly code for an operator for tensor operations is generated based on the basic description information included in the generated abstract syntax tree information and intermediate representation information; the present disclosure can adopt symbolic operation and AST parsing technologies, and can dynamically generate various intermediate small code segments, avoiding the problem in the traditional FFT operator generation scheme that still requires manual writing of small code segment templates for specialization and assembly, and improving the development efficiency of the acceleration library for tensor operations. In addition, since the assembly code is directly generated from the hardware characteristics reflected by the AST and the basic description information, the problem that it is difficult to fully utilize the GPU characteristics due to the limited hardware abstraction of high-level languages is avoided, so that the generated code can fully utilize the GPU for high-performance tensor operations. Therefore, the present disclosure can fully utilize the hardware performance of each processor, and can efficiently generate an operator library for high-performance tensor operations for GPUs.

[0033] Figure 1 FIG. shows a schematic diagram of a computing device 100 for a method of generating an operator for tensor operations according to an embodiment of the present disclosure. AsFigure 1 As shown, the computing device 100 includes: a matrix sequence construction module 110, a symbol representation information and abstract syntax tree information generation module 120, and an assembly code generation module 130. In some embodiments, the computing device 100 may have one or more processing units, including dedicated processing units such as a graphics processing unit (GPU), a field programmable gate array (FPGA), and an application specific integrated circuit (ASIC), as well as a general purpose processing unit such as a central processing unit (CPU).

[0034] Regarding the matrix sequence construction module 110, it is used to parse the input information for generating a matrix sequence. Specifically, the matrix sequence construction module 110 obtains the size and transformation type of the input data to be transformed; performs prime factorization on the size of the input data so as to decompose it into multiple prime factors (more fine-grained data); and performs matrix factorization based on the multiple prime factors and the transformation type; so as to generate a matrix sequence. The way of matrix factorization depends on what form of data transformation is performed.

[0035] Regarding the symbol representation information and abstract syntax tree information generation module 120, it is used to convert the matrix sequence generated by the matrix sequence construction module 110 into symbol representation information, and convert the symbol representation information into abstract syntax tree information and intermediate representation information. The symbol representation information and abstract syntax tree information generation module 120 further includes, for example, a symbol representation information generation module 122 and an abstract syntax tree information and intermediate representation information generation module 124. The symbol representation information generation module 122 is used to generate symbol representation information for operations based on the generated matrix sequence and the input symbol vector. The abstract syntax tree information and intermediate representation information generation module 124 is used to convert the generated symbol representation information into abstract syntax tree information or intermediate representation information. For example, multiplying the current sub-matrix among the multiple sub-matrices included in the matrix sequence by the input symbol vector corresponding to the current sub-matrix to generate symbol representation information for the current-level operation; and multiplying the next sub-matrix among the multiple sub-matrices included in the matrix sequence by the input symbol vector corresponding to the next sub-matrix (where the input symbol vector corresponding to the next sub-matrix is generated by the product of the current sub-matrix and the input symbol vector corresponding to the current sub-matrix) to generate symbol representation information for the next-level operation, and generating a symbol expression into abstract syntax tree information (Abstract Syntax Tree, AST) and intermediate representation information (Intermediate representation, IR). Among them, the intermediate representation (Intermediate representation, IR) includes basic description information. The basic description information can reflect the relevant operations and hardware characteristics of tensor operations. In some embodiments, the basic description information at least includes: the size of the matrix, the operation operations associated with the matrix, and the description parameters of the hardware resources required for the operations associated with the matrix.

[0036] Regarding the assembly code generation module 130, it is used to generate assembly code for the operators of tensor operations based on the basic description information included in the generated abstract syntax tree information and intermediate representation information.

[0037] The following will be combined with Figure 2 to describe the method 200 for generating operators of tensor operations according to an embodiment of the present disclosure. Figure 2 FIG. shows a method for generating operators of tensor operations according to an embodiment of the present disclosure. It should be understood that the method 200 can be executed, for example, at Figure 8 the computing device 800 described. It can also be executed at Figure 1 the computing device 100 described. It should be understood that the method 200 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.

[0038] At step 202, computing device 100 parses the input information for generating a matrix sequence.

[0039] Regarding the method for generating a matrix sequence, for example, it includes: determining a fast Fourier transform algorithm based on the size and transform type of the input data to be transformed; and generating a matrix sequence based on the determined fast Fourier transform algorithm. As another example, computing device 100 obtains the size and transform type of the input data to be transformed; performs prime factorization on the size of the input data to decompose it into multiple prime factors; and performs matrix factorization based on the multiple prime factors and the transform type; so as to generate a matrix sequence. The following will be combined with Figure 7 to illustrate method 700 for generating a matrix sequence, which will not be elaborated here.

[0040] In some embodiments, method 200 further includes that computing device 100 stores information about the matrix sequence in a tree-shaped data structure. The information about the matrix sequence at least includes: multiple sub-matrices included in the matrix sequence, operators and operands associated with the sub-matrices. An operand, that is, operand. An operator, also called an operation operator, that is, operator. Since the tree-shaped data structure stores not only sub-matrices that are easy to understand mathematically, but also the operators and operands associated with the sub-matrices, which are atomic description information for matrix operations or operations, hereby, the present disclosure makes the operations more friendly to hardware.

[0041] In some embodiments, the operators and operands associated with the sub-matrices are generated based on the parsing of a specific type of sub-matrix, and the specific type of sub-matrix is determined based on the size of the sub-matrix, the calculation behavior related to the sub-matrix, and the limited data of the hardware resources. For example, if computing device 10 determines that the range of the operation result associated with the current sub-matrix is less than or equal to a predetermined limit threshold of the hardware resources (such as but not limited to 32) based on the size of the current sub-matrix and the operation step size, use the basic description information of the hardware resources to parse the current sub-matrix to generate the operators and operands associated with the current sub-matrix. The basic description information also includes, for example, the size of the hardware, the number of data that can be processed simultaneously, and so on.

[0042] The following will be combined with Figure 3 to illustrate the storage method of the tree-shaped data structure. Figure 3 FIG. shows a schematic diagram of a tree-shaped data structure 300 according to an embodiment of the present disclosure. Figure 3 The shown tree-shaped data structure 300 stores, for example, information about a matrix sequence, and the information of the matrix sequence indicates multiple sub-matrices, operators and operands generated by performing matrix factorization operations on a predetermined algorithm (I@P@I)*(I@DFT)*T.

[0043] For example, the label 332 indicates the operand of the identity matrix I (i.e., Operand: Identity matrix, I). The label 334 indicates the operand of the primitive matrix P (i.e., Operand: Primitive matrix, P).

[0044] The label 322 indicates the operator @ (i.e., Operator: @), which indicates performing the "@" operation on the data of the associated leaf nodes under the node corresponding to the label 322, that is, performing the "@" operation on the operand of the identity matrix I indicated by the label 332 (i.e., Operand: Identity matrix, I) and the operand of the primitive matrix P indicated by the label 334 (i.e., Operand: Primitive matrix, P). The data of the node corresponding to the label 322 is, for example, "I@P".

[0045] The label 324 indicates the operand of the identity matrix I (i.e., Operand: Identity matrix, I). The operator @ indicated by the label 312 (i.e., Operator: @) indicates performing the "@" operation on the data of the associated leaf nodes under the node corresponding to the label 312, that is, performing the "@" operation on the data of the node corresponding to the label 322 ("I@P") and the operand of the identity matrix I indicated by the label 324 (i.e., Operand: Identity matrix, I). The data of the node corresponding to the label 312 is, for example, "I@P@I".

[0046] The operator @ indicated by the label 326 (i.e., Operator: @) indicates performing the "@" operation on the data of the associated leaf nodes under the node corresponding to the label 326, that is, performing the "@" operation on the operand of the identity matrix I indicated by the label 336 (i.e., Operand: Identity matrix, I) and the data of the node corresponding to the operand of the DFT matrix indicated by the label 338 (i.e., Operand: DFTmatrix, DFT). The data of the node corresponding to the label 326 is, for example, "I@DFT".

[0047] The tag 328 indicates the operand of the rotation matrix T (i.e., Operand: Twiddle matrix, T). The operator * indicated by the tag 316 (i.e., Operator: *) indicates that the "*" operation is performed on the data of the associated leaf nodes below the node corresponding to the tag 316, that is, the "*" operation is performed on the data of the node corresponding to the tag 326 ("I@DFT") and the data of the node corresponding to the operand of the rotation matrix T indicated by the tag 328 (i.e., Operand: Twiddle matrix, T). The data of the node indicated by the tag 316 is, for example, "(I@DFT)*T".

[0048] The operator * indicated by the tag 302 (i.e., Operator: *) indicates that the "*" operation is performed on the data of the associated leaf nodes below the node corresponding to the tag 302, that is, the "*" operation is performed on the data of the node corresponding to the tag 312 ("I@P@I") and the data of the node corresponding to the tag 316 ("(I@DFT)*T"). The data of the node indicated by the tag 302 is, for example, (I@P@I)*(I@DFT)*T.

[0049] By adopting the above means, the present disclosure can simplify the formula for complex calculations, and the above simplification is more friendly to hardware, which is conducive to improving the operation speed and reducing the occupation of computing resources.

[0050] At step 204, the computing device 100 generates symbolic representation information for the operation based on the generated matrix sequence and the input symbol vector. For example, the computing device 100 multiplies the current sub-matrix in the multiple sub-matrices included in the matrix sequence by the input symbol vector corresponding to the current sub-matrix to generate symbolic representation information for the current-level operation; and multiplies the next sub-matrix in the multiple sub-matrices included in the matrix sequence by the input symbol vector corresponding to the next sub-matrix to generate symbolic representation information for the next-level operation, where the input symbol vector corresponding to the next sub-matrix is generated by the product of the current sub-matrix and the input symbol vector corresponding to the current sub-matrix, and so on, to generate symbolic representation information for each-level operation. By adopting the above means, the present disclosure can analyze the tensor operation algorithm to be operated in a symbolic form (e.g., computer algebra method).

[0051] Regarding the input symbol vector, it is the symbolic expression of the input data to be operated. Regarding the symbolic representation information for the operation, it is, for example, a multi-level symbolic expression in the form of an array or a list.

[0052] The following combines Figure 4 to illustrate the method 400 for generating symbolic representation information for the operation. Figure 4A schematic diagram of a method 400 for generating symbolic representation information for operations according to an embodiment of the present disclosure is shown. Reference numeral 402 indicates the generated matrix sequence, and the matrix sequence includes, for example, a plurality of sub-matrices, such as sub-matrices Reference numeral 404 indicates an input symbol matrix, which includes a plurality of input symbol vectors Each sub-matrix in the matrix sequence is multiplied by the input symbol vector corresponding to each sub-matrix (for example, sub-matrix 412 is multiplied by symbol vector 414) to generate symbolic representation information 406 for each level of operation. As Figure 4 shown, sub-matrix is multiplied by the input symbol vector corresponding to sub-matrix (as indicated by reference numeral 422) to generate symbolic representation information for the first level of operation. As Figure 4 shown, reference numeral 416 indicates the result of multiplying sub-matrix 412 by symbol vector 414, that is, the symbolic representation information for the first level of operation. The product of sub-matrix and input symbol vector generates the input symbol vector corresponding to sub-matrix as indicated by reference numeral 423; sub-matrix is then multiplied by the input symbol vector corresponding to sub-matrix to generate symbolic representation information for the second level of operation; and so on until symbolic representation information for each level of operation is generated. It should be understood that the product of the current sub-matrix and the input symbol vector corresponding to the current sub-matrix generates the input symbol vector corresponding to the next sub-matrix.

[0053] For example, the computing device 100 determines whether there are still sub-matrices in the matrix sequence that have not been converted into symbolic representation information; if the computing device 100 determines that there are still sub-matrices in the matrix sequence that have not been converted into symbolic representation information, the unconverted sub-matrix is multiplied by the corresponding input symbol vector matrix to generate symbolic representation information; if the computing device 100 determines that there are no sub-matrices in the matrix sequence that have not been converted into symbolic representation information, the corresponding steps in method 500 below are executed for the overall intermediate representation information to generate assembly code for the operator of the tensor operation.

[0054] At step 206, the computing device 100 converts the generated symbolic representation information into abstract syntax tree information and intermediate representation information. For example, as Figure 4As shown, the computing device 100 parses the expression of the symbolic representation information 406 via the AST parsing algorithm to generate abstract syntax tree information. By adopting the above means, the present disclosure can convert the symbolic representation information into abstract syntax tree information to facilitate the automatic generation of assembly code for operators of tensor operations. In some embodiments, the computing device 100 stores information about a matrix sequence in a tree-shaped data structure. The information about the matrix sequence at least includes: a plurality of sub-matrices included in the matrix sequence, operators and operands associated with the sub-matrices.

[0055] A method for converting the generated symbolic representation information into abstract syntax tree information includes: based on the identification information of the symbolic representation information, determining whether the symbolic representation information is an algebraic expression; in response to determining that the symbolic representation information is an algebraic expression, converting the symbolic representation information into abstract syntax tree information via an abstract syntax tree parsing algorithm; and in response to determining that the symbolic representation information is not an algebraic expression, generating intermediate expression information based on the generated symbolic representation information. The following is combined with Figure 6 Description of the method 600 for converting the generated symbolic representation information into abstract syntax tree information is omitted here.

[0056] At step 208, the computing device 100 generates assembly code for an operator of tensor operations based on the basic description information included in the generated abstract syntax tree information and intermediate expression information.

[0057] The assembly code for an operator of tensor operations is, for example, the assembly code for an operator of fast Fourier transform, that is, the assembly code for completing a Fourier transform of a predetermined size and a predetermined transform type. In some embodiments, the tensor operation may also be a sine operation or a cosine operation, etc.

[0058] In some embodiments, the computing device 100 obtains the abstract syntax tree information and intermediate expression information generated in step 206; traverses the abstract syntax tree information to generate corresponding assembly code; and assembles the corresponding assembly code based on the basic description information included in the intermediate expression information. By adopting the above means, the present disclosure can select different paths for generating assembly code when assembling the basic description information included in the abstract syntax tree information and intermediate expression information.

[0059] For example, first look at the types of basic description information included in the intermediate expression information (the typical types of basic description information include three types: permutation, direct Fourier transform DFT, and rotation factor). Given that the corresponding running behaviors of the three types of basic description information are quite different and the corresponding assembly expression ways are quite different, three independent modules are needed to construct the corresponding target instruction sequences respectively. Specifically, for example, the computing device 100 generates the assembly code of the operator for tensor operation based on the basic description information (such as vector granularity, block granularity, step size, etc.) and the type of basic description information. For example, if it is determined that the type of basic description information is permutation, the target assembly instruction sequence is constructed based on the groups supported on the hardware and the move instructions. If it is determined that the type of basic description information is DFT, the target assembly instruction sequence is constructed based on the operation instructions and the move instructions. If it is determined that the type of basic description information is rotation factor, the target assembly instruction sequence is constructed based on the operation function and the move function of the hardware.

[0060] In the above solution, a matrix sequence is generated by parsing the input information, and a symbolic representation information of a series of operations is generated based on the matrix sequence and the input symbol vector; then, the symbolic representation information for the operation is generated based on the generated matrix sequence and the input symbol vector; then, the generated symbolic representation information is converted into abstract syntax tree information and intermediate expression information; and based on the basic description information included in the generated abstract syntax tree information and intermediate expression information, the assembly code of the operator for tensor operation is generated. The present disclosure can adopt symbolic operation and AST parsing technologies, and can dynamically generate various intermediate small code segments, avoiding the problem that in the traditional tensor operation operator generation schemes such as FFT, small code segment templates still need to be manually written for specialization and assembly, and improving the development efficiency of the acceleration library for tensor operation.

[0061] In addition, in the traditional scheme for generating tensor operation operators, although the transformations optimized for various sizes and types can be pre-compiled and linked into the library. However, it is impossible to put all possible sizes into the library, otherwise the volume of the library will expand infinitely. For some less common transformations, the performance of the GPU cannot be fully released. And the present disclosure generates the assembly code directly from the hardware characteristics reflected by the AST and the basic description information, avoiding the problem that it is difficult to make full use of the GPU characteristics due to the limited hardware abstraction of the high-level language, so that the generated code can make full use of the GPU for high-performance FFT transformation. Therefore, the present disclosure can make full use of the hardware performance of each processor and can efficiently generate an operator library for high-performance tensor operation for the GPU.

[0062] The following will be combined with Figure 5 Describe method 500 for generating the assembly code of the operator for tensor operation according to some embodiments of the present disclosure.Figure 5 FIG. 500 is a flowchart of a method for generating assembly code for an operator for tensor operations according to an embodiment of the present disclosure. It should be understood that the method 500 can be executed, for example, at Figure 8 the computing device 800 described. It can also be executed at Figure 1 the computing device 100 described. It should be understood that the method 500 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.

[0063] At step 502, the computing device 100 performs machine-independent code optimization. For example, the computing device 100 performs multi-level machine-independent code optimization, such as eliminating common expressions, etc.

[0064] In some embodiments, if the computing device 100 determines that there is no un-converted sub-matrix in the matrix sequence, that is, all sub-matrices have been converted, machine-independent code optimization is performed on the overall intermediate expression information. In some embodiments, if the computing device 100 determines that there are still un-converted sub-matrices in the matrix sequence, additional intermediate expression information converted based on the symbolic representation information generated by the current un-converted sub-matrices is added, so that when it is determined that there is no un-converted sub-matrix in the matrix sequence, machine-independent code optimization is performed on the overall intermediate expression information.

[0065] At step 504, the computing device 100 performs register allocation. The intermediate expression information conforms to Static Single Assignment (SSA), that is, a variable name can be assigned only once.

[0066] At step 506, the computing device 100 converts the operation indicated in the intermediate expression information into corresponding assembly instructions.

[0067] At step 508, the computing device 100 parses the abstract syntax tree information to generate intermediate expression information.

[0068] At step 510, the computing device 100 performs scheduling optimization on the generated intermediate expression information for generating assembly code for an operator for fast Fourier transform.

[0069] By adopting the above means, the present disclosure can further improve the development efficiency of the FFT acceleration library. It should be understood that the IR of the present disclosure can be either the expression information from the AST or the expression way of the basic description information. The present disclosure can take different paths during the generation of the assembly code according to whether the IR is obtained from the AST or from the basic description information. After taking different paths, common methods are used for optimization to finally generate the assembly code for the operator of the fast Fourier transform.

[0070] Regarding the method for generating the assembly code for the operator of tensor operations, in some other embodiments, it includes, for example: the computing device 100 performs optimization of machine-independent code; performs register allocation; converts the operation operations indicated in the intermediate expression information into corresponding assembly instructions; converts the abstract syntax tree information into assembly code for scheduling; and generates the assembly code for the operator of the fast Fourier transform.

[0071] The following will be combined with Figure 6 Describe the method 600 for converting the generated symbolic representation information into abstract syntax tree information according to an embodiment of the present disclosure. Figure 6 The flowchart of the method 600 for converting the generated symbolic representation information into abstract syntax tree information according to an embodiment of the present disclosure is shown. It should be understood that the method 600 can be executed, for example, at Figure 8 the computing device 800 described. It can also be executed at Figure 1 the computing device 100 described. It should be understood that the method 600 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.

[0072] At step 602, the computing device 100 determines whether the symbolic representation information is an algebraic expression based on the identification information of the symbolic representation information.

[0073] At step 604, if the computing device 100 determines that the symbolic representation information is an algebraic expression, the symbolic representation information is converted into abstract syntax tree information via the abstract syntax tree parsing algorithm.

[0074] At step 606, if the computing device 100 determines that the symbolic representation information is not an algebraic expression, intermediate expression information is generated based on the generated symbolic representation information.

[0075] By adopting the above means, the present disclosure can write the part related to hardware operations as intermediate expression information, and convert the direct mathematical operations into abstract syntax tree information, thereby facilitating making more full use of the hardware performance of each processor during the process of generating the operator of tensor operations (such as the fast Fourier transform).

[0076] The following will be described in conjunction with Figure 7 a method 700 for generating a matrix sequence according to an embodiment of the present disclosure. Figure 7 FIG. shows a flowchart of a method 700 for generating a matrix sequence according to an embodiment of the present disclosure. It should be understood that the method 700 can be executed, for example, at Figure 8 the computing device 800 described. It can also be executed at Figure 1 the computing device 100 described. It should be understood that the method 700 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.

[0077] At step 702, the computing device 100 obtains the size and transformation type of the input data to be transformed.

[0078] Regarding the transformation type, it is, for example, a transformation method of DFT. There are three types of DFT transformation methods, namely Real-complex, complex-complex, and complex-real. If the transformation type is different, the algorithm expression of DFT is different.

[0079] At step 704, the computing device 100 performs prime factorization on the size of the input data to be decomposed into multiple prime factors.

[0080] It should be understood that the Cooley-Tukey fast Fourier transform (FFT) algorithm is a very common algorithm for accelerating the discrete Fourier transform (DFT). The essence of the Cooley-Tukey fast Fourier algorithm is to recursively split an N-point DFT of a composite number of points into k m-point DFTs.

[0081] The computing device 100 can perform prime factorization on the size of the input data to be transformed in various ways. For example, but not limited to, using the Pollard Rho fast factorization method to decompose the size of the input data into a form of several prime factors multiplied together.

[0082] For example, the computing device 100 determines the smallest prime number as the current prime number; if it is determined that the size of the input data to be transformed is equal to the current prime number, it is determined that the prime factorization ends; if it is determined that the size of the input data to be transformed is greater than the current prime number and the size of the input data to be transformed is divisible by the current prime number, the current prime number is output as one of the prime factors; the quotient of the size of the input data to be transformed divided by the current prime number is obtained to determine whether the quotient is equal to the current prime number; if it is determined that the quotient is equal to the current prime number, the quotient is output as one of the prime factors; if the size of the input data to be transformed is greater than the current prime number and the size of the input data is not divisible by the current prime number, the current prime number is incremented by 1 to update the current prime number; the determination of whether the size of the input data to be transformed is equal to the updated current prime number is repeated. It should be understood that other methods can also be used for prime factorization of the size of the input data to be transformed.

[0083] At step 706, the computing device 100 performs matrix factorization based on multiple prime factors and the transformation type; to generate a matrix sequence. The way of matrix factorization usually depends on what kind of data transformation needs to be performed. For example, the DFT can be decomposed into a matrix sequence including multiple sub-matrices by means of recursive matrix factorization.

[0084] The DFT is a linear transformation from a given input vector (i.e., the sampling sequence of the signal) to an output vector (the frequency spectrum, all of whose elements are complex numbers). The algorithm of the FFT is described below with reference to formula (1).

[0085]

[0086] In the above formula (1), represents the given input vector. represents the output vector. DFT N represents the N-point discrete Fourier transform. The algorithm of the DFT is described below with reference to formula (2).

[0087]

[0088] In the above formula (2), k represents the radix. I m and I k respectively represent the identity matrix. represents the diagonal matrix of the rotation factor. represents the Kronecker product. N represents the size of the input data to be transformed. DFT k represents the k-point discrete Fourier transform. DFT m represents the m-point discrete Fourier transform. DFT N can be factorized into 4 sub-matrices and A matrix sequence formed

[0089] It should be understood that the matrix sequence decomposed by the above formula (2) also includes DFTs of smaller sizes m and DFTs k , therefore, matrix factorization can also be recursively performed for DFTs m and DFTs k until k and m are prime numbers, and thus, a matrix sequence including more sub - matrices can be finally obtained

[0090] By adopting the above means, the present disclosure can decompose the fast Fourier transform algorithm into multiple smaller - sized sub - matrices that can be calculated separately

[0091] Figure 8 FIG. schematically shows a block diagram of an electronic device (or computing device) 800 suitable for implementing the embodiments of the present disclosure. The device 800 can be a device for implementing the methods 200, 500 to 700 shown in Figure 2 , Figures 5 to 7 As shown in the figure, the device 800 includes a central processing unit (CPU) 801, which can execute various appropriate actions and processes according to the computer program instructions stored in the read - only memory (ROM) 802 or the computer program instructions loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM, various programs and data required for the operation of the device 800 can also be stored. The CPU, ROM, and RAM are connected to each other via a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804

[0092] Multiple components in the device 800 are connected to the input / output (I / O) 805, including: an input unit 806, an output unit 807, a storage unit 808. The central processing unit 801 executes the various methods and processes described above, for example, executes the methods 200, 500 to 700. For example, in some embodiments, the methods 200, 500 to 700 can be implemented as computer software programs, which are stored in a machine - readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM and / or the communication unit 809. When the computer program is loaded into the RAM and executed by the CPU, one or more operations of the methods 200, 500 to 700 described above can be performed. Alternatively, in other embodiments, the CPU can be configured to execute one or more actions of the methods 200, 500 to 700 by any other suitable means (e.g., by means of firmware)

[0093] It should be further noted that the present disclosure can be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0094] The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0095] The computer-readable program instructions described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0096] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the C language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0097] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block step of the flowchart and / or block diagram, and combinations of blocks steps in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0098] These computer - readable program instructions can be provided to the processing unit of a processor in a voice interaction device, a general - purpose computer, a special - purpose computer, or other programmable data - processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data - processing device, a device is produced that implements the functions / actions specified in one or more of the block steps of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause the computer, programmable data - processing device, and / or other devices to operate in a specific manner. Thus, the computer - readable medium storing the instructions includes a manufacture, which includes instructions that implement various aspects of the functions / actions specified in one or more of the block steps of the flowchart and / or block diagram.

[0099] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more of the blocks of the flowchart and / or block diagram.

[0100] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0101] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

[0102] The above are only alternative embodiments of the present disclosure and are not intended to limit the present disclosure. Various changes and modifications can be made to the present disclosure by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for generating an operator for tensor operations, comprising: Parsing input information for generating a matrix sequence; Generating symbolic representation information for operations based on the generated matrix sequence and input symbolic vectors, wherein generating symbolic representation information for operations based on the generated matrix sequence and input symbolic vectors includes: multiplying a current sub-matrix among a plurality of sub-matrices included in the matrix sequence by an input symbolic vector corresponding to the current sub-matrix to generate symbolic representation information for the current-level operation; and multiplying a next sub-matrix among the plurality of sub-matrices included in the matrix sequence by an input symbolic vector corresponding to the next sub-matrix to generate symbolic representation information for the next-level operation; Converting the generated symbolic representation information into abstract syntax tree information and intermediate representation information; and Generating assembly code for an operator for tensor operations based on basic description information included in the generated abstract syntax tree information and intermediate representation information.

2. The method according to claim 1, wherein the tensor operation is a fast Fourier transform, and the basic description information at least includes: Description parameters of the size of a matrix, operation operations associated with the matrix, and hardware resources required for operations associated with the matrix.

3. The method according to claim 1, further comprising: Storing information about the matrix sequence in a tree-shaped data structure, where the information about the matrix sequence at least includes: a plurality of sub-matrices included in the matrix sequence, operators and operands associated with the sub-matrices.

4. The method according to claim 1, wherein converting the generated symbolic representation information into abstract syntax tree information and intermediate representation information includes: Determining whether the symbolic representation information is an algebraic expression based on identification information of the symbolic representation information; In response to determining that the symbolic representation information is an algebraic expression, converting the symbolic representation information into abstract syntax tree information via an abstract syntax tree parsing algorithm; And In response to determining that the symbolic representation information is not an algebraic expression, generating intermediate representation information based on the generated symbolic representation information.

5. The method according to claim 2, wherein generating a matrix sequence includes: Obtaining the size and transformation type of input data to be transformed; Performing prime factorization on the size of the input data to decompose it into multiple prime factors; And Performing matrix factorization based on the multiple prime factors and the transformation type; To generate a matrix sequence.

6. The method according to claim 1, wherein the input symbolic vector corresponding to the next sub-matrix is generated by the product of the current sub-matrix and the input symbolic vector corresponding to the current sub-matrix.

7. The method according to claim 3, wherein the operators and operands associated with the sub-matrices are generated based on parsing of a specific type of sub-matrix, and the specific type of sub-matrix is determined based on the size of the sub-matrix, the computational behavior related to the sub-matrix, and limit data of hardware resources.

8. The method according to claim 2, wherein generating assembly code for an operator for tensor operations includes: Performing optimization of machine-independent code; Performing register allocation; Converting the operation operations indicated in the intermediate representation information into corresponding assembly instructions; Parsing the abstract syntax tree information to generate intermediate representation information; Schedule optimization is performed on the generated intermediate representation information for use in generating assembly code for an operator for fast Fourier transform.

9. The method according to claim 2, wherein generating assembly code for an operator for tensor operations includes: Performing machine-independent code optimization; Performing register allocation; Converting the operation operations indicated in the intermediate representation information into corresponding assembly instructions; Converting the abstract syntax tree information into assembly code for scheduling; And Generating assembly code for an operator for fast Fourier transform.

10. The method according to claim 2, wherein generating assembly code for an operator for tensor operations includes: Generating assembly code for an operator for fast Fourier transform based on the basic description information and the type of the basic description information, the type of the basic description information including: permutation, discrete Fourier transform, and twiddle factor.

11. A computing device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a machine, it executes the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Operation data generation method and device and related product

    CN112183735A

  • Code compiling method and code complier

    WO2017035748A1