Method and device for converting CUDA C language into Triton language and medium

Through steps such as parsing, abstract syntax tree construction and semantic analysis, automatic conversion from CUDA C to Triton language is realized, solving the problem of time-consuming, labor-intensive and error-prone problems of manual conversion, improving conversion efficiency and ensuring functional equivalent.

CN119987786APending Publication Date: 2025-05-13SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064626.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Due to the significant differences in CUDA C and Triton in terms of syntax, semantics and parallel computing models, the conversion from CUDA C to Triton usually requires manual operation, which is time-consuming and error-prone, especially when dealing with complex deep learning primitives.

Method used

A conversion method from CUDA C language to Triton language is proposed, including obtaining CUDA C code, parsing and building an abstract syntax tree, performing semantic analysis, determining mapping patterns, loading mapping rules, performing keyword equivalent mapping, and finally performing performance tests to ensure functional equivalent.

Benefits of technology

By automating the conversion process, the conversion efficiency is significantly improved, the time and labor cost required for manual conversion is reduced, and the error rate during the conversion process is reduced, ensuring that the converted Triton code remains functionally equivalent to the original CUDA C code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987786A_ABST
    Figure CN119987786A_ABST
Patent Text Reader

Abstract

The invention discloses a method and equipment for converting a CUDA C language into a Triton language and a medium, and relates to the technical field of electric digital data processing. The method comprises the following steps: analyzing a CUDA C code so as to decompose the CUDA C code into a plurality of lexical units; constructing an abstract syntax tree corresponding to the lexical unit; performing semantic analysis on the abstract syntax tree to determine a keyword corresponding to each node in the abstract syntax tree; determining a mapping mode corresponding to the CUDA C code according to the logic function of the CUDA C code, and loading a mapping rule between a CUDA C language and a Triton language based on the mapping mode; extracting the meaning of a target key element corresponding to the keyword through a mapping rule, and performing equivalent mapping on the keyword according to the meaning of the target key element so as to convert the CUDA C code into a Triton code of which the meaning of the corresponding keyword is the same as the meaning of the target key element; and carrying out performance test on the Triton code, and determining whether the Triton code keeps functional equivalence with the CUDA C code or not according to an obtained test result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic digital data processing, and in particular to a method, device and medium for converting CUDA C language to Triton language. Background Art

[0002] CUDA C (Compute Unified Device Architecture C) is a parallel computing platform and programming model launched by NVIDIA, which allows software developers to use NVIDIA's graphics processing unit (GPU) for general computing. CUDA C language is based on C / C++ and extends a series of features for parallel computing, such as memory management, thread management and synchronization mechanisms, allowing developers to fully utilize the multi-core parallel computing capabilities of GPU.

[0003] At present, in the field of high-performance computing and deep learning, although CUDA C is still widely used, new programming languages ​​such as Triton are gradually gaining attention and application. Compared with CUDA C, Triton provides higher productivity and greater flexibility. It simplifies the complexity of parallel programming, allowing developers to focus more on the implementation and optimization of deep learning algorithms without paying too much attention to the underlying parallel computing details. It is not limited by hardware and can theoretically run on any hardware that supports GPU computing.

[0004] Converting CUDA C to Triton can significantly improve development efficiency and cross-platform compatibility of code. However, due to the significant differences between CUDA C and Triton in syntax, semantics, and parallel computing models, this programming language conversion usually needs to be done manually, which is not only time-consuming and labor-intensive, but also prone to errors, especially when dealing with complex deep learning primitives. Summary of the invention

[0005] In order to solve the above problems, this application proposes a method for converting CUDA C language to Triton language, including:

[0006] Obtaining a CUDA C code to be converted, and parsing the CUDA C code to decompose the CUDA C code into a plurality of lexical units;

[0007] According to the grammatical rules corresponding to the CUDA C code, construct an abstract syntax tree corresponding to the lexical unit; wherein the abstract syntax tree includes a plurality of nodes, each node corresponding to a lexical unit;

[0008] Performing semantic analysis on the abstract syntax tree to determine a keyword corresponding to each node in the abstract syntax tree; wherein the keyword is used to represent a key element in the CUDA C code;

[0009] Determine a mapping mode corresponding to the CUDA C code according to a logical function of the CUDA C code, and load a mapping rule between the CUDA C language and the Triton language based on the mapping mode;

[0010] Extracting the target key element meaning corresponding to the keyword through the mapping rule, and performing equivalent mapping on the keyword according to the target key element meaning, so as to convert the CUDA C code into a Triton code having the same meaning as the target key element as the corresponding keyword;

[0011] A performance test is performed on the Triton code, and based on the obtained test results, it is determined whether the Triton code is functionally equivalent to the CUDA C code.

[0012] In one implementation of the present application, the keyword includes at least one or more of the following: a kernel function, a variable, a scope of the variable, a thread index, and a memory access mode. According to the meaning of the target key element, the keyword is equivalently mapped, specifically including:

[0013] According to the meaning of the target key element, the kernel function is equivalently mapped to the Triton kernel function in the Triton code;

[0014] Mapping the variable to a variable of a specified type in the Triton language, and retaining the scope of the variable;

[0015] Passing the thread index to the Triton kernel function through a function parameter, so as to implement parallel execution of each thread task in the Triton kernel function by calling the thread index;

[0016] The memory access mode is equivalently mapped to a target memory access mode of the same mode type in the Triton language; wherein the memory access mode includes global memory access, array access, and shared access.

[0017] In one implementation of the present application, the memory access mode is equivalently mapped to a target memory access mode of the same mode type in the Triton language, specifically including:

[0018] When the memory access mode is the global memory access and the array access, the global memory access and the array access are equivalently mapped to pointer access in the Triton language;

[0019] When the memory access mode is the shared access, the shared access is equivalently mapped to a shared array access or a pointer access.

[0020] In an implementation of the present application, the mapping type includes centralized mapping and decentralized mapping, and the mapping mode corresponding to the CUDA C code is determined according to the logical function of the CUDA C code, specifically including:

[0021] If the number of the CUDA C code is single, determining that the mapping mode corresponding to the CUDA C code is distributed mapping;

[0022] If there are multiple CUDA C codes, split the CUDA C code into a plurality of code segments according to the code structure of the CUDA C code; wherein the code segments are divided according to function types;

[0023] For different CUDA C codes, analyzing statement attributes corresponding to each statement in the code segment to determine similarities between different statements;

[0024] Extracting the specified statements whose similarity is greater than a preset similarity, obtaining the variable dependency relationship between the specified statements, and determining whether the output values ​​generated by the specified statements when executed are the same according to the variable dependency relationship, so as to determine whether there are functionally equivalent code segments in the CUDA C code;

[0025] If so, it is determined that the mapping mode corresponding to the equivalent code segment is centralized mapping, and the mapping modes corresponding to other code segments in the CUDA C code except the equivalent code segment are distributed mapping.

[0026] In one implementation of the present application, for different CUDA C codes, statement attributes corresponding to each statement in the code segment are analyzed to determine the similarity between different statements, specifically including:

[0027] For different CUDA C codes, obtain statement attributes corresponding to each statement in the code segment; wherein the statement attributes at least include parameter categories and parameter meanings contained in the statement, and statement control scope;

[0028] The similarity between the attributes of the sentences is calculated, and the similarity is used as the similarity between different sentences.

[0029] In one implementation of the present application, obtaining the variable dependency relationship between the specified statements, and determining whether the output values ​​generated by the specified statements when executed are the same according to the variable dependency relationship, specifically includes:

[0030] Determine whether a variable included in the specified statement has a corresponding variable definition and variable reference in the code snippet in which the variable is located;

[0031] If so, directly determine whether the output values ​​generated by the specified statement when executed are the same;

[0032] If not, determine the variable dependency relationship corresponding to the variable according to the variable type corresponding to the variable, and add the code segment where the dependent variable having the variable dependency relationship with the variable is located to the specified statement to obtain a complete execution statement corresponding to the specified statement;

[0033] By executing the complete execution statement, it is determined whether the corresponding output values ​​are the same.

[0034] In one implementation of the present application, based on the mapping mode, the mapping rules between the CUDA C language and the Triton language are loaded, specifically including:

[0035] Based on the centralized mapping mode, identifying common keywords in the equivalent code segments, and generating keyword templates corresponding to the CUDA C code according to the common keywords;

[0036] The mapping rules between the CUDA C language and the Triton language are loaded, and based on the mapping rules, the keyword templates are reused to achieve batch conversion of a plurality of the CUDA C codes.

[0037] In one implementation of the present application, after converting the CUDA C code into the corresponding Triton code, the method further includes:

[0038] Determining the number of thread indices of parallel threads in the Triton code;

[0039] When the number of thread indexes is greater than a preset value, the memory data structure and parallel algorithm of the Triton code are optimized to improve the execution efficiency of the Triton code.

[0040] The embodiment of the present application provides a CUDA C language to Triton language conversion device, the device comprising:

[0041] at least one processor;

[0042] and, a memory communicatively coupled to the at least one processor;

[0043] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a CUDA C language to Triton language conversion method as described in any one of the above items.

[0044] The embodiment of the present application provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as follows:

[0045] A method for converting CUDA C language to Triton language as described in any of the above items.

[0046] The CUDA C language to Triton language method proposed in this application can bring the following beneficial effects:

[0047] By parsing the code, building an abstract syntax tree, and performing semantic analysis, we have achieved automatic conversion from CUDA C to Triton code, significantly improving conversion efficiency and reducing the time and labor costs required for manual conversion. Secondly, by building an abstract syntax tree and performing semantic analysis, we can accurately capture the key elements in the CUDA C code, ensuring that the converted Triton code is functionally equivalent to the original CUDA C code, reducing the error rate during the conversion process. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0049] Figure 1 A flowchart of a method for converting CUDA C language to Triton language provided in an embodiment of the present application;

[0050] Figure 2 A schematic diagram of the structure of a CUDA C language to Triton language conversion device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0052] CUDA C is an extension of C, and most of its syntax is the same as C. In CUDA C, a kernel function is a function executed on the GPU and is defined by adding the __global__ identifier before the function. The thread configuration information is defined by <<<grid,block> >> syntax, where grid and block represent the size of the thread grid and thread block, respectively.

[0053] Triton is a programming language similar to Python with a concise syntax structure. The built-in compiler compiles the code into efficient machine code to run on modern GPU hardware. The functions in the code can be compiled and accelerated through the @triton.jit decorator. Triton emphasizes the operation of tensors. It simplifies the difficulty of parallel programming by introducing high-level abstraction mechanisms, automatically handles memory management and shared memory, reduces the complexity of memory management for developers, and provides automatic optimization functions. It can optimize according to the specific situation of the code to improve operating efficiency. When writing Triton code, users do not need to pay attention to the underlying hardware characteristics, but only need to describe the specific calculation logic. The use of hardware such as shared memory, global memory, etc. is completely specified by the compiler during the compilation process. The compilation process of Triton code is as follows:

[0054] 1. Convert the kernel written by the user in Python or the TritonKernel generated by Inductor in PyTorch 2.0 to the corresponding Triton IR (intermediate representation).

[0055] 2. Through various passes, Triton IR is gradually converted and optimized to Triton GPU IR. At the Triton IR level, the compiler can apply some advanced optimizations such as dead code elimination, constant folding, etc.

[0056] 3. Convert Triton GPU IR to LLVM IR step by step. LLVM IR is a platform-independent intermediate representation that can be used on different hardware platforms.

[0057] 4. For NVIDIA GPU, LLVM IR will eventually be compiled into cubin.

[0058] In order to improve development efficiency and enhance the cross-platform compatibility of the code, the embodiment of the present application proposes a CUDA C language to Triton language conversion method based on the above Triton coding, so as to convert the CUDA C language to the Triton language. The technical solutions provided by the embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0059] like Figure 1 As shown, a method for converting CUDA C language to Triton language provided in an embodiment of the present application includes:

[0060] S101: Obtain CUDA C code to be converted, and parse the CUDA C code to decompose the CUDA C code into a plurality of lexical units.

[0061] Get the CUDA C code to be converted, parse the CUDA C code through the parser, and decompose the CUDA C code into several lexical units. Lexical units refer to the basic elements that constitute the source code, such as keywords (such as __global__, void, float, etc.), identifiers (such as vectorAdd, A, B, etc.), operators (such as *, +, etc.) and separators.

[0062] S102: constructing an abstract syntax tree corresponding to the lexical unit according to the grammatical rules corresponding to the CUDA C code; wherein the abstract syntax tree includes a plurality of nodes, each node corresponding to a lexical unit.

[0063] In the syntax analysis phase, lexical units are organized into a tree structure to reflect the grammatical structure of the source code. For CUDA C code, building an abstract syntax tree first requires identifying grammatical rules (such as statements, expressions, declarations, etc.) and mapping these grammatical rules to the tree structure. The abstract syntax tree includes multiple nodes, each of which corresponds to a lexical unit.

[0064] In one embodiment, the following CUDAC code is parsed:

[0065]

[0066] After identifying each lexical unit, we can finally identify that vectorAdd is a function declaration with a return type of void, a function name of vectorAdd, a parameter list including four parameters (float*A, float*B, float*C, and int numElements), and the function has a __global__ attribute, indicating that it is a CUDA kernel function. Using the grammar rules of CUDA C, define a node class for each grammar construct. These classes will contain the node type, child node list, and any other relevant information. Write a recursive descent parser or use a parser generation tool to parse the lexical unit stream according to the grammar rules and build an abstract syntax tree.

[0067] S103: Perform semantic analysis on the abstract syntax tree to determine a keyword corresponding to each node in the abstract syntax tree; wherein the keyword is used to represent a key element in the CUDA C code.

[0068] Semantic analysis is performed based on the abstract syntax tree output by the above code parsing. The semantic analyzer checks whether the type of each variable is consistent and whether the parameter type of each function call matches the function declaration. At the same time, for each node, its type is checked to determine whether it corresponds to a keyword. If the node corresponds to a keyword, then the keyword needs to be identified. Keywords are used to represent key elements in CUDA C code, including at least one or more of the following: kernel function, variable, scope of the variable, thread index, and memory access mode.

[0069] S104: Determine a mapping mode corresponding to the CUDA C code according to the logical function of the CUDA C code, and load a mapping rule between the CUDA C language and the Triton language based on the mapping mode.

[0070] There are differences in the logical functions between different CUDA C codes. According to the logical functions of the CUDA C code, the embodiments of the present application provide two different mapping modes, namely, distributed mapping and centralized mapping. In different mapping modes, it is necessary to load the mapping rules between the CUDA C language and the Triton language to achieve flexible conversion between the CUDA C code and the Triton code. It should be noted that the mapping rules under different mapping modes have different corresponding execution logics. In the distributed mapping mode, each CUDA C code needs to be equivalently mapped in turn and converted into the corresponding Triton code. In the centralized mapping, it is necessary to call a mapping template that can be shared to achieve batch conversion of CUDA C code to Triton code, which helps to improve the code conversion efficiency.

[0071] In one embodiment, if the number of CUDA C codes is single, the conversion logic between codes only needs to be executed once, and there is no need to perform centralized mapping on the codes. Therefore, the mapping mode corresponding to the CUDA C code is distributed mapping.

[0072] If there are multiple CUDA C codes, it is necessary to perform a logical function analysis on each CUDA C code to determine whether the functions and keywords of different codes are the same, so as to determine whether there are code segments in the CUDA C code that can be centrally mapped.

[0073] Specifically, according to the code structure of the CUDA C code, the CUDA C code is split into several code segments; wherein the code segments are divided according to the function types, for example, the code segments can be divided into kernel functions, memory operation functions, auxiliary functions, etc.

[0074] For different CUDAC codes, the statement attributes corresponding to each statement in the code segment are analyzed to determine the similarity between different statements. The statement attributes here include at least the parameter category and parameter meaning contained in the statement, and the statement control scope. The statement control scope refers to the statement type in which the code statement is located, such as loop statements, branch statements, etc. Calculate the similarity between statement attributes, and use the similarity as the similarity between different statements. It should be noted that when analyzing statement attributes, a statement can be a single statement or a code fragment consisting of multiple independent statements. The specific splitting logic can be distinguished according to the actual situation of the code. For example, if there are multiple nested statements in the code, the nested statements can be evaluated as a statement for similarity.

[0075] After obtaining the similarity between different statements through the above process, extract the specified statements whose similarity is greater than the preset similarity, and obtain the variable dependency between the specified statements. The variable dependency is used to determine whether a certain statement can independently complete the input and output of the variable. Specifically, it is determined whether the variable contained in the specified statement has a corresponding variable definition in the code snippet in which it is located. If there is a variable definition corresponding to the variable in the code snippet, the variable can be determined as an input variable. If there is a corresponding variable reference in the code snippet, the variable can be determined as an output variable. In this case, it can be determined that this code snippet can fully realize the input and output of the variable. At this time, the output values ​​generated by the specified statements in different CUDAC codes when they are executed can be directly compared.

[0076] If the input variables and output variables cannot be located directly, it is necessary to determine the variable dependency corresponding to the variable according to the variable type corresponding to the variable. The variable type here is used to indicate the location of the variable. In one case, the variable exists in the parameter list, and in another case, the variable exists in the function declaration. For different variable types, the corresponding variable dependency is different. After clarifying the variable dependency, the code segment where the dependent variable with which the variable has a variable dependency relationship is located can be added to the specified statement, and finally a complete execution statement corresponding to the specified statement is formed. At this time, by executing the complete execution statement, the input value and output value of the variable can be clarified, and then it can be determined whether the output value is the same. If the output values ​​are the same, it means that there are equivalent code segments with equivalent functions in the current CUDAC code, and the equivalent code segments are composed of statements filtered out with a similarity greater than a preset similarity and the same output value.

[0077] It should be noted that after the complete execution statement is generated, the specified statement in the CUDA C code that participates in the similarity comparison has changed. At this time, the similarity of the statement needs to be recalculated. If the similarity is still greater than the preset similarity, the complete execution statement corresponding to the above specified statement can be used as an equivalent code segment with equivalent functionality.

[0078] If there are equivalent code segments in different CUDA C codes, when the codes are equivalently mapped, the mapping mode corresponding to the equivalent code segments is centralized mapping, and the mapping mode corresponding to other code segments except the equivalent code segments is distributed mapping.

[0079] In the distributed mapping mode, the mapping rules require that each CUDA C code needs to execute the corresponding code conversion logic to complete the language conversion. In the centralized mapping mode, each CUDA C code no longer needs to execute the above equivalent mapping operation to complete the conversion of all codes. At this time, it is necessary to identify the common keywords in the equivalent code segments and generate the keyword template corresponding to the CUDA C code based on the common keywords. The mapping rules between the CUDA C language and the Triton language are loaded. The mapping rules require different CUDA C codes to reuse the keyword templates when performing language conversion. In this way, for equivalent code segments, only one keyword template needs to be called to achieve batch conversion of the corresponding code, which greatly improves the code conversion efficiency.

[0080] S105: extracting the target key element meaning corresponding to the keyword through the mapping rule, and performing equivalent mapping on the keyword according to the target key element meaning, so as to convert the CUDA C code into Triton code whose corresponding keyword meaning is the same as the target key element meaning.

[0081] The equivalent mapping of keywords is not a simple replacement of keywords, but a code reconstruction process based on the meaning of keywords. That is to say, in the equivalent mapping process, it is first necessary to extract the target key element meaning corresponding to the keyword through the mapping rules. For example, for the keyword vectorAdd, which exists in the form of an identifier, its corresponding target key element meaning is the kernel function. After extracting the target key element meaning, the keyword can be equivalently mapped, and the CUDA C code can be converted into Triton code with the same meaning as the corresponding keyword and the target key element through the compiler of the Triton language. For example, convert vectorAdd into a Triton kernel function in Triton (a function decorated with @triton.jit).

[0082] Specifically, according to the meaning of the target key element, the kernel function is equivalently mapped to the Triton kernel function in the Triton code. At the same time, the variables need to be equivalently mapped to the variables of the specified type in the Triton language, and the scope of the variables is retained. Both CUDA C and Triton can implement parallel computing of multiple threads through parallel models. Each thread corresponds to a thread index. Each thread will execute the code in the function body, but the value of idx will be different, depending on the position of the thread in the grid. Therefore, for the key element of thread index, the thread index needs to be passed to the Triton kernel function through the function parameter to realize the parallel execution of each thread task in the Triton kernel function by calling the thread index. By using the @triton.jit decorator, the vector_add_kernel function is compiled into code that can be efficiently executed on the GPU or other parallel hardware. When this function is called, Triton will allocate threads according to the provided grid and block size.

[0083] Memory access modes also need to be mapped during the language conversion process. Therefore, the memory access modes must be equivalently mapped to target memory access modes with the same mode type in the Triton language; among them, memory access modes include global memory access, array access, and shared access.

[0084] Different memory access modes have different corresponding conversion methods. Triton emphasizes explicit memory management and pointer operations. Therefore, for the two memory access modes of global memory access and array access, they need to be equivalently mapped to pointer access in the Triton language. For the memory access mode of shared access, specific memory allocation and synchronization mechanisms can be used to simulate the behavior of shared memory. Correspondingly, shared access also needs to be equivalently mapped to shared array access or pointer access.

[0085] S106: Perform a performance test on the Triton code, and determine whether the Triton code is functionally equivalent to the CUDA C code based on the obtained test results.

[0086] After the Triton code is generated, it needs to be optimized to improve the code performance and efficiency. First, the redundant code needs to be removed. Secondly, after removing the redundant code, it is necessary to determine whether the parallel computing performance of the current Triton code is relatively good. Specifically, the number of thread indexes of parallel threads in the Triton code is determined. When the number of indexes is greater than the preset value, it means that the number of parallel threads in the current code is large, which may have a certain impact on the execution efficiency of the code. At this time, the memory data structure and parallel algorithm of the Triton code can be optimized. By selecting a cache-friendly data structure or a more effective parallel algorithm, the execution efficiency of the Triton code can be improved.

[0087] After the optimization of the Triton code is completed, the performance test of the Triton code needs to be performed. By comparing the test results with the corresponding running results of the CUDA C code, it can be determined whether the converted Triton code is functionally equivalent to the CUDA C code. If they are consistent, it means that the converted code has the same functions, and the conversion of the CUDA C code to the Triton code is completed. At the same time, during the test process, the corresponding code test performance needs to be recorded, mainly including the code running time and the code memory occupied. Through the above code test performance, the code can be further optimized and adjusted.

[0088] In one embodiment, the CUDA C code is as follows:

[0089]

[0090] Identify the lexical units and construct an abstract syntax tree. Perform semantic analysis on the abstract syntax tree and finally identify the following information:

[0091] 1. vectorAdd is a CUDA kernel function that accepts four parameters.

[0092] 2. B and C are global memory pointers pointing to float type.

[0093] 3. numElements is an int type variable used to specify the number of elements to be processed.

[0094] 4. Use blockIdx.x and threadIdx.x in the kernel function to calculate the index idx of each thread.

[0095] 5. If idx is less than numElements, perform array access and addition.

[0096] Based on the above keyword meanings, an equivalent mapping is performed. For example, vectorAdd is mapped to a function decorated with @triton.jit, and the final converted Triton code is as follows:

[0097]

[0098] The converted Triton code has the same functions as the CUDAC code, and maintains good code performance while having parallel computing capabilities.

[0099] The above are embodiments of the method proposed in this application. Based on the same idea, some embodiments of this application also provide devices and non-volatile computer storage media corresponding to the above methods.

[0100] Figure 2 A schematic diagram of a CUDA C language to Triton language conversion device provided in an embodiment of the present application. Figure 2 As shown, including:

[0101] at least one processor; and,

[0102] at least one processor is communicatively connected to a memory; wherein,

[0103] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can execute a CUDA C language to Triton language conversion method as described in any one of the above items.

[0104] The embodiment of the present application provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as follows:

[0105] A method for converting CUDA C language to Triton language as described in any of the above items.

[0106] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0107] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0108] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0109] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0110] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0112] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0113] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0114] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0115] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0116] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for converting CUDA C language to Triton language, characterized in that: The method comprises: Obtaining a CUDA C code to be converted, and parsing the CUDA C code to decompose the CUDA C code into a plurality of lexical units; According to the grammatical rules corresponding to the CUDA C code, construct an abstract syntax tree corresponding to the lexical unit; wherein the abstract syntax tree includes a plurality of nodes, each node corresponding to a lexical unit; Performing semantic analysis on the abstract syntax tree to determine a keyword corresponding to each node in the abstract syntax tree; wherein the keyword is used to represent a key element in the CUDA C code; Determine a mapping mode corresponding to the CUDA C code according to a logical function of the CUDA C code, and load a mapping rule between the CUDA C language and the Triton language based on the mapping mode; Extracting the target key element meaning corresponding to the keyword through the mapping rule, and performing equivalent mapping on the keyword according to the target key element meaning, so as to convert the CUDA C code into a Triton code having the same meaning as the target key element as the corresponding keyword; A performance test is performed on the Triton code, and based on the obtained test results, it is determined whether the Triton code is functionally equivalent to the CUDA C code.

2. The method for converting CUDAC language to Triton language according to claim 1, characterized in that: The keywords include at least one or more of the following: kernel function, variable, scope of the variable, thread index, and memory access mode. According to the meaning of the target key element, the keywords are equivalently mapped, specifically including: According to the meaning of the target key element, the kernel function is equivalently mapped to the Triton kernel function in the Triton code; Mapping the variable to a variable of a specified type in the Triton language, and retaining the scope of the variable; Passing the thread index to the Triton kernel function through a function parameter, so as to implement parallel execution of each thread task in the Triton kernel function by calling the thread index; The memory access mode is equivalently mapped to a target memory access mode of the same mode type in the Triton language; wherein the memory access mode includes global memory access, array access, and shared access.

3. The method for converting CUDA C language to Triton language according to claim 2, characterized in that: Equivalently mapping the memory access pattern to a target memory access pattern of the same pattern type in the Triton language specifically includes: When the memory access mode is the global memory access and the array access, the global memory access and the array access are equivalently mapped to pointer access in the Triton language; When the memory access mode is the shared access, the shared access is equivalently mapped to a shared array access or a pointer access.

4. The method for converting CUDAC language to Triton language according to claim 1, characterized in that: The mapping type includes centralized mapping and decentralized mapping. According to the logical function of the CUDA C code, the mapping mode corresponding to the CUDA C code is determined, specifically including: If the number of the CUDAC code is single, determining that the mapping mode corresponding to the CUDA C code is distributed mapping; If there are multiple CUDA C codes, split the CUDA C code into a plurality of code segments according to the code structure of the CUDA C code; wherein the code segments are divided according to function types; For different CUDA C codes, analyzing statement attributes corresponding to each statement in the code segment to determine similarities between different statements; Extracting the specified statements whose similarity is greater than a preset similarity, obtaining the variable dependency relationship between the specified statements, and determining whether the output values ​​generated by the specified statements when executed are the same according to the variable dependency relationship, so as to determine whether there are functionally equivalent code segments in the CUDA C code; If so, it is determined that the mapping mode corresponding to the equivalent code segment is centralized mapping, and the mapping modes corresponding to other code segments in the CUDA C code except the equivalent code segment are distributed mapping.

5. The method for converting CUDA C language to Triton language according to claim 4, characterized in that: For different CUDA C codes, the statement attributes corresponding to each statement in the code segment are analyzed to determine the similarity between different statements, specifically including: For different CUDA C codes, obtain statement attributes corresponding to each statement in the code segment; wherein the statement attributes at least include parameter categories and parameter meanings contained in the statement, and statement control scope; The similarity between the attributes of the sentences is calculated, and the similarity is used as the similarity between different sentences.

6. The method for converting CUDA C language to Triton language according to claim 5, characterized in that: Obtaining the variable dependency relationship between the specified statements, and determining whether the output values ​​generated by the specified statements when executed are the same according to the variable dependency relationship, specifically includes: Determine whether a variable included in the specified statement has a corresponding variable definition and variable reference in the code snippet in which the variable is located; If so, directly determine whether the output values ​​generated by the specified statement when executed are the same; If not, determine the variable dependency relationship corresponding to the variable according to the variable type corresponding to the variable, and add the code segment where the dependent variable having the variable dependency relationship with the variable is located to the specified statement to obtain a complete execution statement corresponding to the specified statement; By executing the complete execution statement, it is determined whether the corresponding output values ​​are the same.

7. The method for converting CUDA C language to Triton language according to claim 4, characterized in that: Based on the mapping mode, the mapping rules between CUDA C language and Triton language are loaded, specifically including: Based on the centralized mapping mode, identifying common keywords in the equivalent code segments, and generating keyword templates corresponding to the CUDA C code according to the common keywords; The mapping rules between the CUDA C language and the Triton language are loaded, and based on the mapping rules, the keyword templates are reused to achieve batch conversion of a plurality of the CUDA C codes.

8. The method for converting CUDA C language to Triton language according to claim 1, characterized in that: After converting the CUDA C code into the corresponding Triton code, the method further includes: Determining the number of thread indices of parallel threads in the Triton code; When the number of thread indexes is greater than a preset value, the memory data structure and parallel algorithm of the Triton code are optimized to improve the execution efficiency of the Triton code.

9. A CUDA C language to Triton language conversion device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a CUDA C language to Triton language conversion method as described in any one of claims 1-8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: A method for converting CUDA C language to Triton language as described in any one of claims 1 to 8.