Method and device for converting CUDA code into OpenCL code, equipment and storage medium

By automatically converting the conversion method, obtaining CUDA code and converting it into OpenCL code, the problem of a large amount of time and resource investment required to manually port CUDA programs to OpenCL is solved, and the efficiency and accuracy of code conversion is improved.

CN120234012APending Publication Date: 2025-07-01XIAN XINTONG SEMICON TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267991.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Migrating from CUDA to OpenCL requires a lot of time and resource investment, and is prone to errors, resulting in low code conversion efficiency and accuracy.

Method used

By obtaining the CUDA code to be converted and its programming pattern, matching the preset OpenCL code framework, extracting the target information in the CUDA code, and converting it into the corresponding OpenCL code, and finally adding the OpenCL code to the target framework.

Benefits of technology

The automated process of converting CUDA code into OpenCL code is realized, which improves the efficiency and accuracy of code conversion and reduces developers' dependence on CUDA and OpenCL technology stacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234012A_ABST
    Figure CN120234012A_ABST
Patent Text Reader

Abstract

The invention discloses a method, device and equipment for converting a CUDA code into an OpenCL code and a storage medium, relates to the technical field of code conversion, and can improve the efficiency and accuracy of converting the CUDA code into the OpenCL code. According to the specific scheme, the method comprises the steps of obtaining a to-be-converted CUDA code and obtaining a programming mode of the CUDA code; matching a target code framework from a plurality of preset OpenCL code frameworks according to the programming mode; extracting target information in the CUDA code, wherein the target information comprises a kernel calling function, a kernel function, thread information, thread block configuration information, memory management information and API calling information; target information in the CUDA code is converted into a corresponding OpenCL code; and adding the OpenCL code into the target code framework to obtain a target OpenCL code converted by the CUDA code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of code conversion, and particularly to a method, device, equipment and storage medium for converting CUDA code into OpenCL code. Background Art

[0002] With the rapid growth of the demand for parallel computing, the Compute Unified Device Architecture (CUDA) platform can significantly improve computing performance by leveraging the processing power of the Graphics Processing Unit (GPU), and thus has a wide range of applications. However, the CUDA platform is limited to NVIDIA's GPU devices, which restricts the breadth of its applications. To address this issue, OpenCL has been proposed as a unified hardware abstraction, programming interface, and kernel development language, aiming to eliminate the significant investment by vendors in creating and maintaining proprietary software ecosystems.

[0003] OpenCL provides the ability for rapid deployment of computing devices, enabling them to be supported by chips from more vendors. However, migrating from CUDA to OpenCL is not straightforward. Currently, manually porting CUDA programs to OpenCL requires a large amount of time and resource investment, demands that developers have in-depth knowledge of the CUDA and OpenCL technology stacks, and this porting process is very time-consuming and error-prone. Summary of the Invention

[0004] This application provides a method, device, equipment and storage medium for converting CUDA code into OpenCL code, which can improve the efficiency and accuracy of converting CUDA code into OpenCL code.

[0005] To achieve the above object, this application adopts the following technical solutions:

[0006] In the first aspect of the embodiments of this application, a method for converting CUDA code into OpenCL code is provided, and the method includes:

[0007] Obtain the CUDA code to be converted and obtain the programming mode of the CUDA code;

[0008] Match a target code framework from a preset plurality of OpenCL code frameworks according to the programming mode;

[0009] Extract target information from the CUDA code, where the target information includes: kernel call function, kernel function, thread information, thread block configuration information, memory management information, and API call information;

[0010] Convert the target information in the CUDA code into corresponding OpenCL code;

[0011] Add the OpenCL code to the target code framework to obtain the target OpenCL code after CUDA code conversion.

[0012] As a possible implementation, after obtaining the programming mode of the CUDA code, the method further includes:

[0013] Obtain the thread memory hierarchy of the CUDA code;

[0014] Determine whether the CUDA code can be converted to OpenCL code according to the programming mode and the thread memory hierarchy;

[0015] The matching of the target code framework from a plurality of preset OpenCL code frameworks according to the programming mode includes:

[0016] If it is determined that the CUDA code can be converted to OpenCL code, then match the target code framework from a plurality of preset OpenCL code frameworks according to the programming mode.

[0017] As a possible implementation, if the target information is API call information, the conversion of the target information in the CUDA code into the corresponding OpenCL code includes:

[0018] Call a preset first conversion function to convert the API call information into the API call information corresponding to the OpenCL code;

[0019] If the target information is memory management information, the conversion of the target information in the CUDA code into the corresponding OpenCL code includes:

[0020] Call a preset second conversion function to convert the memory management information in the CUDA code into the memory management information corresponding to the OpenCL code.

[0021] As a possible implementation, if the target information is thread information, the conversion of the target information in the CUDA code into the corresponding OpenCL code includes:

[0022] Call a third conversion function to convert the thread information in the CUDA code into work items in the OpenCL code;

[0023] If the target information is thread block configuration information, the conversion of the target information in the CUDA code into the corresponding OpenCL code includes:

[0024] Call the fourth conversion function to convert the thread block configuration information in the CUDA code into the work group configuration in the OpenCL code.

[0025] As a possible implementation, if the target information is a kernel call function, the conversion of the target information in the CUDA code into the corresponding OpenCL code includes:

[0026] Obtain the function name of the kernel call function and modify the function name to the name corresponding to the OpenCL code format;

[0027] Traverse the kernel parameters in the kernel call function and convert the kernel parameters into the kernel parameters corresponding to the OpenCL code format;

[0028] Traverse the thread block information in the kernel call function and convert the thread block information into the work group information in the OpenCL code, where the thread block information includes thread block parameters and grid size;

[0029] Generate a kernel call function of the OpenCL code according to the function name, the kernel parameters, and the thread block information;

[0030] Use the kernel call function of the OpenCL code to replace the kernel call function in the CUDA code.

[0031] As a possible implementation, when the target information is a kernel function, the conversion of the target information in the CUDA code into the corresponding OpenCL code includes:

[0032] Detect the function attribute of the kernel function. If the function attribute indicates that the kernel function is the main function, convert the main function and the sub-functions in the main function into the kernel function in the OpenCL code format.

[0033] As a possible implementation, after obtaining the target OpenCL code converted from the CUDA code, the method further includes:

[0034] Perform unit testing, integration testing, and performance testing on the target OpenCL code, and cache the target OpenCL code after the target OpenCL code passes the testing.

[0035] In the second aspect of the embodiments of the present application, a device for converting CUDA code into OpenCL code is provided. The device includes:

[0036] An acquisition module, configured to acquire the CUDA code to be converted and acquire the programming mode of the CUDA code;

[0037] A matching module, configured to match a target code framework from a plurality of preset OpenCL code frameworks according to the programming mode;

[0038] An extraction module, configured to extract target information from the CUDA code, where the target information includes: kernel call functions, kernel functions, thread information, thread block configuration information, memory management information, and API call information;

[0039] A conversion module, configured to convert the target information in the CUDA code into corresponding OpenCL code;

[0040] A processing module, configured to add the OpenCL code to the target code framework to obtain the target OpenCL code converted from the CUDA code.

[0041] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the method for converting CUDA code into OpenCL code in the first aspect of the embodiments of the present application is implemented.

[0042] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for converting CUDA code into OpenCL code in the first aspect of the embodiments of the present application is implemented.

[0043] The beneficial effects brought by the technical solutions provided by the embodiments of the present application at least include:

[0044] The method for converting CUDA code into OpenCL code provided by the embodiments of the present application includes obtaining the CUDA code to be converted and the programming mode of the CUDA code, matching a target code framework from a plurality of preset OpenCL code frameworks according to the programming mode, then extracting the target information in the CUDA code, where the target information includes: kernel call functions, kernel functions, thread information, thread block configuration information, memory management information, and API call information, then converting the target information in the CUDA code into corresponding OpenCL code, and finally adding the OpenCL code to the target code framework to obtain the target OpenCL code converted from the CUDA code. Through the automated conversion method, the present application realizes accurate conversion and reliable underlying computing, thereby helping developers convert the original CUDA code into OpenCL code suitable for wider hardware support, which solves the large amount of time and resource investment required for manually porting CUDA programs to OpenCL and improves the efficiency and accuracy of code conversion. Description of the Drawings

[0045] Figure 1 A flowchart of a method for converting CUDA code into OpenCL code provided by an embodiment of the present application;

[0046] Figure 2 A structural diagram of a device for converting CUDA code into OpenCL code provided by an embodiment of the present application;

[0047] Figure 3 A structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0048] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.

[0049] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "a plurality" is two or more.

[0050] In addition, the use of "based on" or "according to" means open and inclusive, because a process, step, calculation, or other action based on or according to one or more conditions or values may, in practice, be based on additional conditions or values beyond those.

[0051] With the rapid growth of the demand for parallel computing, the Compute Unified Device Architecture (CUDA) platform can significantly improve computing performance by leveraging the processing power of the Graphics Processing Unit (GPU), and thus has a wide range of applications. However, the CUDA platform is limited to NVIDIA's GPU devices, restricting the extensiveness of its applications. To address this issue, OpenCL has been proposed as a unified hardware abstraction, programming interface, and kernel development language, aiming to eliminate the significant investment by vendors in creating and maintaining proprietary software ecosystems.

[0052] As two major GPU programming models, CUDA and OpenCL have extensive applications in fields such as scientific computing, deep learning, and big data processing. With the increasing demand for cross-platform and hardware diversification, the need to port existing CUDA applications to OpenCL-supported platforms or enhance the cross-platform capabilities and flexibility of applications is also growing. The maturity and accuracy of automated conversion technologies are key factors for market adoption. If the technology can ensure that the converted OpenCL code is consistent with the original CUDA code in terms of functionality and performance, and can efficiently convert complex CUDA code structures, it will greatly enhance its market attractiveness. In a highly competitive technology market, having efficient, accurate, and reliable automated conversion technologies will have strong market competitiveness. The market prospect of the CUDA-to-OpenCL automated conversion technology solution is broad, especially in scenarios that require cross-platform porting and optimization of existing GPU applications. The market size expands with the growth of GPU computing applications, covering a wide range of industries and application fields, providing sufficient market opportunities and potential for the technology.

[0053] However, migrating from CUDA to OpenCL is not straightforward. Currently, manually porting CUDA programs to OpenCL requires a large amount of time and resource investment and is error-prone, thus being inefficient. The porting process requires developers to have in-depth knowledge of the CUDA and OpenCL technology stacks, which places high demands on developers. Manual porting is prone to errors, which may cause problems when the converted code runs on different hardware. The porting process is time-consuming and cumbersome, and is not conducive to quickly migrating existing CUDA code to the OpenCL platform. Due to the complexity of manual porting, the accuracy of code conversion may not be high, affecting the reliability of the final calculation results.

[0054] Based on the above problems, the embodiments of this application provide a method for converting CUDA code to OpenCL code. This method realizes accurate conversion and reliable underlying calculations through an automated conversion method, thereby helping developers convert the original CUDA code into OpenCL code suitable for a wider range of hardware support. This solves the large amount of time and resource investment required for manually porting CUDA programs to OpenCL and improves the efficiency and accuracy of code conversion.

[0055] The method for converting CUDA code to OpenCL code provided by the embodiments of this application converts the host-side and device-side code, generates separate output files suitable for standard C++ compilers, and output files suitable for OpenCL kernel compilers. The supported conversions mainly focus on core runtime API calls for memory, device, thread, stream, and event management, as well as CUDA-style kernel calls and complete kernel function conversions.

[0056] It should be noted that each process or step in the method for converting CUDA code to OpenCL code provided by the embodiments of the present application is executed by calling the corresponding processing function, script or tool, and the whole process of executing the method for converting CUDA code to OpenCL code in the present application is an automated process.

[0057] The embodiments of the present application provide a method for converting CUDA code to OpenCL code, as Figure 1 shown, the method includes the following steps:

[0058] Step 101, obtain the CUDA code to be converted and obtain the programming mode of the CUDA code.

[0059] Optionally, after obtaining the CUDA code to be converted, verify the CUDA code.

[0060] In the actual execution process, after obtaining the CUDA code to be converted, it is necessary to confirm that the input program is indeed a cuda program and that it has no obvious syntax errors. This step can use the basic commands of the Clang compiler tool to verify the errors or architecture support versions of the cuda program itself. At the same time, relevant scripts can be used to complete the inspection of the basic cuda code before this.

[0061] In addition, when the relevant script compiles and performs a syntax check on the CUDA code, during the compilation process, the Clang compiler will automatically perform a syntax check and output any detected errors or warnings. Through this error or warning information, the developer needs to modify it manually.

[0062] Optionally, after the CUDA code passes the verification, obtain the code identifier of the CUDA code, and check whether the code identifier is included in the pre-stored historical code conversion data according to the code identifier. If it is included, it indicates that the CUDA code has been converted to OpenCL code before, and then directly obtain the OpenCL code corresponding to the CUDA code.

[0063] Step 102, match the target code framework from a plurality of preset OpenCL code frameworks according to the programming mode.

[0064] During the actual execution process, matching patterns of relevant and commonly used OpenCL codes can be prepared in advance. According to the programming mode, the tool will automatically select a matching template that it deems to be the best, use this pattern for matching, and create an OpenCL code file that conforms to this pattern. For example, only create a.cpp file and write the.cl kernel code as a macro string, or generate a.cpp and a.cl, etc. patterns. At the same time, the main CUDA code parameters, such as workgroups and context command queues, etc., will be identified in advance.

[0065] Step 103: Extract the target information in the CUDA code. The target information includes: kernel call function, kernel function, thread information, thread block configuration information, memory management information, and API call information.

[0066] It should be noted that the process of step 103 mainly extracts the key parts in the CUDA code, such as information on kernel functions, thread block configurations, memory management, API calls, etc., to prepare for subsequent conversions. This step is the preparation stage of the data structure in the code conversion tool. At the beginning of the complete process of this tool, that is, the cuda code to be converted needs to be input first. After the tool receives this file input, it will prepare the code content and data it needs by itself, and complete the preparation such as the storage and filling of the data structure.

[0067] For the implementation of the process of reading the CUDA code for the entire code conversion, it is through a parsing function. The principle of this function is triggered by the AST traversal mechanism of Clang. Whenever a global scope declaration is encountered, whether it is a function, variable, class, or other declarations, it will be called. This function is responsible for identifying the source file where each declaration is located, and when processing files in.cu and.cuh formats, generate the host-side and kernel-side include files of *-cl.h and *-cl.cl as needed. This function first identifies the type of the declaration, and finally hands over the control to the corresponding host-side and kernel-side rewriters for processing.

[0068] Step 104: Convert the target information in the CUDA code into the corresponding OpenCL code.

[0069] Step 105: Add the OpenCL code to the target code framework to obtain the target OpenCL code converted from the CUDA code.

[0070] The method for converting CUDA code into OpenCL code provided by the embodiments of this application includes obtaining the CUDA code to be converted and the programming mode of the CUDA code, matching a target code framework from multiple preset OpenCL code frameworks according to the programming mode, then extracting target information from the CUDA code, where the target information includes: kernel call functions, kernel functions, thread information, thread block configuration information, memory management information, and API call information. Then, convert the target information in the CUDA code into corresponding OpenCL code, and finally add the OpenCL code to the target code framework to obtain the target OpenCL code converted from the CUDA code. Through the automated conversion method, this application achieves accurate conversion and reliable underlying computing, thereby helping developers convert the original CUDA code into OpenCL code suitable for a wider range of hardware support, which solves the large amount of time and resource investment required for manually porting CUDA programs to OpenCL and improves the efficiency and accuracy of code conversion.

[0071] Optionally, after obtaining the programming mode of the CUDA code, the method further includes:

[0072] Obtaining the thread memory hierarchy of the CUDA code;

[0073] Judging whether the CUDA code can be converted into OpenCL code according to the programming mode and the thread memory hierarchy;

[0074] The step of matching a target code framework from multiple preset OpenCL code frameworks according to the programming mode includes:

[0075] If it is determined that the CUDA code can be converted into OpenCL code, then match a target code framework from multiple preset OpenCL code frameworks according to the programming mode.

[0076] That is to say, after step 101 of obtaining the programming mode of the CUDA code and before step 102 of matching a target code framework from multiple preset OpenCL code frameworks according to the programming mode, the method further includes: obtaining the thread memory hierarchy of the CUDA code; judging whether the CUDA code can be converted into OpenCL code according to the programming mode and the thread memory hierarchy.

[0077] In actual practice, before starting the conversion of CUDA code, it is necessary to first analyze the programming mode and thread memory hierarchy of the cuda code to be converted to preliminarily judge whether the conversion can be completed. If not, the tool will automatically terminate and give a prompt.

[0078] Optionally, if the target information is API call information, the conversion of the target information in the CUDA code into corresponding OpenCL code includes:

[0079] Call a preset first conversion function to convert the API call information into API call information corresponding to the OpenCL code.

[0080] During the actual execution process, convert the API calls in the CUDA code into corresponding API calls in the OpenCL code. Here, an example of creating a buffer and writing to the buffer is given. For example, cudaMalloc is converted to clCreateBuffer. First, it is judged whether the input function name is cudaMalloc. If so, enter the processing flow. The memory is created and applied for by the cuda program, and the memory size is passed in. The overall processing idea is mainly to judge the type of the processing function to determine whether it includes the API call respectively. If it does, enter the corresponding processing program logic respectively.

[0081] If the target information is memory management information, the conversion of the target information in the CUDA code into corresponding OpenCL code includes:

[0082] Call a preset second conversion function to convert the memory management information in the CUDA code into memory management information corresponding to the OpenCL code. In addition, it is also necessary to ensure the correctness of memory allocation and release.

[0083] Optionally, if the target information is thread information, the conversion of the target information in the CUDA code into corresponding OpenCL code includes:

[0084] Call a third conversion function to convert the thread information in the CUDA code into work items in the OpenCL code;

[0085] If the target information is thread block configuration information, the conversion of the target information in the CUDA code into corresponding OpenCL code includes:

[0086] Call a fourth conversion function to convert the thread block configuration information in the CUDA code into work group configuration in the OpenCL code.

[0087] During the actual execution process, the thread and thread block configurations in CUDA are converted into the work item and work group configurations in OpenCL, and threadIdx, blockIdx, etc. are converted into get_global_id, get_local_id, etc. Therefore, we implemented a function to complete this conversion. The general logic is to match the cl code written this time based on the results of the previous analysis, along with the incoming parameters it depends on. This parameter is obtained by analyzing and identifying the CUDA code extraction to complete the conversion.

[0088] Optionally, if the target information is a kernel call function, converting the target information in the CUDA code into the corresponding OpenCL code includes:

[0089] Obtain the function name of the kernel call function and modify the function name to the name corresponding to the OpenCL code format;

[0090] Traverse the kernel parameters in the kernel call function and convert the kernel parameters into the kernel parameters corresponding to the OpenCL code format;

[0091] Traverse the thread block information in the kernel call function and convert the thread block information into the work group information in the OpenCL code. The thread block information includes thread block parameters and grid size;

[0092] Generate a kernel call function for the OpenCL code based on the function name, the kernel parameters, and the thread block information;

[0093] Use the kernel call function of the OpenCL code to replace the kernel call function in the CUDA code.

[0094] During the actual execution process, extract the kernel call function name and modify its name to the OpenCL style, such as the cu2cl_Kernel_ prefix. Obtain the kernel launch configuration: Use kernelCall->getConfig() to obtain the CUDA launch configuration parameters, namely the grid and block information. Traverse the kernel parameters: Loop through the list of parameters passed to the CUDA kernel. Process parameter passing: If the parameter is a literal: Create a temporary variable for the literal parameter and set the parameter using clSetKernelArg. If the parameter is a variable: Directly pass the address of the variable. If it is a device memory variable, handle it as the cl_mem type. Set the workgroup size: Process the Block (thread block) size: By checking the block parameter, determine whether it is a dim3 variable or an integer type and convert it to the local work size of OpenCL. Process the Grid (grid) size: Similar to processing the grid parameter, convert it to the global work size of OpenCL and multiply it by the local work size to obtain the final workgroup size. Generate the kernel call code for the OpenCL code: Based on the previously parsed information, generate the corresponding clEnqueueNDRangeKernel call statement to launch the kernel in OpenCL. Return the generated OpenCL code: Return the generated OpenCL code as a string to replace the original CUDA kernel call.

[0095] Optionally, the target information is a kernel function, and converting the target information in the CUDA code into the corresponding OpenCL code includes:

[0096] Detect the function attributes of the kernel function. If the function attributes indicate that the kernel function is the main function, convert the main function and the sub-functions in the main function into kernel functions in the OpenCL code format.

[0097] During the actual execution process, CUDA kernels are usually directly embedded in the binary at compile time, while OpenCL kernels need to be dynamically compiled through the API. We implemented a RewriteKernelFunction function to automatically complete the conversion of cuda kernel programs to opencl c kernel code. Specifically, the process can be as follows:

[0098] Check kernel function attributes: The code first checks whether the function has the main function attribute (__global__) and has a function body (hasBody()). Only such functions will be processed because they need to be converted into OpenCL kernel functions. Obtain the kernel file name and store the kernel function name: Obtain the source file name containing the kernel function definition and store the name of the kernel function in a list associated with that file name (Kernels[r]). Manage global variable declarations: Check whether the declaration of the cl_kernel variable has been created for this kernel function. If not, add the declaration to the global declaration list to ensure that each kernel function has a corresponding declaration in OpenCL.

[0099] Convert the __global__ attribute to __kernel: Convert the __global__ attribute of CUDA to the __kernel attribute of OpenCL, which is the identifier of kernel functions in OpenCL. Remove the __device__ attribute, sub-functions: The __device__ attribute of CUDA has no direct corresponding attribute in OpenCL, so it is removed here. Remove the __host__ attribute, CPU: OpenCL does not have the __host__ attribute because these attributes are only valid on the host side and do not need to be retained during the conversion process.

[0100] Rewrite the parameters of the kernel function: Traverse all the parameters of the kernel function and convert each parameter into a form suitable for OpenCL. The RewriteKernelParam function is responsible for handling this process and decides how to convert according to whether the parameter belongs to a global function (CUDAGlobalAttr).

[0101] Rewrite the main code of the kernel function: If the kernel function has a function body (hasBody()), call RewriteKernelStmt to convert each statement in the function body to ensure that they conform to the syntax and semantics of OpenCL.

[0102] Clean up the reference parameter variables: After completing the rewriting of the kernel function, empty the list CurRefParmVars that stores the reference parameter variables to prepare for the processing of the next function.

[0103] After the above process, write the contents of the rewritten host-side and kernel-side files to the corresponding output files. If the file has been rewritten, write the new content to the output file; if the file has not been modified, directly copy the original content. The code also handles the case where the file ID is invalid and provides some error handling and logging functions.

[0104] Traverse all rewritten host-side files: Use the iterator i to traverse the OutFiles map, which stores the file information to be output.

[0105] Obtain file information: Obtain the file entry (FileEntry) through Files.getFile((*i).first). Use RewriteSM.translateFile(FE) to convert the file entry to a file ID (FileID). Check if the file ID is valid: If the file ID is invalid (possibly because the file has not been modified), attempt to force the creation of the file ID.

[0106] If the creation of the file ID is successful, continue processing; otherwise, skip the file and issue an error message. Process the content of the rewritten file: Check if there is a rewrite buffer (RewriteBuffer) for this file in GlobalHostRewrite. If there is, write the buffer content to the output file. If the buffer is not found, directly copy the original content of the file to the output file and issue a message indicating that the file has not been modified. Clean up the output file: Call the clearOutputFile(outFile, &Files) function to clean up the relevant resources of the output file.

[0107] Traverse all rewritten kernel files: In the same way as processing host-side files, traverse the KernelOutFiles map. Obtain file information: Obtain the file entry through Files.getFile((*i).first) and convert it to a file ID (FileID). Check if the file ID is valid: If the file ID is invalid, issue an error message and push the file to the reprocessing list for future use. Process the content of the rewritten file: Check if there is a rewrite buffer for this file in GlobalKernRewrite. If there is, write the buffer content to the output file. If the buffer is not found, issue a message indicating that the file has not been modified. Clean up the output file: Call the clearOutputFile(outFile, &Files) function to clean up the relevant resources of the output file.

[0108] Optionally, after obtaining the target OpenCL code for the CUDA code conversion, the method further includes:

[0109] Perform unit testing, integration testing, and performance testing on the target OpenCL code, and cache the target OpenCL code after it passes the tests. Optimize and debug the generated target OpenCL code to ensure that the code functions correctly and the performance meets the expectations.

[0110] Testing and Validation: Rigorously test and validate the converted target OpenCL code, including unit tests, integration tests, and performance tests, to ensure that its functionality is consistent with the original CUDA code. Through the above steps, the automated conversion of CUDA code to OpenCL code can be achieved, improving development efficiency and ensuring the correctness and performance of the code.

[0111] This application realizes the automated code conversion process from CUDA to OpenCL. By combining the Clang compiler infrastructure to parse the CUDA source code and generate an abstract syntax tree, and then generating an OpenCL code framework based on the extracted information, while handling CUDA API calls and memory management, finally optimizing, debugging, and validating the converted OpenCL code. This process not only saves a large amount of time for developers in manual conversion but also ensures that the converted code is consistent with the original CUDA code in terms of functionality and performance, thus improving code reusability and cross-platform portability efficiency. Through a series of systematic and automated steps, the seamless conversion of CUDA code to OpenCL code is achieved.

[0112] This application provides a systematic parsing and conversion process: It proposes a complete process from identifying and parsing CUDA source code to generating and optimizing OpenCL code, including syntax and semantic analysis, key element extraction, and identification of specific programming patterns.

[0113] Precise API Replacement: Aiming at the characteristics of CUDA API calls, a precise replacement strategy is proposed. By finding equivalent APIs in OpenCL and replacing the corresponding CUDA API calls, seamless conversion at the API level is achieved.

[0114] Memory Management Conversion: Identify and convert the CUDA memory management mechanism, using OpenCL memory objects and operations to replace CUDA memory management methods to ensure the correctness of memory allocation, release, and data transfer.

[0115] Thread and Block Configuration Conversion: A method is proposed to convert CUDA thread and block configurations to OpenCL workgroups and global offsets. By updating the kernel function to use OpenCL built-in functions, automated conversion of thread and block configurations is achieved.

[0116] Optimization and Debugging: Through optimization and debugging steps, check and improve the performance of the converted OpenCL code to ensure that its performance characteristics are similar to the original CUDA code, and at the same time, a method for debugging using OpenCL debugging tools is provided.

[0117] Testing and Verification: Systematic testing and verification steps are carried out. Test cases are written to verify the correctness of the OpenCL code, and the outputs and performances of the CUDA and OpenCL codes are compared to ensure that the converted code has consistent functionality and reliable performance.

[0118] This application is applicable to application programs with high performance requirements that need to run on multiple GPU architectures. For example, in the fields of scientific computing, deep learning training, large-scale data processing, etc., projects that require parallel computing using CUDA or OpenCL. It may be applicable to the situation where existing CUDA application programs need to be ported to a platform supporting OpenCL to enhance the cross-platform capabilities and flexibility of the application programs. For example: meteorological simulation, seismic simulation, etc., which need to run on different GPU architectures to utilize the performance advantages of different hardware. Deep learning training application programs: In different deep learning frameworks, it may be necessary to run training tasks on different GPU architectures to utilize the computing resources of different hardware.

[0119] The device for converting CUDA code into OpenCL code provided by the embodiments of this application, as Figure 2 shown, the device includes:

[0120] An acquisition module 11, configured to acquire the CUDA code to be converted and the programming mode of the CUDA code;

[0121] A matching module 12, configured to match a target code framework from a plurality of preset OpenCL code frameworks according to the programming mode;

[0122] An extraction module 13, configured to extract target information from the CUDA code, where the target information includes: kernel call functions, kernel functions, thread information, thread block configuration information, memory management information, and API call information;

[0123] A conversion module 14, configured to convert the target information in the CUDA code into corresponding OpenCL code;

[0124] A processing module 15, configured to add the OpenCL code to the target code framework to obtain the target OpenCL code converted from the CUDA code.

[0125] In one embodiment, the acquisition module 11 is further configured to acquire the thread memory hierarchy of the CUDA code; and determine whether the CUDA code can be converted into OpenCL code according to the programming mode and the thread memory hierarchy;

[0126] The matching module 12 is further configured to, if it is determined that the CUDA code can be converted into OpenCL code, match a target code framework from a plurality of preset OpenCL code frameworks according to the programming mode.

[0127] In one embodiment, if the target information is API call information, the conversion module 14 is specifically configured to:

[0128] Call a preset first conversion function to convert the API call information into API call information corresponding to OpenCL code;

[0129] If the target information is memory management information, the conversion module 14 is specifically configured to:

[0130] Call a preset second conversion function to convert the memory management information in the CUDA code into memory management information corresponding to the OpenCL code.

[0131] In one embodiment, if the target information is thread information, the conversion module 14 is specifically configured to:

[0132] Call a third conversion function to convert the thread information in the CUDA code into work items in OpenCL code;

[0133] If the target information is thread block configuration information, the conversion module 14 is specifically configured to:

[0134] Call a fourth conversion function to convert the thread block configuration information in the CUDA code into work group configuration in OpenCL code.

[0135] In one embodiment, if the target information is a kernel call function, the conversion module 14 is specifically configured to:

[0136] Obtain the function name of the kernel call function, and modify the function name to a name corresponding to the OpenCL code format;

[0137] Traverse the kernel parameters in the kernel call function, and convert the kernel parameters into kernel parameters corresponding to the OpenCL code format;

[0138] Traverse the thread block information in the kernel call function, and convert the thread block information into work group information in OpenCL code, where the thread block information includes thread block parameters and grid size;

[0139] Generate a kernel call function of OpenCL code according to the function name, the kernel parameters, and the thread block information;

[0140] Replace the kernel call function in the CUDA code with the kernel call function of the OpenCL code.

[0141] In one embodiment, the target information is a kernel function, and the conversion module 14 is specifically configured to:[[]]

[0142] Detect the function attributes of the kernel function. If the function attributes indicate that the kernel function is the main function, convert the main function and the sub-functions in the main function into kernel functions in OpenCL code format.

[0143] In one embodiment, the processing module 15 is further configured to:[[]]

[0144] Perform unit testing, integration testing, and performance testing on the target OpenCL code, and cache the target OpenCL code after the target OpenCL code passes the testing.

[0145] The apparatus for converting CUDA code to OpenCL code provided in this embodiment can execute the method embodiment of converting CUDA code to OpenCL code, and its implementation principle and technical effects are similar, so details are not described herein again.

[0146] For the specific limitations on the apparatus for converting CUDA code to OpenCL code, reference can be made to the limitations on the method of converting CUDA code to OpenCL code in the foregoing text, which are not described herein again. Each module in the above-mentioned apparatus for converting CUDA code to OpenCL code can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the electronic device in hardware form or be independent of it, or be stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0147] The execution subject of the method for converting CUDA code to OpenCL code provided in the embodiments of the present application can be an electronic device, which can be a computer device, a terminal device, a server, or a server cluster. The embodiments of the present application do not make specific limitations on this.

[0148] Figure 3 It is a schematic internal structure diagram of an electronic device provided in the embodiments of the present application. As Figure 3As shown in the figure, the electronic device includes a processor and a memory connected through a system bus. Among them, the processor is used to provide computing and control capabilities. The memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The computer program can be executed by the processor to implement the steps of the method for converting CUDA code into OpenCL code provided in each of the above embodiments. The internal memory provides a cache operating environment for the operating system and the computer program in the non-volatile storage medium.

[0149] Those skilled in the art can understand that Figure 3 the internal structure diagram of the electronic device shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0150] In another embodiment of this application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for converting CUDA code into OpenCL code as in the embodiments of this application are implemented.

[0151] In another embodiment of this application, a computer program product is further provided. The computer program product includes computer instructions, and when the computer instructions execute each step of the method for converting CUDA code into OpenCL code in the method flow shown in the above method embodiments.

[0152] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that contains one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0153] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0154] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for converting CUDA code into OpenCL code, characterized in that: The method comprises: Obtaining a CUDA code to be converted and a programming mode for obtaining the CUDA code; Matching a target code frame from a plurality of preset OpenCL code frames according to the programming mode; Extracting target information in the CUDA code, the target information including: kernel call function, kernel function, thread information, thread block configuration information, memory management information and API call information; Convert the target information in the CUDA code into the corresponding OpenCL code; The OpenCL code is added to the target code framework to obtain the target OpenCL code converted from the CUDA code.

2. The method according to claim 1, characterized in that After acquiring the programming mode of the CUDA code, the method further includes: Obtaining the thread memory hierarchy of the CUDA code; Determining whether the CUDA code can be converted into OpenCL code according to the programming mode and the thread memory hierarchy; The matching of a target code framework from a plurality of preset OpenCL code frameworks according to the programming mode includes: If it is determined that the CUDA code can be converted into an OpenCL code, a target code framework is matched from a plurality of preset OpenCL code frameworks according to the programming mode.

3. The method according to claim 1, characterized in that If the target information is API call information, converting the target information in the CUDA code into a corresponding OpenCL code includes: Calling a preset first conversion function to convert the API call information into API call information corresponding to the OpenCL code; If the target information is memory management information, converting the target information in the CUDA code into a corresponding OpenCL code includes: A preset second conversion function is called to convert the memory management information in the CUDA code into memory management information corresponding to the corresponding OpenCL code.

4. The method according to claim 1, characterized in that: If the target information is thread information, converting the target information in the CUDA code into a corresponding OpenCL code includes: Calling a third conversion function to convert the thread information in the CUDA code into a work item in the OpenCL code; If the target information is thread block configuration information, converting the target information in the CUDA code into a corresponding OpenCL code includes: The fourth conversion function is called to convert the thread block configuration information in the CUDA code into a work group configuration in the OpenCL code.

5. The method according to claim 1, characterized in that If the target information is a kernel call function, converting the target information in the CUDA code into a corresponding OpenCL code includes: Obtaining a function name of the kernel call function, and modifying the function name to a name corresponding to an OpenCL code format; Traversing the kernel parameters in the kernel calling function, and converting the kernel parameters into kernel parameters corresponding to the OpenCL code format; Traversing the thread block information in the kernel call function, and converting the thread block information into work group information in the OpenCL code, wherein the thread block information includes thread block parameters and a grid size; Generate a kernel call function of the OpenCL code according to the function name, the kernel parameters and the thread block information; The kernel calling function in the CUDA code is replaced by the kernel calling function in the OpenCL code.

6. The method according to claim 1, characterized in that The target information is a kernel function, and converting the target information in the CUDA code into a corresponding OpenCL code includes: The function attribute of the kernel function is detected, and if the function attribute indicates that the kernel function is a main function, the main function and the sub-functions in the main function are converted into the kernel function in the OpenCL code format.

7. The method according to claim 1, characterized in that After obtaining the target OpenCL code converted from the CUDA code, the method further includes: Unit testing, integration testing and performance testing are performed on the target OpenCL code, and after the target OpenCL code passes the test, the target OpenCL code is cached.

8. A device for converting CUDA code into OpenCL code, characterized in that: The device comprises: An acquisition module, used for acquiring the CUDA code to be converted and acquiring the programming mode of the CUDA code; A matching module, used for matching a target code frame from a plurality of preset OpenCL code frames according to the programming mode; An extraction module, used to extract target information in the CUDA code, wherein the target information includes: kernel call function, kernel function, thread information, thread block configuration information, memory management information and API call information; A conversion module, used for converting the target information in the CUDA code into a corresponding OpenCL code; The processing module is used to add the OpenCL code to the target code framework to obtain the target OpenCL code converted from the CUDA code.

9. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for converting a CUDA code into an OpenCL code according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method for converting a CUDA code into an OpenCL code according to any one of claims 1 to 7 is implemented.