Code Compilation Method, Electronic Device, and Storage Medium

By allocating the memory address of thread-local storage variables to the GPU architecture during compilation, the development complexity of thread-local storage under the GPU architecture is solved, and efficient memory management and performance optimization are achieved.

CN120122957BActive Publication Date: 2025-07-08SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510534557.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-08
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The lack of support for existing thread-local storage implementations in the GPU architecture has led to developers needing to manually manage thread data, which increases development complexity and performance losses. The LLVM compiler lacks support for GPU thread-local storage and cannot automatically generate efficient memory allocation and access code.

Method used

Provides a code compilation method, which can be used to store variables locally in the target thread, allocate them to thread-local memory, and fix the memory address during compilation, generate optimized intermediate representations and machine code, simplify the programming model and reduce the memory address acquisition at runtime.

Benefits of technology

It realizes efficient thread-local storage management under the GPU architecture, simplifies the development process, improves development efficiency, optimizes performance, and reduces register and runtime overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122957B_ABST
    Figure CN120122957B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a code compilation method, an electronic device, and a storage medium. The code compilation method is used to compile code including at least one thread-local storage variable, and the method includes: in response to the existence of a target thread-local storage variable, based on the attributes of at least one target thread-local storage variable, allocating at least one target thread-local storage variable to thread-local memory to obtain a thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is a thread-local storage variable used by a first function; updating the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; generating machine code corresponding to the code based on the updated intermediate representation. The code compilation method provides support for a thread-local storage model in a parallel computing architecture, and solves the problem that traditional thread-local storage implementations cannot be directly applied to the GPU architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a code compilation method, an electronic device, and a storage medium. Background Art

[0002] The Low-Level Virtual Machine (LLVM) is an open-source compiler framework, which is designed as a collection of modular and reusable compiler and toolchain technologies. LLVM supports the Thread Local Storage (TLS) model and can define thread-local storage variables, enabling each thread to have an independent copy of the variable, thereby avoiding data competition problems between threads. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a code compilation method, where the code includes at least one thread-local storage variable, and the code compilation method includes: in response to the existence of a target thread-local storage variable, based on the attributes of at least one target thread-local storage variable, allocating the at least one target thread-local storage variable to thread-local memory to obtain a thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is a thread-local storage variable used by a first function among the at least one thread-local storage variable; updating an intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; and generating machine code corresponding to the code based on the updated intermediate representation.

[0004] The code compilation method provided by at least one embodiment of the present disclosure further includes: inserting an instruction for initializing the target thread-local storage variable used by each first function into the intermediate representation corresponding to the code.

[0005] The code compilation method provided by at least one embodiment of the present disclosure further includes: for each target thread-local storage variable, determining a first function that uses the target thread-local storage variable to construct a variable-function mapping relationship.

[0006] In the code compilation method provided by at least one embodiment of the present disclosure, the first function directly uses the target thread-local storage variable or uses the target thread-local storage variable by calling a second function.

[0007] In the code compilation method provided by at least one embodiment of the present disclosure, the first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0008] In the code compilation method provided by at least one embodiment of the present disclosure, allocating the at least one target thread-local storage variable to the thread-local memory based on the attributes of the at least one target thread-local storage variable includes: determining a variable allocation order based on the attributes of the at least one target thread-local storage variable; and sequentially allocating the at least one target thread-local storage variable to the thread-local memory according to the variable allocation order.

[0009] In the code compilation method provided by at least one embodiment of the present disclosure, the attributes of the target thread-local storage variable include at least one of the number of first functions corresponding to the target thread-local storage variable, the alignment parameter of the target thread-local storage variable, the size of the target thread-local storage variable, and the definition order of the target thread-local storage variable.

[0010] In the code compilation method provided by at least one embodiment of the present disclosure, determining the variable allocation order based on the attributes of the at least one target thread-local storage variable includes: arranging the at least one target thread-local storage variable in descending order according to the number of first functions corresponding to the at least one target thread-local storage variable; arranging the at least one target thread-local storage variable in descending order according to the alignment parameter of the target thread-local storage variable; arranging the at least one target thread-local storage variable in ascending order according to the size of the target thread-local storage variable; or arranging the at least one target thread-local storage variable according to the definition order of the target thread-local storage variable.

[0011] In the code compilation method provided by at least one embodiment of the present disclosure, determining the variable allocation order based on the attributes of the at least one target thread-local storage variable includes: arranging the at least one target thread-local storage variable in descending order according to the number of first functions corresponding to the at least one target thread-local storage variable; in response to a plurality of first thread-local storage variables having the same number of first functions among the at least one target thread-local storage variable, arranging the plurality of first thread-local storage variables in descending order according to the alignment parameter of the target thread-local storage variable; in response to a plurality of second thread-local storage variables having the same alignment parameter among the plurality of first thread-local storage variables, arranging the plurality of second thread-local storage variables in ascending order according to the size of the target thread-local storage variable; and in response to a plurality of third thread-local storage variables having the same size among the plurality of second thread-local storage variables, arranging the plurality of third thread-local storage variables according to the definition order of the target thread-local storage variable.

[0012] In the code compilation method provided by at least one embodiment of the present disclosure, for the at least one target thread-local storage variable, based on the attributes of the at least one target thread-local storage variable, allocating the at least one target thread-local storage variable to thread-local memory further includes: for the target thread-local storage variables corresponding to multiple first functions, using the maximum value among the allocated offsets corresponding to the multiple first functions as the allocation starting point of the target thread-local storage variable in the thread-local memory.

[0013] In the code compilation method provided by at least one embodiment of the present disclosure, for updating the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable, it includes: for each target thread-local storage variable, replacing the instruction for obtaining the address of the target thread-local storage variable with the thread-local memory address corresponding to the target thread-local storage variable, where the intermediate representation corresponding to the code includes the instruction for obtaining the address of the target thread-local storage variable.

[0014] The code compilation method provided by at least one embodiment of the present disclosure further includes: obtaining a record of the thread-local memory size required for all target thread-local storage variables used by each first function in the code, so as to pre-allocate a fixed thread-local memory space for each first function according to the record.

[0015] In the code compilation method provided by at least one embodiment of the present disclosure, the life cycle of the pre-allocated thread-local memory space is consistent with the execution cycle of the corresponding first function.

[0016] The present disclosure provides a code compilation apparatus at least in one embodiment. The code includes at least one thread-local storage variable. The code compilation apparatus includes: an allocation module configured to, in response to the existence of a target thread-local storage variable, based on the attributes of at least one target thread-local storage variable, allocate the at least one target thread-local storage variable to thread-local memory to obtain the thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is the thread-local storage variable used by the first function among the at least one thread-local storage variable; an update module configured to update the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; and a generation module configured to generate the machine code corresponding to the code based on the updated intermediate representation.

[0017] One embodiment of the present disclosure provides an electronic device, which includes: at least one processor; at least one memory, including one or more computer program modules; wherein, the one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the code compilation method provided by at least one embodiment of the present disclosure.

[0018] One embodiment of the present disclosure provides a non-transitory computer-readable storage medium, on which computer instructions are stored, wherein the computer instructions, when executed by at least one processor, execute the code compilation method provided by at least one embodiment of the present disclosure.

[0019] The code compilation method, code compilation device, electronic device and non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure provide support for the thread local storage model in a parallel computing architecture, and solve the problem that the traditional thread local storage implementation cannot be directly applied to the GPU architecture. Developers can directly use high-level syntax for programming without manually managing thread data. The programming model is simple and easy to operate, greatly improving the development efficiency. Moreover, through compilation optimization, the corresponding thread local memory address can be fixed directly during compilation, without loading the thread local memory address corresponding to each thread local storage variable from the register, nor dynamically obtaining the thread local memory address at runtime, achieving better performance in obtaining the thread local memory address. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0021] Figure 1 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0022] Figure 2 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0023] Figure 3 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0024] Figure 4A It is an exemplary diagram of a code compilation method provided by at least one embodiment of the present disclosure;

[0025] Figure 4B It is an exemplary diagram of a code compilation method provided by at least one embodiment of the present disclosure;

[0026] Figure 5 Flow chart of a code compilation method provided by at least one embodiment of the present disclosure;

[0027] Figure 6 Schematic block diagram of a code compilation apparatus provided by at least one embodiment of the present disclosure;

[0028] Figure 7 Schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure;

[0029] Figure 8 Schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure; and

[0030] Figure 9 Schematic block diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0031] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0032] Flowcharts are used in the present disclosure to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations described above or below do not necessarily need to be performed precisely in order. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0033] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second" and similar terms used in the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Terms such as "include" or "comprise" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. Terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left", "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0034] The following provides an illustration of the present disclosure through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of an embodiment of the present disclosure appears in more than one drawing, the component is denoted by the same or similar reference numeral in each drawing.

[0035] LLVM is an open-source compiler framework, which is designed as a collection of modular and reusable compiler and toolchain technologies. LLVM includes a front end, a middle end, and a back end.

[0036] The front end of LLVM is responsible for converting the source code of different programming languages or programming models into a general intermediate representation (IR). For example, Clang is a compiler front end natively supported by LLVM for C, C++, and Objective-C. When a user declares a variable using thread_local (C++11) or __thread (GCC extension), Clang can mark the variable as a thread-local storage variable (TLS variable), for example, it can be marked through the thread_local attribute and a specific linkage type (such as dso_local) in the generated intermediate representation.

[0037] The middle end of LLVM is responsible for optimizing the intermediate representation generated by the front end to improve the performance of the finally generated code. For example, it can simplify the thread-local storage access pattern and merge multiple accesses into one address calculation. For example, it can also implicitly select the thread-local storage model according to the scope of the variable or explicitly specify the thread-local storage model in the code.

[0038] The back end of LLVM is responsible for converting the optimized intermediate representation into machine code or assembly code according to the characteristics of the target hardware platform. The specific tasks of the back end can include aspects such as instruction selection, register allocation, instruction scheduling, and code generation. For example, instruction selection refers to mapping the intermediate representation to the instruction set of the target hardware platform; register allocation refers to mapping the virtual registers in the intermediate representation to the actual registers of the target hardware; instruction scheduling refers to rearranging the instruction order to maximize the utilization of hardware resources; code generation refers to generating the machine code or assembly code corresponding to the target hardware platform.

[0039] LLVM supports the thread-local storage model. Through the thread_local keyword, thread-local storage variables can be defined, enabling each thread to have an independent copy of the variable, thereby avoiding data competition problems between threads. That is to say, thread-local storage variables are not shared between threads, which means that even if multiple threads access the same variable name, they are actually operating on their respective independent memory areas.

[0040] However, the inventors of the present disclosure have noticed that the implementation of thread - local storage depends on the support of the target platform. Currently, in traditional Central Processing Unit (CPU) architectures (such as x86 and ARM), the operating system and the Application Binary Interface (ABI) provide support for thread - local storage. For example, it can support the implementation of thread - local storage through specific segments in the Executable and Linkable Format (ELF) file (such as the.tdata segment for storing uninitialized thread - local storage variables and the.tbss segment for storing initialized thread - local storage variables). However, in Graphics Processing Unit (GPU), General - Purpose Graphics Processing Unit (GPGPU), or other back - ends related to parallel computing architectures (such as CUDA and OpenCL), there is no similar mechanism to support the implementation of thread - local storage, which poses challenges for application scenarios that require large - scale parallel processing, such as high - performance computing and machine learning.

[0041] The parallel thread models of GPUs (such as CUDA and OpenCL) usually require each thread to independently access its private data, but the traditional implementation methods of thread - local storage (such as the __thread keyword in CPUs) cannot be directly mapped to the GPU architecture, making it difficult to meet the requirements of creating thread - private data and supporting related data management and optimization. Moreover, GPUs lack hardware - level support for thread - local storage, resulting in developers having to manually manage thread - local data (such as calculating offsets through thread indices), increasing the development complexity and being error - prone. In existing GPU programming practices, developers need to explicitly pass thread indices (such as threadIdx.x in CUDA) and manually allocate global memory to simulate the storage of thread - private data, which leads to code redundancy and performance loss. In addition, the existing LLVM compiler has insufficient support for thread - local storage in GPUs and cannot automatically generate efficient memory allocation and access code.

[0042] The inventors of the present disclosure have also noticed that the thread - local storage models supported by current mainstream CPU architectures have performance limitations. For example, the static model has instruction and register overhead when accessing thread - local storage variables, and the dynamic model has runtime resolution overhead when accessing thread - local storage variables. The specific analysis is as follows:

[0043] (1) Static model

[0044] In the CPU architecture, the static model can access thread-local storage variables by using segment registers (such as the FS segment register in Linux and the GS segment register in Windows) plus an offset. An example can be referred to the following assembly code:

[0045] movq %fs:0, %rax ; / / LINE 1

[0046] addq $TLS_var@tpoff, %rax ; / / LINE 2

[0047] Line LINE 1 means loading the TLS base address to the RAX register from the address pointed to by the FS segment register. Line LINE 2 means adding the offset (@tpoff) calculated at compile time to the base address. The actual memory address of the TLS_var variable of the current thread is stored in the RAX register. At static linking, the linker determines the final offset of the TLS_var variable (e.g., 0x10) and replaces @tpoff with the actual value of the offset. The above method requires executing movq to load the base address every time accessing a thread-local storage variable, and needs to occupy the RAX register to temporarily store the base address, increasing the register pressure. Therefore, there are relatively large instruction overhead and register overhead.

[0048] (2) Dynamic model

[0049] In the CPU architecture, the dynamic model obtains the address of thread-local storage variables by calling __tls_get_addr() at runtime or dynamically calculating the offset, which may bring function call overhead or memory access overhead and has low performance.

[0050] At least one embodiment of the present disclosure provides a code compilation method for compiling code including at least one thread-local storage variable. The code compilation method includes: in response to the existence of target thread-local storage variables, based on the attributes of at least one target thread-local storage variable, allocating at least one target thread-local storage variable to thread-local memory to obtain the thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is the thread-local storage variable used by the first function among at least one thread-local storage variable; updating the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; generating the machine code corresponding to the code based on the updated intermediate representation.

[0051] The code compilation method provided by at least one embodiment of the present disclosure can be executed by a compiler. For example, the compiler can be located at the host side, and the code compilation method can be executed by the host side.

[0052] In the code compilation method provided by at least one embodiment of the present disclosure, thread-local storage model support under a parallel computing architecture is provided, solving the problem that traditional thread-local storage implementations cannot be directly applied to the GPU architecture. Through the above code compilation method, developers can directly use high-level syntax (such as __thread) for programming without manually managing thread data. The programming model is simple and easy to operate, greatly improving development efficiency. Moreover, through compilation optimization, the corresponding thread-local memory addresses are directly fixed during compilation in this code compilation method, without the need to load the thread-local memory addresses corresponding to each thread-local storage variable from registers, nor to dynamically obtain the thread-local memory addresses at runtime, achieving better performance in obtaining thread-local memory addresses.

[0053] Figure 1 The flowchart of a code compilation method provided by at least one embodiment of the present disclosure is used to compile code including at least one thread-local storage variable. For example, as Figure 1 shown, the code compilation method provided by at least one embodiment of the present disclosure includes the following steps S101 to S103.

[0054] Step S101: In response to the existence of target thread-local storage variables, based on the attributes of at least one target thread-local storage variable, allocate at least one target thread-local storage variable to thread-local memory to obtain the thread-local memory address corresponding to each target thread-local storage variable.

[0055] For example, in step S101, the target thread-local storage variable refers to the thread-local storage variable used by the first function among the above at least one thread-local storage variable. That is, the above-mentioned "at least one thread-local storage variable" may include thread-local storage variables used by the first function and may also include thread-local storage variables not used by any first function. Step S101 allocates the space of thread-local memory for the used thread-local storage variables, while the unused thread-local storage variables will not be allocated to thread-local memory.

[0056] In the code compilation method provided by at least one embodiment of the present disclosure, the code may include multiple first functions and multiple second functions. The first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0057] For example, in at least one embodiment of the present disclosure, the host side may include a central processing unit (CPU), and the device side may include a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), etc.

[0058] For example, in at least one embodiment of the present disclosure, the first function may be a kernel function, and the second function may be a device function. A kernel function refers to a function called by the host side and executed by the device side, and a device function refers to a function called by the device side and executed by the device side. For example, a kernel function may be called by a host function (such as a host function) running on a CPU, and a device function may be called by a kernel function executed at a GPU or other device functions. The code usually includes multiple kernel functions and device functions. Each kernel function may call multiple device functions, and each device function may also be called by multiple kernel functions.

[0059] The "thread-local storage variable used by the first function" refers to the thread-local storage variable directly used by the first function and the thread-local storage variable used by the first function by calling the second function. The above thread-local storage variables are all target thread-local storage variables.

[0060] Thread Local Memory (TLM) refers to the memory area private to each thread, which is mainly used to store thread-private data and cannot be directly accessed by other threads. For example, the thread local memory may be the thread local memory in a GPU. By storing the target thread-local storage variable in the thread local memory, the burden on the registers in the GPU can be reduced. For example, the thread local memory may also be the thread local memory in other parallel computing architectures, and the embodiments of the present disclosure do not limit this.

[0061] It should be noted that thread local memory and the thread local storage introduced above are two different concepts. Thread local memory refers to a memory area, while thread local storage refers to a storage method of variables.

[0062] For example, in step S101, the target thread-local storage variable may be allocated to the thread local memory based on the attributes of the target thread-local storage variable, so that the thread local memory address corresponding to each target thread-local storage variable can be obtained. The attributes of the target thread-local storage variable may include, for example, variable type, variable size, alignment requirements, etc., and the embodiments of the present disclosure do not limit this.

[0063] Step S102: Update the intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread-local storage variable.

[0064] For example, the Clang front-end of LLVM can automatically generate the intermediate representation corresponding to the code. In LLVM, the TLSSupported member variable of the TargetInfo class can be used to control whether to enable the thread-local storage function. Setting the value of the TLSSupported member variable to true enables the support for the thread-local storage function in the LLVM front-end.

[0065] For example, assume the user writes the following code:

[0066] __device__ __thread int a = 1; / / LINE1

[0067] __global__ void kernel() {

[0068] c[1] = 2; / / LINE3

[0069] }

[0070] The above code line LINE1 is the declaration and definition of the thread-local storage variable a. In some examples, according to code line LINE1, the intermediate representation automatically generated by the Clang front-end of LLVM is as follows:

[0071] @a = thread_local addrspace(1) externally_initialized global i32 1,align 4

[0072] In the above intermediate representation, @a indicates that the variable a is defined as a global variable, thread_local indicates that the variable a is a thread-local storage variable, addrspace(1) specifies that the address space of the variable is global, externally_initialized indicates that the initialization of the variable is done externally (e.g., at runtime, rather than during compilation), global i32 1 indicates that a global variable is defined, the type is a 32-bit integer, the initial value is 1, and align 4 indicates that the alignment of the variable is 4-byte alignment.

[0073] The above code line LINE3 implements the assignment operation of the thread-local storage variable c[1]. In some examples, according to code line LINE3, the intermediate representation automatically generated by the Clang front-end of LLVM is as follows:

[0074] %1 = tail call ptr @llvm.threadlocal.address.p0(ptr addrspacecast(ptr addrspace(1) @c to ptr))

[0075] arrayidx1 = getelementptr inbounds [2 x float], ptr %1, i64 0, i64 1

[0076] store float 2.000000e+00, ptr %arrayidx1, align 4

[0077] In the above intermediate representation, the llvm.threadlocal.address.p0 instruction is used to obtain the address of the thread-local storage variable c. What %1 stores is the address of the thread-local storage variable c, and what %arrayidx1 stores is the address of the thread-local storage variable c[1]. The store instruction stores the floating-point number 2.0 into the storage location pointed to by %arrayidx1, that is, the storage location corresponding to c[1].

[0078] In step S102, the intermediate representation automatically generated by the Clang front end can be updated based on the result of variable allocation in step S101 to obtain an optimized intermediate representation.

[0079] In the code compilation method provided by at least one embodiment of the present disclosure, the intermediate representation corresponding to the code includes an instruction for obtaining the address of the target thread-local storage variable. An example of step S102 can be: for each target thread-local storage variable, replace the instruction for obtaining the address of the target thread-local storage variable with the thread-local memory address corresponding to the target thread-local storage variable.

[0080] For example, the instruction for obtaining the address of the target thread-local storage variable is, for example, llvm.threadlocal.address.p0. Traverse the call functions of each target thread-local storage variable, and directly replace all llvm.threadlocal.address.p0 instructions with the thread-local memory addresses corresponding to the target thread-local storage variables to obtain an optimized intermediate representation.

[0081] In some examples, the intermediate representation before optimization (that is, the intermediate representation automatically generated by the Clang front end) is as follows:

[0082] %0 = tail call ptr @llvm.threadlocal.address.p0(ptr addrspacecast(ptr addrspace(1) @th to ptr))

[0083] store i32 66, ptr %0, align 4

[0084] The corresponding optimized intermediate representation is as follows:

[0085] store i32 66, ptr addrspacecast (ptr addrspace(4) inttoptr (i32 24 to ptr addrspace(4)) to ptr), align 4

[0086] In the optimized intermediate representation, the llvm.threadlocal.address.p0 instruction is no longer needed. Here, 24 represents the thread-local memory address of the target thread-local storage variable th obtained through allocation in step S101. addrspace(4) specifies that the address space of the variable is local. addrspacecast represents address space conversion to generate correct thread-local memory read and write instructions, meeting the requirements of the target architecture.

[0087] It can be seen that the optimized intermediate representation fixes the corresponding thread-local memory address during compilation, eliminating the need to load the thread-local memory addresses corresponding to each target thread-local storage variable from registers and dynamically obtain thread-local memory addresses at runtime, achieving better performance in obtaining thread-local memory addresses.

[0088] Step S103: Generate the machine code corresponding to the code based on the updated intermediate representation.

[0089] For example, in step S103, the optimized intermediate representation obtained in step S102 can be converted into machine code suitable for the target hardware platform through steps such as instruction selection, register allocation, instruction scheduling, and code generation as introduced above. The optimization of the intermediate representation in step S102 helps reduce the workload of instruction selection and improve performance.

[0090] Specifically, in the CPU architecture, the instruction selection for thread-local storage variable access instructions (such as ISD::GlobalTLSAddress) is performed in the TargetLowering phase. The TargetLowering phase is mainly responsible for converting the intermediate representation into a low-level representation related to the target hardware platform. Specifically, during the process of converting the intermediate representation into machine code, access to TLS variables is achieved by using segment registers plus offsets, and corresponding read and write operations are completed. However, the code compilation method provided by at least one embodiment of the present disclosure replaces the instruction for obtaining the address of the target thread-local storage variable with the thread-local memory address corresponding to the target thread-local storage variable before instruction selection, effectively reducing the workload of instruction selection.

[0091] The code compilation method provided by at least one embodiment of the present disclosure further includes the following step S104. Step S104 may be executed before step S101.

[0092] Step S104: For each target thread local storage variable, determine the first function that uses the target thread local storage variable to build a variable-function mapping relationship.

[0093] In step S104, the transfer path of the variable in the function can be traced according to the call graph (also known as the call relationship graph) to determine which functions use the target thread local storage variable. The call graph is a directed graph that can represent the call relationship between functions in a program. The compiler can generate the call graph by means of, for example, static analysis. For example, the call graph builder integrated in the compiler can parse the call relationship between functions and generate the call graph. For example, the variable-function mapping relationship records which first functions each target thread local storage variable is used by, and the variable-function mapping relationship can be presented in the form of a mapping graph.

[0094] It should be noted that since some thread local storage variables in the code may only be defined but not used by any first function, not every thread local storage variable can find a corresponding first function in the variable-function mapping relationship.

[0095] In step S104, the "first function that uses the target thread local storage variable" includes the first function that directly or indirectly uses the target thread local storage variable. For example, "directly use" means that the first function directly uses the target thread local storage variable, and "indirectly use" means that the first function uses the target thread local storage variable by calling a second function.

[0096] Correspondingly, an example of step S104 may be: for each target thread local storage variable, determine the kernel function that directly or indirectly uses the target thread local storage variable to build a variable-function mapping relationship.

[0097] Figure 2 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure.

[0098] As Figure 2 shown, in the code compilation method provided by at least one embodiment of the present disclosure, an example of step S101 may include the following steps S201 to S202.

[0099] Step S201: Based on the attributes of at least one target thread local storage variable, determine the variable allocation order.

[0100] For example, in step S201, multiple target thread local storage variables can be sorted according to the attributes of the target thread local storage variables.

[0101] Step S202: Allocate at least one target thread local storage variable to the thread local memory in sequence according to the variable allocation order.

[0102] For example, in step S202, according to the variable allocation order determined in step S201, multiple target thread local storage variables can be allocated to corresponding positions in the thread local memory in sequence.

[0103] In the code compilation method provided in at least one embodiment of the present disclosure, an example of step S101 may further include the following step S203.

[0104] Step S203: For the target thread local storage variables corresponding to multiple first functions, use the maximum value among the allocated offsets corresponding to the multiple first functions as the allocation starting point of the target thread local storage variable in the thread local memory.

[0105] For example, in step S203, the target thread local storage variables corresponding to multiple first functions mean that the target thread local storage variable is directly or indirectly used by the multiple first functions. In some examples, direct use may be, for example, that the kernel function directly uses the target thread local storage variable, and indirect use may be, for example, that the kernel function uses the target thread local storage variable by calling a device function.

[0106] For example, assume that the thread local storage variable b is only used by the kernel function Kernel 1 and the kernel function Kernel 2 (the thread local storage variable b is a target thread local storage variable), and the allocated offset corresponding to the kernel function Kernel 1 is 4 bytes, and the allocated offset corresponding to the kernel function Kernel 2 is 0 bytes. It can be seen that the maximum value of the allocated offsets is 4 bytes, and it is used as the allocation starting point of the thread local storage variable b in the thread local memory, that is, the allocation starting address of the thread local storage variable b in the thread local memory is 4, rather than 0.

[0107] It should be noted that in the code compilation method provided by at least one embodiment of the present disclosure, for each kernel function, an independent thread-local memory space is allocated. The "allocation starting point in the thread-local memory" mentioned above refers to the offset relative to the starting address of each thread-local memory space. For example, kernel function Kernel 1 corresponds to a thread-local memory space T1, and kernel function Kernel 2 corresponds to a thread-local memory space T2. To simplify the program logic, for each kernel function, the starting address of its corresponding thread-local memory space is regarded as 0, that is, the starting addresses of T1 and T2 are both regarded as 0, and subsequent hardware processing can map them to different physical addresses. Therefore, in the above example, the storage position of the thread-local storage variable b in the thread-local memory spaces T1 and T2 starts from the starting address 0 and is offset backward by 4 bytes.

[0108] Since multiple kernel functions and multiple device functions may use the same thread-local storage variable, and the device function cannot determine which kernel functions it will be called by, it is required that the addresses allocated to the same thread-local storage variable in all kernel functions and device functions are the same, that is, no matter which kernel function or device function it is, it can find the thread-local storage variable in the thread-local memory through the same address. Through the method provided in step S203, the consistency of the address allocation of the same thread-local storage variable required above can be achieved, that is, both the kernel function and the device function can find the thread-local storage variable b in the thread-local memory through the address 4.

[0109] It should be noted that for the target thread-local storage variable used only by a single first function, there is no need to determine the allocation starting point in the thread-local memory according to step S203, and it can be allocated sequentially in order.

[0110] In the code compilation method provided by at least one embodiment of the present disclosure, the attributes of the target thread-local storage variable may include at least one of the number of first functions corresponding to the target thread-local storage variable, the alignment parameter of the target thread-local storage variable, the size of the target thread-local storage variable, and the definition order of the target thread-local storage variable.

[0111] For example, the number of first functions corresponding to the target thread-local storage variable can be understood as the number of first functions using the target thread-local storage variable. For example, assuming that there are n kernel functions directly or indirectly using the target thread-local storage variable, the number of first functions corresponding to the target thread-local storage variable is n.

[0112] For example, the alignment parameter of a target thread-local storage variable is a parameter used to explicitly specify the alignment manner of the variable in storage, usually in bytes, such as 4 bytes, 8 bytes, 16 bytes, etc. For example, when the alignment parameter of a target thread-local storage variable is 4 bytes, it indicates that the target thread-local storage variable needs to be aligned in 4-byte units during storage.

[0113] For example, the size of a target thread-local storage variable can be understood as the size of the storage space that the target thread-local storage variable needs to occupy, which is related to the type of the variable and is usually in bytes. For example, the size of an int-type target thread-local storage variable is 4 bytes. It should be noted that in different systems or compilers, the size of the same type of target thread-local storage variable may be different.

[0114] For example, the definition order of target thread-local storage variables refers to the order in which each target thread-local storage variable is defined in the code, which is determined by the user when writing the code.

[0115] For example, when determining the variable allocation order, one or more of the following attributes can be considered: the number of first functions corresponding to the target thread-local storage variable, the alignment parameter of the target thread-local storage variable, the size of the target thread-local storage variable, and the definition order of the target thread-local storage variable. For example, in some examples, the variable allocation order can be determined solely based on the number of first functions corresponding to at least one target thread-local storage variable, or can be determined solely based on the alignment parameter of at least one target thread-local storage variable, or can be determined solely based on the size of at least one target thread-local storage variable, or can be determined solely based on the definition order of at least one target thread-local storage variable. For example, in other examples, the variable allocation order can be determined based on the number of first functions corresponding to at least one target thread-local storage variable, the alignment parameter of at least one target thread-local storage variable, the size of at least one target thread-local storage variable, and the definition order of at least one target thread-local storage variable. For example, two or three of the above four attributes can also be arbitrarily selected as the basis for determining the variable allocation order.

[0116] It should be noted that the above three attributes are only one example, and the variable allocation order can also be determined according to other attributes according to actual needs, and the embodiments of the present disclosure do not limit this.

[0117] In the code compilation method provided by at least one embodiment of the present disclosure, an example of step S201 may include any one of steps S301 to S304 below. For example, if only one attribute is considered when determining the variable allocation order, the corresponding step can be selected from steps S301 to S304 according to this attribute for execution.

[0118] Step S301: Arrange at least one target thread local storage variable in descending order according to the number of first functions corresponding to the at least one target thread local storage variable.

[0119] Step S302: Arrange at least one target thread local storage variable in descending order according to the alignment parameter of the at least one target thread local storage variable.

[0120] Step S303: Arrange at least one target thread local storage variable in ascending order according to the size of the at least one target thread local storage variable.

[0121] Step S304: Arrange at least one target thread local storage variable in the order of definition of the at least one target thread local storage variable.

[0122] For example, in step S301, the target thread local storage variables can be sorted from largest to smallest according to the number of first functions corresponding to the target thread local storage variables. For example, assume that there are three target thread local storage variables a, b, and c to be sorted, where the number of first functions corresponding to the target thread local storage variable a is 4, the number of first functions corresponding to the target thread local storage variable b is 5, and the number of first functions corresponding to the target thread local storage variable c is 2. Then the sorting result is b, a, c. For example, in some examples, if there are multiple target thread local storage variables with the same corresponding number of first functions, they can be further sorted according to the definition order of the variables.

[0123] For example, in step S302, the target thread local storage variables can be sorted from largest to smallest according to the alignment parameter of the target thread local storage variables. For example, assume that there are three target thread local storage variables a, b, and c to be sorted, where the alignment parameter of the target thread local storage variable a is 4 bytes, the alignment parameter of the target thread local storage variable b is 8 bytes, and the alignment parameter of the target thread local storage variable c is 16 bytes. Then the sorting result is c, b, a. For example, in some examples, if there are multiple target thread local storage variables with the same alignment parameter, they can be further sorted according to the definition order of the variables.

[0124] For example, in step S303, the target thread-local storage variables can be sorted in ascending order of their sizes. For example, assume there are three target thread-local storage variables a, b, and c to be sorted, where the size of target thread-local storage variable a is 8 bytes, the size of target thread-local storage variable b is 1 byte, and the size of target thread-local storage variable c is 4 bytes. Then the sorting result is b, c, a. For example, in some examples, if there are multiple target thread-local storage variables of the same size, they can be further sorted in the order of variable definition.

[0125] For example, in step S304, the target thread-local storage variables can be sorted in the order of their definitions. For example, assume there are three target thread-local storage variables a, b, and c to be sorted, and the definition order is a, b, c. Then the sorting result is a, b, c.

[0126] It should be noted that the above sorting method is to make the allocated thread-local memory gap as small as possible and improve the utilization rate of thread-local memory. According to actual needs, sorting can also be performed in other ways, and the embodiments of the present disclosure do not limit this.

[0127] Figure 3 This is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure.

[0128] As Figure 3 shown, in the code compilation method provided by at least one embodiment of the present disclosure, if all four attributes above are considered when determining the variable allocation order, one example of the above step S201 may include the following steps S401 to S404.

[0129] Step S401: Arrange at least one target thread-local storage variable in descending order of the corresponding first function count.

[0130] Step S402: In response to there being multiple first thread-local storage variables with the same corresponding first function count among at least one target thread-local storage variable, arrange the multiple first thread-local storage variables in descending order of the alignment parameter of the target thread-local storage variable.

[0131] Step S403: In response to there being multiple second thread-local storage variables with the same alignment parameter among the multiple first thread-local storage variables, arrange the multiple second thread-local storage variables in ascending order of the size of the target thread-local storage variable.

[0132] Step S404: In response to the existence of multiple third thread - local storage variables with the same size among the multiple second thread - local storage variables, arrange the multiple third thread - local storage variables in the defined order of the target thread - local storage variables.

[0133] For example, in step S401, the target thread - local storage variables can be sorted in descending order according to the number of first functions corresponding to the target thread - local storage variables.

[0134] For example, in step S402, according to the sorting result of step S401, if it is determined that there are multiple first thread - local storage variables with the same number of first functions among the above - mentioned target thread - local storage variables, then arrange the multiple first thread - local storage variables in descending order according to the alignment parameters of the target thread - local storage variables. It should be noted that step S402 only processes the variables whose order has not been determined in step S401 (i.e., the first thread - local storage variables), and the variables whose order has been determined are not considered.

[0135] For example, in step S403, according to the sorting result of step S402, if it is determined that there are multiple second thread - local storage variables with the same alignment parameters among the multiple first thread - local storage variables, then arrange the multiple second thread - local storage variables in ascending order according to the size of the target thread - local storage variables. It should be noted that step S403 only processes the variables whose order has not been determined in step S402 (i.e., the second thread - local storage variables), and the variables whose order has been determined are not considered.

[0136] For example, in step S404, according to the sorting result of step S403, if it is determined that there are multiple third thread - local storage variables with the same size among the multiple second thread - local storage variables, then arrange the multiple third thread - local storage variables in the defined order of the target thread - local storage variables to obtain the final sorting result. Here, the defined order refers to the order in which the variables are defined when the user writes the code. It should be noted that step S404 only processes the variables whose order has not been determined in step S403 (i.e., the third thread - local storage variables), and the variables whose order has been determined are not considered.

[0137] It should be noted that the above steps S401 to S404 are only examples, providing an attribute priority as follows: 1. The number of first functions corresponding to the target thread local storage variable; 2. The alignment parameter of the target thread local storage variable; 3. The size of the target thread local storage variable; 4. The definition order of the target thread local storage variable. This setting of attribute priority is to make the allocated thread local memory gap as small as possible and improve the utilization rate of thread local memory. According to actual needs, sorting can also be performed according to other attribute priorities. For example, sorting can be first performed according to size and then further according to alignment parameters, etc. The embodiments of the present disclosure do not limit this.

[0138] The following is a specific example of the above steps S401 to S404.

[0139] For example, assume that there are six thread local storage variables a, b, c, d, e, f to be sorted (all are target thread local storage variables). Among them, the number of first functions corresponding to the thread local storage variable a is 4, the number of first functions corresponding to the thread local storage variable b is 5, and the number of first functions corresponding to the thread local storage variables c, d, e, f is 2. Executing step S401, the sorting result is b, a, (c, d, e, f), where (c, d, e, f) means that the relative order between the variables c, d, e, f has not been determined yet, and currently it can only be confirmed that these four variables are all after the variable a.

[0140] Since the number of first functions of the variables c, d, e, f is the same (that is, c, d, e, f are all first thread local storage variables), step S402 needs to be executed to determine the relative order between these four variables according to the alignment parameter. For example, further assume that the alignment parameters of the thread local storage variables c, d, e are all 4 bytes, and the alignment parameter of the thread local storage variable f is 8 bytes. Executing step S402, the further sorting result is b, a, f, (c, d, e), where (c, d, e) means that the relative order between the variables c, d, e has not been determined yet, and currently it can only be confirmed that these three variables are all after the variable f.

[0141] Since the alignment parameters of the variables c, d, e are the same (that is, c, d, e are all second thread local storage variables), step S403 needs to be executed to determine the relative order between these three variables according to the size. For example, further assume that the size of the thread local storage variable c is 2 bytes, and the sizes of the thread local storage variables d and e are both 1 byte. Executing step S403, the further sorting result is b, a, f, (d, e), c, where (d, e) means that the relative order between the variables d, e has not been determined yet, and currently it can only be confirmed that these two variables are all after the variable f and before the variable c.

[0142] Since the variables d and e are of the same size (i.e., both d and e are third-thread local storage variables), step S404 needs to be executed to determine the relative order between these two variables according to the definition order. For example, further assume that the thread local storage variable d is defined first and the thread local storage variable e is defined later. After executing step S404, the further sorting result is b, a, f, d, e, c, which is the final sorting result.

[0143] Figure 4A Exemplary schematic diagram of a code compilation method provided by at least one embodiment of the present disclosure.

[0144] For example, as Figure 4A shown, the code includes three kernel functions and six device functions. The three kernel functions are respectively denoted as Kernel 1 to Kernel 3, and the six device functions are respectively denoted as Func 1 to Func 6. As Figure 4A shown in the right call graph, the kernel function Kernel 1 calls the device functions Func 1 and Func 2, the kernel function Kernel 2 calls the device functions Func3 and Func4, and the kernel function kernel3 calls the device functions Func5 and Func6. Among them, the device function Func1 uses the thread local storage variable a, the device function Func2 uses the thread local storage variable b, the device function Func3 uses the thread local storage variable b, the device function Func4 uses the thread local storage variable c, the device function Func5 uses the thread local storage variable c, and the device function Func6 uses the thread local storage variable a. The alignment parameters of the thread local storage variables a, b, and c are all 4 bytes, and the sizes are also all 4 bytes. The above thread local storage variables a, b, and c are all target thread local storage variables.

[0145] Execute the above steps S401 to S404 to sort the thread local storage variables a, b, and c. As can be seen from the above, the number of first functions corresponding to the thread local storage variable a is 2 (kernel functions Kernel 1 and Kernel 3), the number of first functions corresponding to the thread local storage variable b is 2 (kernel functions Kernel 1 and Kernel 2), and the number of first functions corresponding to the thread local storage variable c is 2 (kernel functions Kernel 2 and Kernel 3). Therefore, the number of first functions, alignment parameters, and sizes of the thread local storage variables a, b, and c are all the same, and the final sorting result cannot be obtained according to steps S401 to S403. Execute step S404 to sort the thread local storage variables a, b, and c according to the definition order, and the final sorting result is a, b, c.

[0146] First, allocate the thread-local storage variable a. The thread-local storage variable a is used by kernel function Kernel 1 and kernel function Kernel 3, and the allocated offsets corresponding to kernel functions Kernel 1 and Kernel 3 are both 0 bytes, that is, the maximum value of the allocated offset is 0 bytes. Therefore, the starting address of the allocation of the thread-local storage variable a in the thread-local memory is 0, and 4 bytes of thread-local memory space are allocated.

[0147] Next, allocate the thread-local storage variable b. The thread-local storage variable b is used by kernel function Kernel 1 and kernel function Kernel 2, and the allocated offset corresponding to kernel function Kernel 1 is 4 bytes, and the allocated offset corresponding to kernel function Kernel 2 is 0 bytes, that is, the maximum value of the allocated offset is 4 bytes. Therefore, the starting address of the allocation of the thread-local storage variable b in the thread-local memory is 4, and 4 bytes of thread-local memory space are allocated.

[0148] Finally, allocate the thread-local storage variable c. The thread-local storage variable c is used by kernel function Kernel 2 and kernel function Kernel 3, and the allocated offset corresponding to kernel function Kernel 2 is 8 bytes, and the allocated offset corresponding to kernel function Kernel 3 is 4 bytes, that is, the maximum value of the allocated offset is 8 bytes. Therefore, the starting address of the allocation of the thread-local storage variable c in the thread-local memory is 8, and 4 bytes of thread-local memory space are allocated.

[0149] Figure 4B Exemplary schematic diagram of a code compilation method provided by at least one embodiment of the present disclosure.

[0150] For example, as Figure 4B shown, the code includes two kernel functions and four device functions. The two kernel functions are respectively denoted as Kernel 1~Kernel 2 for example, and the four device functions are respectively denoted as Func 1~Func 4 for example. As Figure 4B shown in the right call graph, kernel function Kernel 1 calls device functions Func 1 and Func 2, and kernel function Kernel 2 calls device functions Func3 and Func4. Among them, device function Func1 uses thread-local storage variable a, device function Func2 uses thread-local storage variable b, device function Func3 uses thread-local storage variable c, and device function Func4 uses thread-local storage variable d. The above thread-local storage variables a, b, c, and d are all target thread-local storage variables. It can be seen that there are no shared thread-local storage variables between kernel functions Kernel 1 and Kernel 2, and they can be compactly allocated in order. In this case, there are no gaps in the thread-local memory, and there is no waste of thread-local memory space.

[0151] In the code compilation method provided by at least one embodiment of the present disclosure, an independent thread-local memory space is allocated for each kernel function. Therefore, the "allocation start address in thread-local memory" mentioned above refers to the offset relative to the start address of each thread-local memory space. It should be noted that Figure 4A and Figure 4B in which a stack is used to illustrate the allocation of target thread-local storage variables, the stack here can be regarded as a virtual concept. In LLVM, when the intermediate representation is converted into machine code, the allocation on the stack will be mapped to the actual storage space of the target architecture, that is, mapped to the corresponding thread-local memory space according to the position of the target thread-local storage variables on the stack.

[0152] It should be noted that thread-local memory can also be used to store local variables other than target thread-local storage variables. For example, as Figure 4A and Figure 4B shown, the stack corresponding to each kernel function can also include local variables.

[0153] The compiler can automatically manage and optimize memory allocation. By optimizing the frame layout of target thread-local storage variables, the utilization rate of the stack can be effectively improved, and the overall occupied space of thread-local memory can be reduced.

[0154] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S105. Step S105 may be executed before step S102.

[0155] Step S105: Insert an instruction for initializing the target thread-local storage variables used by each first function into the intermediate representation corresponding to the code.

[0156] For example, in step S105, the code may include multiple first functions and multiple second functions. According to the call graph, the target thread-local storage variables directly or indirectly used by each first function can be determined. For each first function, an instruction for initializing the target thread-local storage variables directly or indirectly used by it can be inserted at the starting position of its corresponding intermediate representation.

[0157] For example, taking the target thread-local storage variable a as an example, an example of an instruction for initializing the target thread-local storage variable a is as follows:

[0158] %0 = call ptr @llvm.threadlocal.address.p0(ptr addrspacecast (ptr addrspace(1) @a to ptr))

[0159] store i32 1, ptr %0, align 4

[0160] The meaning of the above intermediate representation is that first, the address of the target thread-local storage variable a is obtained through the llvm.threadlocal.address.p0 instruction, and then the initial value of variable a (which is 1 in this example) is stored in this address through the store instruction. It should be noted that the above initialization instruction is only an example, and different initialization instructions can be inserted according to actual needs.

[0161] It should be noted that the above initialization instruction will also be updated together in step S102, that is, the llvm.threadlocal.address.p0 instruction therein will be replaced with the thread-local memory address corresponding to the target thread-local storage variable a. Therefore, the initialization instruction introduced here will not bring memory access overhead for reading and writing addresses from registers.

[0162] In the CPU architecture, thread-local storage variables are usually initialized by the runtime environment when the program starts. Different from the above initialization method, the code compilation method provided by at least one embodiment of the present disclosure inserts instructions for initializing the target thread-local storage variables during compilation, rather than initializing the target thread-local storage variables at runtime, which can effectively avoid refreshing the entire thread-local memory at runtime and achieve performance optimization. In addition, this method increases more opportunities for compilation optimization, reduces the coupling between compilation and runtime, and reduces the workload during compilation and runtime.

[0163] For example, in at least one embodiment of the present disclosure, before step S105, the usage of each thread-local storage variable can also be analyzed for each first function first. For thread-local storage variables that are not used, their corresponding initialization instructions do not need to be inserted. For example, the above analysis operation can be implemented by the compiler. By the above method, unnecessary initialization operations can be reduced and performance improvement can be achieved.

[0164] For example, in order to further optimize performance, redundant instruction optimization can also be performed on the intermediate representation corresponding to the code before step S101.

[0165] For example, in some examples, for unused thread-local storage variables, the instructions for obtaining the addresses of thread-local storage variables can be removed from the intermediate representation corresponding to the code.

[0166] For example, the intermediate representation automatically generated by the Clang front-end of LLVM can be analyzed, that is, for each first function, the usage of thread-local storage variables by each thread is analyzed. For thread-local storage variables that are not used, the instructions for obtaining the addresses of thread-local storage variables are removed from the intermediate representation corresponding to the first function.

[0167] For example, in some other examples, duplicate instructions for obtaining the addresses of thread-local storage variables can be removed from the intermediate representation corresponding to the code, so that the address of each thread-local storage variable within each first function is obtained only once. For example, if there are multiple identical instructions for obtaining the addresses of thread-local storage variables in the intermediate representation corresponding to a certain first function, the duplicate instructions need to be removed, and only the first instruction is retained.

[0168] A special example is as follows. Assume that c is a variable of array type. For a kernel function that uses thread-local storage variables c[1] and c[2], the instructions for obtaining the addresses of thread-local storage variables c[1] and c[2] in the intermediate representation need to be removed, and only the instruction for obtaining the base address of thread-local storage variable c is retained. The addresses of c[1] and c[2] can be calculated by adding corresponding offsets to the base address of c.

[0169] By the above method, duplicate address acquisition operations can be avoided, significantly improving the running efficiency and reducing the overhead of redundant instructions.

[0170] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S106. Step S106 may be executed before step S101.

[0171] Step S106: Obtain a record of the thread-local memory size required for all target thread-local storage variables used by each first function in the code, so as to pre-allocate a fixed thread-local memory space for each first function according to the record.

[0172] Frame Lowering is a stage of the LLVM backend, which is responsible for converting the abstract stack frame operations in the intermediate representation into machine code suitable for the target hardware platform. Its main task is to generate the prologue and epilogue of function calls, that is, the stack frame management code at the entrance and exit of the function. These instructions are mainly used to set up and restore the stack frame, so that the function can run correctly and maintain the integrity of the call chain.

[0173] For example, during the compilation stage, the size of the thread-local memory required for each kernel function can be determined through step S101. In step S106, the size of the thread-local memory corresponding to all target thread-local storage variables used by each kernel function can be recorded through metadata. It should be noted that the thread-local memory is also used to store local variables other than the target thread-local storage variables. What is recorded here is only the size of the thread-local memory required by all target thread-local storage variables used by the kernel function, excluding the thread-local memory size occupied by other local variables. For example, when the compiler generates the executable file corresponding to the first function, it usually embeds a piece of metadata. Metadata can be regarded as a kind of additional information, which can be used to guide compiler optimization, debugging, code analysis, or runtime behavior. These metadata records of the thread-local memory size required by each kernel function can be obtained at the initial stage of backend code generation, and based on this, a fixed thread-local memory space can be pre-allocated for each kernel function. Among them, the lifecycle of the pre-allocated thread-local memory space is consistent with the execution cycle of the corresponding kernel function, that is, the lifecycle lasts until the corresponding kernel function ends. For device functions, the target thread-local storage variables can also be obtained from the thread-local memory space corresponding to the kernel function. Therefore, through the above pre-allocation method, in the frame degradation stage, there is no need to perform additional processing on the prologue and epilogue of each function to manage the target thread-local storage variables, because the space for these variables has been reserved in advance.

[0174] Figure 5 A flowchart of a code compilation method provided by at least one embodiment of the present disclosure is used to compile code including at least one thread-local storage variable. For example, as Figure 5 shown, the code compilation method provided by at least one embodiment of the present disclosure includes the following steps S501 to S506.

[0175] Step S501: For each target thread-local storage variable, determine the first function that uses the target thread-local storage variable to construct a variable-function mapping relationship. The description of step S501 can refer to the description of step S104 in the above embodiment, and will not be elaborated here.

[0176] Step S502: In response to the existence of target thread-local storage variables, allocate at least one target thread-local storage variable to the thread-local memory based on the attributes of at least one target thread-local storage variable to obtain the thread-local memory address corresponding to each target thread-local storage variable. The description of step S502 can refer to the description of step S101 in the above embodiment, and will not be elaborated here.

[0177] Step S503: Obtain a record of the thread-local memory size required for all target thread-local storage variables used by each first function in the code, so as to pre-allocate a fixed thread-local memory space for each first function according to the record. For the description of step S503, reference can be made to the description of step S106 in the above embodiment, which will not be elaborated here.

[0178] Step S504: Insert instructions for initializing the target thread-local storage variables used by each first function in the intermediate representation corresponding to the code. For the description of step S504, reference can be made to the description of step S105 in the above embodiment, which will not be elaborated here.

[0179] Step S505: Update the intermediate representation corresponding to the code based on the thread-local memory addresses corresponding to each target thread-local storage variable. For the description of step S505, reference can be made to the description of step S102 in the above embodiment, which will not be elaborated here.

[0180] Step S506: Generate the machine code corresponding to the code based on the updated intermediate representation. For the description of step S506, reference can be made to the description of step S103 in the above embodiment, which will not be elaborated here.

[0181] It should also be noted that in various embodiments of the present disclosure, the execution order of the steps of the code compilation method is not limited. Although the execution processes of the steps are described in a specific order above, this does not constitute a limitation on the embodiments of the present disclosure. The steps in the code compilation method can be executed serially or in parallel, which can be determined according to actual needs.

[0182] For example, compared with the above description, the code compilation method provided by at least one embodiment of the present disclosure may further include more or fewer steps, and the embodiments of the present disclosure do not limit this.

[0183] Figure 6 It is a schematic block diagram of a code compilation device provided by at least one embodiment of the present disclosure. The code compilation device may be, for example, a compiler for compiling code including at least one thread-local storage variable.

[0184] For example, as Figure 6 shown, the code compilation device provided by at least one embodiment of the present disclosure includes an allocation module 601, an update module 602, and a generation module 603.

[0185] For example, the allocation module 601 is configured to, in response to the presence of target thread-local storage variables, allocate at least one target thread-local storage variable to the thread-local memory based on the attributes of at least one target thread-local storage variable, to obtain a thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is a thread-local storage variable used by a first function among at least one thread-local storage variable. For the related content of the allocation module 601, reference may be made to the related description of step S101 in the embodiment of the above code compilation method, which will not be elaborated herein.

[0186] For example, the update module 602 is configured to update the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable. For the related content of the update module 602, reference may be made to the related description of step S102 in the embodiment of the above code compilation method, which will not be elaborated herein.

[0187] For example, the generation module 603 is configured to generate machine code corresponding to the code based on the updated intermediate representation. For the related content of the generation module 603, reference may be made to the related description of step S103 in the embodiment of the above code compilation method, which will not be elaborated herein.

[0188] For example, in at least one embodiment of the present disclosure, the code compilation device 600 may further include an initialization module, configured to insert instructions for initializing the target thread-local storage variables used by each first function into the intermediate representation corresponding to the code.

[0189] For example, in at least one embodiment of the present disclosure, the code compilation device 600 may further include a construction module, configured to, for each target thread-local storage variable, determine the first function using the target thread-local storage variable, so as to construct a variable-function mapping relationship.

[0190] For example, in at least one embodiment of the present disclosure, the first function directly uses the target thread-local storage variable, or uses the target thread-local storage variable by calling a second function.

[0191] For example, in at least one embodiment of the present disclosure, the first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0192] For example, in at least one embodiment of the present disclosure, the allocation module includes a determination unit and an allocation unit. The determination unit is configured to determine the variable allocation order based on the attributes of at least one target thread-local storage variable. The allocation unit is configured to sequentially allocate at least one target thread-local storage variable to the thread-local memory according to the variable allocation order.

[0193] For example, in at least one embodiment of the present disclosure, the attributes of the target thread local storage variable include at least one of the following: the number of first functions corresponding to the target thread local storage variable, the alignment parameter of the target thread local storage variable, the size of the target thread local storage variable, and the definition order of the target thread local storage variable.

[0194] For example, in at least one embodiment of the present disclosure, the determining unit is further configured to: arrange at least one target thread local storage variable in descending order according to the number of first functions corresponding to at least one target thread local storage variable; arrange at least one target thread local storage variable in descending order according to the alignment parameter of at least one target thread local storage variable; arrange at least one target thread local storage variable in ascending order according to the size of at least one target thread local storage variable; or arrange at least one target thread local storage variable according to the definition order of at least one target thread local storage variable.

[0195] For example, in at least one embodiment of the present disclosure, the determining unit is further configured to: arrange at least one target thread local storage variable in descending order according to the number of first functions corresponding to at least one target thread local storage variable; in response to there being multiple first thread local storage variables with the same number of corresponding first functions in a target thread local storage variable, arrange the multiple first thread local storage variables in descending order according to the alignment parameter of the target thread local storage variable; in response to there being multiple second thread local storage variables with the same alignment parameter among the multiple first thread local storage variables, arrange the multiple second thread local storage variables in ascending order according to the size of the target thread local storage variable; in response to there being multiple third thread local storage variables with the same size among the multiple second thread local storage variables, arrange the multiple third thread local storage variables according to the definition order of the target thread local storage variable.

[0196] For example, in at least one embodiment of the present disclosure, the allocation module may further include a starting address determining unit, configured to, for the target thread local storage variable corresponding to multiple first functions, use the maximum value among the allocated offsets corresponding to the multiple first functions as the allocation starting point of the target thread local storage variable in the thread local memory.

[0197] For example, in at least one embodiment of the present disclosure, the intermediate representation corresponding to the code includes an instruction for obtaining the address of the target thread local storage variable, and the updating module is further configured to: for each target thread local storage variable, replace the instruction for obtaining the address of the target thread local storage variable with the thread local memory address corresponding to the target thread local storage variable.

[0198] For example, in at least one embodiment of the present disclosure, the code compilation device 600 may further include a record acquisition module configured to acquire a record of the thread-local memory size required for all target thread-local storage variables used by each first function in the code, so as to pre-allocate a fixed thread-local memory space for each first function according to the record.

[0199] For example, in at least one embodiment of the present disclosure, the life cycle of the pre-allocated thread-local memory space is consistent with the execution cycle of the corresponding first function.

[0200] For example, in at least one embodiment of the present disclosure, the thread-local memory is located in the graphics processing unit.

[0201] It should be noted that the above various modules and units can be implemented by software, hardware, firmware, or any combination thereof. For example, the allocation module, the update module, and the generation module can be respectively implemented as an allocation circuit, an update circuit, and a generation circuit. The embodiments of the present disclosure do not limit their specific implementation manners.

[0202] It should be understood that the code compilation device 600 provided in at least one embodiment of the present disclosure can be used to implement the foregoing code compilation method, and can also achieve technical effects similar to those of the foregoing code compilation method, which will not be elaborated herein.

[0203] It should be noted that in the embodiments of the present disclosure, the code compilation device 600 may include more or fewer modules or units, and the connection relationships between the various modules or units are not limited and can be determined according to actual needs. The specific composition manners of the various modules or units are not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or in other applicable manners.

[0204] Figure 7 It is a schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure.

[0205] For example, as Figure 7 shown, the electronic device 700 includes at least one processor 701 and at least one memory 702. The at least one memory 702 includes one or more computer program modules. The one or more computer program modules are stored in the memory 702 and are configured to be executed by the at least one processor 701. The one or more computer program modules include instructions for executing the foregoing code compilation method. When executed by the at least one processor 701, they can execute one or more steps in the code compilation method provided in at least one embodiment of the present disclosure. The memory 702 and the processor 701 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0206] For example, the processor 701 can be a central processing unit (CPU), a digital signal processor (DSP), an image processor (GPU), a general-purpose graphics processing unit (GPGPU), an artificial intelligence (AI) accelerator, or other forms of processing units with data processing capabilities and / or program execution capabilities, such as a field-programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) can be of the X86, ARM, RISC-V architectures, etc. The processor 701 can be a general-purpose processor or a dedicated processor, and can control other components in the electronic device 700 to perform desired functions.

[0207] For example, the memory 702 can include any combination of one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.

[0208] Figure 8 Schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure.

[0209] The electronic device in at least one embodiment of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0210] The electronic device includes at least one processor and a memory. The processor here can be referred to as the processing device 801 described below, and the memory can include at least one of the read-only memory (ROM), random access memory (RAM), and storage device 808 described below. The memory is used to store programs for executing the methods described in the above method embodiments; the processor is configured to execute the programs stored in the memory. The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.

[0211] As Figure 8As shown, the electronic device 800 may include a processing device 801 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device 808 into a random access memory (RAM). In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface is also connected to the bus 804.

[0212] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a display, a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0213] Specifically, according to at least one embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, at least one embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the method of at least one embodiment of the present disclosure are executed.

[0214] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In at least one embodiment of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device. In at least one embodiment of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0215] The above computer-readable medium can be included in the above electronic device 800; or it can exist separately without being assembled into the electronic device 800.

[0216] Figure 9 It is a schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure.

[0217] For example, as Figure 9 shown, computer-readable instructions 901 are stored on the non-transitory computer-readable storage medium 900, and when the computer-readable instructions 901 are executed by at least one processor, one or more steps of the above code compilation method are executed.

[0218] For example, the storage medium may include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, and may also be other applicable storage media. For example, the readable storage medium may also be the Figure 7 memory 702 in, and for the relevant description, reference may be made to the foregoing content, which will not be elaborated herein.

[0219] Although the present disclosure has been described in detail with general descriptions and specific embodiments above, based on the embodiments of the present disclosure, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure all fall within the scope of protection required by the present disclosure.

[0220] The following points need to be noted regarding the present disclosure:

[0221] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.

[0222] (2) For the sake of clarity, in the accompanying drawings used to describe the embodiments of the present disclosure, the thickness of the layer or region is enlarged or reduced, that is, these drawings are not drawn according to the actual ratio.

[0223] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0224] As described above, the above is only the specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A code compilation method, wherein, The code includes at least one thread-local storage variable, and the method for compiling the code includes: In response to the existence of target thread-local storage variables, based on the attributes of at least one target thread-local storage variable, allocate the at least one target thread-local storage variable to thread-local memory to obtain the thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is the thread-local storage variable used by a first function among the at least one thread-local storage variable; Update the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; Generate the machine code corresponding to the code based on the updated intermediate representation.

2. The method for compiling the code according to claim 1, further includes: Insert instructions for initializing the target thread-local storage variables used by each first function in the intermediate representation corresponding to the code.

3. The method for compiling the code according to claim 1, further includes: For each target thread-local storage variable, determine the first function that uses the target thread-local storage variable to construct a variable-function mapping relationship.

4. The code compilation method according to claim 3, wherein, The first function directly uses the target thread-local storage variable, or uses the target thread-local storage variable by calling a second function.

5. The code compilation method according to claim 4, wherein, The first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

6. The code compilation method according to claim 1, wherein, The allocating the at least one target thread-local storage variable to thread-local memory based on the attributes of at least one target thread-local storage variable includes: Determine the variable allocation order based on the attributes of the at least one target thread-local storage variable; Allocate the at least one target thread-local storage variable to the thread-local memory in sequence according to the variable allocation order.

7. The code compilation method according to claim 6, wherein, The attributes of the target thread-local storage variable include at least one of the number of first functions corresponding to the target thread-local storage variable, the alignment parameter of the target thread-local storage variable, the size of the target thread-local storage variable, and the definition order of the target thread-local storage variable.

8. The code compilation method according to claim 7, wherein, The determining the variable allocation order based on the attributes of the at least one target thread-local storage variable includes: Arrange the at least one target thread-local storage variable in descending order according to the number of first functions corresponding to the at least one target thread-local storage variable; Arrange the at least one target thread-local storage variable in descending order according to the alignment parameter of the at least one target thread-local storage variable; Arrange the at least one target thread-local storage variable in ascending order according to the size of the at least one target thread-local storage variable; or Arrange the at least one target thread-local storage variable according to the definition order of the at least one target thread-local storage variable.

9. The code compilation method according to claim 7, wherein, The determining the variable allocation order based on the attributes of the at least one target thread-local storage variable includes: Arrange the at least one target thread-local storage variable in descending order according to the number of first functions corresponding to the at least one target thread-local storage variable; In response to the existence of a plurality of first thread - local storage variables with the same number of corresponding first functions in the at least one target thread - local storage variable, arranging the plurality of first thread - local storage variables in descending order according to the alignment parameter of the target thread - local storage variable; In response to the existence of a plurality of second thread - local storage variables with the same alignment parameter in the plurality of first thread - local storage variables, arranging the plurality of second thread - local storage variables in ascending order according to the size of the target thread - local storage variable; In response to the existence of a plurality of third thread - local storage variables with the same size in the plurality of second thread - local storage variables, arranging the plurality of third thread - local storage variables in the defined order of the target thread - local storage variable.

10. The code compilation method according to claim 6, wherein, The allocating the at least one target thread - local storage variable to the thread - local memory based on the attributes of the at least one target thread - local storage variable further includes: For the target thread - local storage variable corresponding to a plurality of first functions, taking the maximum value of the allocated offsets corresponding to the plurality of first functions as the allocation starting point of the target thread - local storage variable in the thread - local memory.

11. The code compilation method according to claim 1, wherein, The updating the intermediate representation corresponding to the code based on the thread - local memory address corresponding to each target thread - local storage variable includes: For each target thread - local storage variable, replacing the instruction for obtaining the address of the target thread - local storage variable with the thread - local memory address corresponding to the target thread - local storage variable, wherein the intermediate representation corresponding to the code includes the instruction for obtaining the address of the target thread - local storage variable.

12. The code compilation method according to claim 1, further includes: Obtaining a record of the thread - local memory sizes required by all target thread - local storage variables used by each first function in the code, so as to pre - allocate a fixed thread - local memory space for each first function according to the record.

13. The code compilation method according to claim 12, wherein, The life cycle of the pre - allocated thread - local memory space is consistent with the execution cycle of the corresponding first function.

14. An electronic device, comprising: At least one processor; At least one memory, including one or more computer program modules; wherein the one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the code compilation method according to any one of claims 1 - 13.

15. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein, When the computer instructions are executed by the at least one processor, the code compilation method according to any one of claims 1 - 13 is executed.

Citation Information

Patent Citations

  • Method and device for transplanting code from process model to thread model

    CN105786525A

  • Data processing method and device, storage medium and electronic equipment

    CN119127345A