Code compiling method, electronic equipment and storage medium

Through a code compilation method, the problem that traditional thread-local storage implementation cannot be applied to GPU architecture is solved, and the thread-local storage model is supported is implemented, which simplifies the development process and improves performance.

CN120122957AActive Publication Date: 2025-06-10SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510534557.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-10
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional thread-local storage implementations cannot be directly applied to GPU architecture, making it difficult to support thread-local storage models under parallel computing architectures, increasing development complexity and performance losses.

Method used

Provides a code compilation method, which solves the allocation and access of thread-local storage variables on the GPU architecture by responding to target thread-local storage variables, allocating them to thread-local storage variables based on their attributes, updating the intermediate representation, and generating machine code.

Benefits of technology

It realizes thread-local storage model support under the parallel computing architecture, simplifies the development process, improves development efficiency, and achieves better performance in obtaining thread-local memory addresses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122957A_ABST
    Figure CN120122957A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a code compiling method, electronic equipment and a storage medium. The code compiling method is used for compiling a code comprising at least one thread local storage variable, and comprises the following steps: in response to existence of a target thread local storage variable, allocating the at least one target thread local storage variable to a thread local memory based on an attribute of the at least one target thread local storage variable, obtaining a thread local memory address corresponding to each target thread local storage variable, wherein the target thread local storage variable is the thread local storage variable used by the first function; updating an intermediate representation corresponding to the code based on a thread local memory address corresponding to each target thread local storage variable; and generating a machine code corresponding to the code based on the updated intermediate representation. According to the code compiling method, thread local storage model support under a parallel computing architecture is provided, and the problem that traditional thread local storage implementation cannot be directly applied to a GPU architecture is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a code compilation method, an electronic device, and a storage medium. Background Art

[0002] The Low-Level Virtual Machine (LLVM) is an open-source compiler framework, which is designed as a collection of modular and reusable compiler and toolchain technologies. LLVM supports the Thread Local Storage (TLS) model and can define thread-local storage variables, enabling each thread to have an independent copy of the variable, thereby avoiding data competition problems between threads. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a code compilation method, where the code includes at least one thread-local storage variable, and the code compilation method includes: in response to the existence of a target thread-local storage variable, based on the attributes of at least one target thread-local storage variable, allocating the at least one target thread-local storage variable to thread-local memory to obtain a thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is a thread-local storage variable used by a first function among the at least one thread-local storage variable; updating an intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; and generating machine code corresponding to the code based on the updated intermediate representation.

[0004] The code compilation method provided by at least one embodiment of the present disclosure further includes: inserting an instruction for initializing a target thread-local storage variable used by each first function in the intermediate representation corresponding to the code.

[0005] The code compilation method provided by at least one embodiment of the present disclosure further includes: for each target thread-local storage variable, determining a first function that uses the target thread-local storage variable to construct a variable-function mapping relationship.

[0006] In the code compilation method provided by at least one embodiment of the present disclosure, the first function directly uses the target thread-local storage variable or uses the target thread-local storage variable by calling a second function.

[0007] In the code compilation method provided by at least one embodiment of the present disclosure, the first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0008] In the code compilation method provided by at least one embodiment of the present disclosure, allocating the at least one target thread-local storage variable to thread-local memory based on the attributes of the at least one target thread-local storage variable includes: determining a variable allocation order based on the attributes of the at least one target thread-local storage variable; and sequentially allocating the at least one target thread-local storage variable to the thread-local memory according to the variable allocation order.

[0009] In the code compilation method provided by at least one embodiment of the present disclosure, the attributes of the target thread-local storage variable include at least one of the number of first functions corresponding to the target thread-local storage variable, the alignment parameter of the target thread-local storage variable, the size of the target thread-local storage variable, and the definition order of the target thread-local storage variable.

[0010] In the code compilation method provided by at least one embodiment of the present disclosure, determining the variable allocation order based on the attributes of the at least one target thread-local storage variable includes: arranging the at least one target thread-local storage variable in descending order according to the number of first functions corresponding to the at least one target thread-local storage variable; arranging the at least one target thread-local storage variable in descending order according to the alignment parameter of the at least one target thread-local storage variable; arranging the at least one target thread-local storage variable in ascending order according to the size of the at least one target thread-local storage variable; or arranging the at least one target thread-local storage variable according to the definition order of the at least one target thread-local storage variable.

[0011] In the code compilation method provided by at least one embodiment of the present disclosure, determining the variable allocation order based on the attributes of the at least one target thread-local storage variable includes: arranging the at least one target thread-local storage variable in descending order according to the number of first functions corresponding to the at least one target thread-local storage variable; in response to there being multiple first thread-local storage variables with the same number of first functions among the at least one target thread-local storage variable, arranging the multiple first thread-local storage variables in descending order according to the alignment parameter of the target thread-local storage variable; in response to there being multiple second thread-local storage variables with the same alignment parameter among the multiple first thread-local storage variables, arranging the multiple second thread-local storage variables in ascending order according to the size of the target thread-local storage variable; in response to there being multiple third thread-local storage variables with the same size among the multiple second thread-local storage variables, arranging the multiple third thread-local storage variables according to the definition order of the target thread-local storage variable.

[0012] In the code compilation method provided by at least one embodiment of the present disclosure, for allocating the at least one target thread-local storage variable to the thread-local memory based on the attributes of the at least one target thread-local storage variable, it further includes: for the target thread-local storage variables corresponding to multiple first functions, using the maximum value among the allocated offsets corresponding to the multiple first functions as the allocation starting point of the target thread-local storage variables in the thread-local memory.

[0013] In the code compilation method provided by at least one embodiment of the present disclosure, for updating the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable, it includes: for each target thread-local storage variable, replacing the instruction for obtaining the address of the target thread-local storage variable with the thread-local memory address corresponding to the target thread-local storage variable, where the intermediate representation corresponding to the code includes the instruction for obtaining the address of the target thread-local storage variable.

[0014] The code compilation method provided by at least one embodiment of the present disclosure further includes: obtaining a record of the thread-local memory sizes required by all the target thread-local storage variables used by each first function in the code, so as to pre-allocate a fixed thread-local memory space for each first function according to the record.

[0015] In the code compilation method provided by at least one embodiment of the present disclosure, the life cycle of the pre-allocated thread-local memory space is consistent with the execution cycle of the corresponding first function.

[0016] The present disclosure provides a code compilation device in at least one embodiment, where the code includes at least one thread-local storage variable, and the code compilation device includes: an allocation module configured to, in response to the existence of target thread-local storage variables, allocate the at least one target thread-local storage variable to the thread-local memory based on the attributes of the at least one target thread-local storage variable, to obtain the thread-local memory address corresponding to each target thread-local storage variable, where the target thread-local storage variable is the thread-local storage variable used by a first function among the at least one thread-local storage variable; an update module configured to update the intermediate representation corresponding to the code based on the thread-local memory address corresponding to each target thread-local storage variable; and a generation module configured to generate the machine code corresponding to the code based on the updated intermediate representation.

[0017] At least one embodiment of the present disclosure provides an electronic device, which includes: at least one processor; at least one memory, including one or more computer program modules; wherein, the one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the code compilation method provided by at least one embodiment of the present disclosure.

[0018] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, on which computer instructions are stored, wherein the computer instructions, when executed by at least one processor, execute the code compilation method provided by at least one embodiment of the present disclosure.

[0019] The code compilation method, code compilation device, electronic device and non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure provide support for the thread local storage model in the parallel computing architecture, and solve the problem that the traditional thread local storage implementation cannot be directly applied to the GPU architecture. Developers can directly use high-level syntax for programming without manually managing thread data. The programming model is simple and easy to operate, greatly improving the development efficiency. Moreover, through compilation optimization, the corresponding thread local memory address can be directly fixed during compilation, without loading the thread local memory address corresponding to each thread local storage variable from the register, nor dynamically obtaining the thread local memory address at runtime, achieving better performance in obtaining the thread local memory address. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0021] Figure 1 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0022] Figure 2 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0023] Figure 3 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0024] Figure 4A It is an exemplary diagram of a code compilation method provided by at least one embodiment of the present disclosure;

[0025] Figure 4B It is an exemplary diagram of a code compilation method provided by at least one embodiment of the present disclosure;

[0026] Figure 5 A flowchart of a code compilation method provided for at least one embodiment of the present disclosure;

[0027] Figure 6 A schematic block diagram of a code compiling device provided for at least one embodiment of the present disclosure;

[0028] Figure 7 A schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure;

[0029] Figure 8 A schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure; and

[0030] Figure 9 A schematic block diagram of a non-transitory computer-readable storage medium is provided for at least one embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0032] Flowcharts are used in this disclosure to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. At the same time, other operations may also be added to these processes, or one or more operations may be removed from these processes.

[0033] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0034] The present disclosure is described below through several specific embodiments. In order to keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components (elements) may be omitted. When any component (component) of the embodiments of the present disclosure appears in more than one drawing, the component (component) is represented by the same or similar reference numeral in each drawing.

[0035] LLVM is an open source compiler framework that is designed as a modular, reusable collection of compiler and toolchain technologies. LLVM includes the front end, middle end, and back end.

[0036] The front end of LLVM is responsible for converting source code of different programming languages ​​or programming models into a common intermediate representation (IR). For example, Clang is the compiler front end of LLVM that natively supports C, C++, and Objective-C. When a user declares a variable using thread_local (C++11) or __thread (GCC extension), Clang can mark the variable as a thread local storage variable (TLS variable), for example, it can be marked in the generated intermediate representation through the thread_local attribute and a specific link type (such as dso_local).

[0037] The middle end of LLVM is responsible for optimizing the intermediate representation generated by the front end to improve the performance of the final generated code. For example, it can simplify the thread local storage access mode and merge multiple accesses into one address calculation. For example, it can also implicitly select the thread local storage model based on the scope of the variable, or explicitly specify the thread local storage model in the code.

[0038] The backend of LLVM is responsible for converting the optimized intermediate representation into machine code or assembly code according to the characteristics of the target hardware platform. The specific tasks of the backend may include instruction selection, register allocation, instruction scheduling, code generation, etc. For example, instruction selection refers to mapping the intermediate representation to the instruction set of the target hardware platform; register allocation refers to mapping the virtual registers in the intermediate representation to the actual registers of the target hardware; instruction scheduling refers to rearranging the order of instructions to maximize the utilization of hardware resources; code generation refers to generating machine code or assembly code corresponding to the target hardware platform.

[0039] LLVM supports the thread local storage model. Through the thread_local keyword, you can define thread local storage variables so that each thread has an independent copy of the variable, thus avoiding data competition problems between threads. In other words, thread local storage variables are not shared between threads, which means that even if multiple threads access the same variable name, they actually operate on their own independent memory areas.

[0040] However, the inventors of the present disclosure note that the implementation of thread local storage depends on the support of the target platform. Currently, in traditional central processing unit (CPU) architectures (such as x86 and ARM), operating systems and application binary interfaces (ABI) provide support for thread local storage, for example, thread local storage implementation can be supported through specific segments in executable and linkable format (ELF) files (such as .tdata segments for storing uninitialized thread local storage variables and .tbss segments for storing initialized thread local storage variables). However, there is no similar mechanism to support thread local storage implementation in graphics processing units (GPUs), general-purpose graphics processing units (GPGPUs) or other parallel computing architecture-related backends (such as CUDA and OpenCL), which brings challenges to the implementation of application scenarios that require large-scale parallel processing such as high-performance computing and machine learning.

[0041] The parallel thread model of the GPU (such as CUDA and OpenCL) usually requires each thread to access its private data independently, but the traditional thread local storage implementation method (such as the CPU's __thread keyword) cannot be directly mapped to the GPU architecture, and it is difficult to meet the requirements of creating thread private data and supporting related data management and optimization. In addition, the GPU lacks hardware-level thread local storage support, which requires developers to manually manage thread local data (such as calculating offsets through thread indexes), increasing development complexity and making it prone to errors. In existing GPU programming practices, developers need to explicitly pass thread indexes (such as threadIdx.x in CUDA) and manually allocate global memory to simulate the storage of thread private data, which leads to code redundancy and performance loss. In addition, the existing LLVM compiler does not provide sufficient support for GPU thread local storage and cannot automatically generate efficient memory allocation and access code.

[0042] The inventors of the present disclosure also noticed that the thread local storage model supported by the current mainstream CPU architecture has performance limitations. For example, the static model has instruction and register overhead when accessing thread local storage variables, and the dynamic model has runtime parsing overhead when accessing thread local storage variables. The specific analysis is as follows:

[0043] (1) Static model

[0044] In the CPU architecture, the static model can access thread local storage variables by using segment registers (such as Linux's FS segment register, Windows' GS segment register) plus offsets. An example can refer to the following assembly code:

[0045] movq %fs:0, %rax ; / / LINE 1

[0046] addq $TLS_var@tpoff, %rax; / / LINE 2

[0047] Code line LINE 1 indicates loading the TLS base address from the address pointed to by the FS segment register to the RAX register, and code line LINE 2 indicates adding the offset (@tpoff) calculated at compile time to the base address. The actual memory address of the TLS_var variable of the current thread is stored in register RAX. During static linking, the linker determines the final offset of the TLS_var variable (for example, 0x10) and replaces @tpoff with the actual value of the offset. The above method requires executing movq to load the base address each time a thread-local storage variable is accessed, and the RAX register needs to be occupied to temporarily store the base address, which increases register pressure, so there is a large instruction overhead and register overhead.

[0048] (2) Dynamic model

[0049] In the CPU architecture, the dynamic model obtains the address of the thread local storage variable by calling __tls_get_addr() or dynamically calculating the offset at runtime, which may bring function call overhead or memory access overhead and low performance.

[0050] At least one embodiment of the present disclosure provides a code compilation method for compiling code including at least one thread local storage variable, the code compilation method comprising: in response to the existence of a target thread local storage variable, based on the attribute of the at least one target thread local storage variable, allocating the at least one target thread local storage variable to a thread local memory, obtaining a thread local memory address corresponding to each target thread local storage variable, wherein the target thread local storage variable is a thread local storage variable used by a first function in the at least one thread local storage variable; updating an intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread local storage variable; and generating a machine code corresponding to the code based on the updated intermediate representation.

[0051] The code compiling method provided by at least one embodiment of the present disclosure may be executed by a compiler. For example, the compiler may be located at a host end, and the code compiling method may be executed by the host end.

[0052] In the code compilation method provided in at least one embodiment of the present disclosure, thread local storage model support under the parallel computing architecture is provided, which solves the problem that the traditional thread local storage implementation cannot be directly applied to the GPU architecture. Through the above code compilation method, developers can directly use high-level syntax (such as __thread) for programming, without manually managing thread data. The programming model is simple and easy to operate, which greatly improves development efficiency. In addition, the code compilation method directly fixes the corresponding thread local memory address during compilation through compilation optimization, without loading the thread local memory address corresponding to each thread local storage variable from the register, and without dynamically obtaining the thread local memory address at runtime, thereby achieving better performance in obtaining the thread local memory address.

[0053] Figure 1 A flowchart of a code compilation method provided in at least one embodiment of the present disclosure is used to compile code including at least one thread local storage variable. Figure 1 As shown, the code compilation method provided by at least one embodiment of the present disclosure includes the following steps S101-S103.

[0054] Step S101: In response to the existence of a target thread local storage variable, at least one target thread local storage variable is allocated to a thread local memory based on an attribute of the at least one target thread local storage variable, and a thread local memory address corresponding to each target thread local storage variable is obtained.

[0055] For example, in step S101, the target thread local storage variable refers to the thread local storage variable used by the first function among the at least one thread local storage variable. That is, the "at least one thread local storage variable" mentioned above may include thread local storage variables used by the first function, and may also include thread local storage variables not used by any first function. Step S101 allocates thread local memory space for the used thread local storage variables, while unused thread local storage variables will not be allocated to the thread local memory.

[0056] In the code compilation method provided in at least one embodiment of the present disclosure, the code may include multiple first functions and multiple second functions. The first function is a function called by the host end and executed by the device end, and the second function is a function called by the device end and executed by the device end.

[0057] For example, in at least one embodiment of the present disclosure, the host side may include a central processing unit (CPU), and the device side may include a graphics processing unit (GPU), a general purpose graphics processing unit (GPGPU), and the like.

[0058] For example, in at least one embodiment of the present disclosure, the first function may be a kernel function, and the second function may be a device function. A kernel function refers to a function called by a host and executed by a device, and a device function refers to a function called by a device and executed by a device. For example, a kernel function may be called by a host function (such as a host function) running in a CPU, while a device function may be called by a kernel function or other device function executed at a GPU. The code usually includes multiple kernel functions and device functions, each kernel function may call multiple device functions, and each device function may also be called by multiple kernel functions.

[0059] “Thread local storage variables used by the first function” refers to thread local storage variables directly used by the first function and thread local storage variables used by the first function by calling the second function, and the above thread local storage variables are all target thread local storage variables.

[0060] Thread Local Memory (TLM) refers to a memory area private to each thread, which is mainly used to store thread-private data and cannot be directly accessed by other threads. For example, thread local memory can be thread local memory in a GPU. By saving the target thread local storage variables in thread local memory, the burden of registers in the GPU can be reduced. For example, thread local memory can also be thread local memory in other parallel computing architectures, which is not limited in the embodiments of the present disclosure.

[0061] It should be noted that thread local memory and the thread local storage introduced above are two different concepts. Thread local memory refers to a memory area, while thread local storage refers to a way of storing variables.

[0062] For example, in step S101, the target thread local storage variables may be allocated to the thread local memory based on the attributes of the target thread local storage variables, so that the thread local memory address corresponding to each target thread local storage variable may be obtained. The attributes of the target thread local storage variables may include, for example, variable type, variable size, alignment requirements, etc., which are not limited in the embodiments of the present disclosure.

[0063] Step S102: updating the intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread local storage variable.

[0064] For example, LLVM's Clang front end can automatically generate an intermediate representation for the code. In LLVM, you can use the TLSSupported member variable of the TargetInfo class to control whether the thread local storage feature is enabled. Setting the value of the TLSSupported member variable to true enables the thread local storage feature support of the LLVM front end.

[0065] For example, suppose a user writes the following code:

[0066] __device__ __thread int a=1; / / LINE1

[0067] __global__ void kernel() {

[0068] c[1]=2; / / LINE3

[0069] }

[0070] The above code line LINE1 is the declaration definition of the thread local storage variable a. In some examples, according to the code line LINE1, the intermediate representation automatically generated by the Clang front end of LLVM is as follows:

[0071] @a = thread_local addrspace(1) externally_initialized global i32 1,align 4

[0072] In the above intermediate representation, @a indicates that variable a is defined as a global variable, thread_local indicates that variable a is a thread local storage variable, addrspace(1) specifies that the address space of the variable is global, externally_initialized indicates that the variable is initialized externally (for example, at runtime, not during compilation), global i32 1 indicates that a global variable is defined, the type is a 32-bit integer, the initial value is 1, and align 4 indicates that the variable is aligned to 4 bytes.

[0073] The above code line LINE3 implements the assignment operation of the thread local storage variable c[1]. In some examples, according to the code line LINE3, the intermediate representation automatically generated by the Clang front end of LLVM is as follows:

[0074] %1 = tail call ptr @llvm.threadlocal.address.p0(ptr addrspacecast(ptr addrspace(1) @c to ptr))

[0075] arrayidx1 = getelementptr inbounds [2 x float], ptr %1, i64 0, i64 1

[0076] store float 2.000000e+00, ptr %arrayidx1, align 4

[0077] In the above intermediate representation, the llvm.threadlocal.address.p0 instruction is used to obtain the address of the thread local storage variable c, %1 stores the address of the thread local storage variable c, %arrayidx1 stores the address of the thread local storage variable c[1], and the store instruction stores the floating point number 2.0 into the storage location pointed to by %arrayidx1, which is the storage location corresponding to c[1].

[0078] In step S102, the intermediate representation automatically generated by the Clang front end may be updated based on the result of the variable allocation in step S101 to obtain an optimized intermediate representation.

[0079] In the code compilation method provided by at least one embodiment of the present disclosure, the intermediate representation corresponding to the code includes instructions for obtaining the address of the target thread local storage variable. An example of step S102 may be: for each target thread local storage variable, the instruction for obtaining the address of the target thread local storage variable is replaced with the thread local memory address corresponding to the target thread local storage variable.

[0080] For example, the instruction for obtaining the address of the target thread local storage variable is llvm.threadlocal.address.p0, and the calling function of each target thread local storage variable is traversed, and all llvm.threadlocal.address.p0 instructions are directly replaced with the thread local memory address corresponding to the target thread local storage variable to obtain an optimized intermediate representation.

[0081] In some examples, the intermediate representation before optimization (that is, the intermediate representation automatically generated by the Clang front end) is as follows:

[0082] %0 = tail call ptr @llvm.threadlocal.address.p0(ptr addrspacecast(ptr addrspace(1) @th to ptr))

[0083] store i32 66, ptr %0, align 4

[0084] The corresponding optimized intermediate representation is as follows:

[0085] store i32 66, ptr addrspacecast (ptr addrspace(4) inttoptr (i32 24 toptr addrspace(4)) to ptr), align 4

[0086] In the optimized intermediate representation, the llvm.threadlocal.address.p0 instruction is no longer needed, where 24 represents the thread local memory address of the target thread local storage variable th allocated by step S101, addrspace(4) specifies the address space of the variable as local, and addrspacecast represents the address space conversion to generate the correct thread local memory read and write instructions that meet the requirements of the target architecture.

[0087] It can be seen that the optimized intermediate representation fixes the corresponding thread local memory address during compilation. There is no need to load the thread local memory address corresponding to each target thread local storage variable from the register, nor is there any need to dynamically obtain the thread local memory address at runtime, thus achieving better performance in obtaining the thread local memory address.

[0088] Step S103: Generate machine code corresponding to the code based on the updated intermediate representation.

[0089] For example, in step S103, the optimized intermediate representation obtained in step S102 can be converted into machine code suitable for the target hardware platform through the steps of instruction selection, register allocation, instruction scheduling, code generation, etc. The optimization of the intermediate representation in step S102 helps to reduce the workload of instruction selection and improve performance.

[0090] Specifically, in the CPU architecture, instruction selection for thread local storage variable access instructions (such as ISD::GlobalTLSAddress) is performed in the target lowering stage. The target lowering stage is mainly responsible for converting the intermediate representation into a low-level representation related to the target hardware platform. Specifically, in the process of converting the intermediate representation into machine code, access to TLS variables is achieved by using segment registers plus offsets, and corresponding read and write operations are completed. The code compilation method provided by at least one embodiment of the present disclosure replaces the instruction for obtaining the address of the target thread local storage variable with the thread local memory address corresponding to the target thread local storage variable before instruction selection, effectively reducing the workload of instruction selection.

[0091] The code compilation method provided by at least one embodiment of the present disclosure further includes the following step S104. Step S104 may be performed before step S101.

[0092] Step S104: for each target thread local storage variable, determine a first function that uses the target thread local storage variable to construct a variable-function mapping relationship.

[0093] In step S104, the transfer path of the variable in the function can be tracked according to the call graph (also called the call relationship graph), so as to determine which functions use the local storage variables of the target thread. The call graph is a directed graph that can represent the call relationship between functions in the program. The compiler can generate the call graph by, for example, static analysis. For example, the call relationship between functions can be parsed by the call graph builder (Call Graph Builder) integrated in the compiler, and the call graph can be generated. For example, the variable-function mapping relationship records which first functions use the local storage variables of each target thread, and the variable-function mapping relationship can be presented in the form of a mapping graph.

[0094] It should be noted that, since some thread local storage variables in the code may only be defined but not used by any first function, not every thread local storage variable can find a corresponding first function in the variable-function mapping relationship.

[0095] In step S104, the "first function using the target thread local storage variable" includes the first function using the target thread local storage variable directly or indirectly. For example, "direct use" means that the first function directly uses the target thread local storage variable, and "indirect use" means that the first function uses the target thread local storage variable by calling the second function.

[0096] Correspondingly, an example of step S104 may be: for each target thread local storage variable, determining a kernel function that directly or indirectly uses the target thread local storage variable to construct a variable-function mapping relationship.

[0097] Figure 2 A flowchart of a code compiling method provided for at least one embodiment of the present disclosure.

[0098] like Figure 2 As shown, in the code compilation method provided by at least one embodiment of the present disclosure, an example of step S101 may include the following steps S201 to S202.

[0099] Step S201: Determine a variable allocation order based on the attributes of at least one target thread local storage variable.

[0100] For example, in step S201, multiple target thread local storage variables may be sorted according to the attributes of the target thread local storage variables.

[0101] Step S202: allocating at least one target thread local storage variable to the thread local memory in sequence according to the variable allocation order.

[0102] For example, in step S202, according to the variable allocation order determined in step S201, multiple target thread local storage variables may be allocated to corresponding positions in the thread local memory in sequence.

[0103] In the code compilation method provided in at least one embodiment of the present disclosure, an example of step S101 may further include the following step S203.

[0104] Step S203: For target thread local storage variables corresponding to multiple first functions, use the maximum value of the allocated offsets corresponding to the multiple first functions as the allocation starting point of the target thread local storage variable in the thread local memory.

[0105] For example, in step S203, the target thread local storage variable corresponding to the multiple first functions means that the target thread local storage variable is used directly or indirectly by the multiple first functions. In some examples, direct use may be, for example, a kernel function directly using the target thread local storage variable, and indirect use may be, for example, a kernel function using the target thread local storage variable by calling a device function.

[0106] For example, suppose that thread local storage variable b is only used by kernel function Kernel 1 and kernel function Kernel 2 (thread local storage variable b is the target thread local storage variable), and the allocated offset corresponding to kernel function Kernel 1 is 4 bytes, and the allocated offset corresponding to kernel function Kernel 2 is 0 bytes. It can be seen that the maximum value of the allocated offset is 4 bytes, and it is used as the allocation starting point of thread local storage variable b in thread local memory, that is, the allocation starting address of thread local storage variable b in thread local memory is 4, not 0.

[0107] It should be noted that in the code compilation method provided in at least one embodiment of the present disclosure, an independent thread-local memory space is allocated for each kernel function, and the "allocation starting point in the thread-local memory" mentioned above refers to the offset relative to the starting address of each thread-local memory space. For example, kernel function Kernel 1 corresponds to a thread-local memory space T1, and kernel function Kernel 2 corresponds to a thread-local memory space T2. In order to simplify the program logic, for each kernel function, the starting address of its corresponding thread-local memory space is regarded as 0, that is, the starting addresses of T1 and T2 are both regarded as 0, and can be mapped to different physical addresses through hardware processing later. Therefore, in the above example, the storage location of the thread-local storage variable b in the thread-local memory space T1 and T2 is offset 4 bytes backward from the starting address 0.

[0108] Since multiple kernel functions and multiple device functions may use the same thread local storage variable, and the device function cannot determine which kernel functions will call it, it is required that the addresses assigned to the same thread local storage variables in all kernel functions and device functions are consistent, that is, no matter which kernel function or device function can find the thread local storage variable in the thread local memory through the same address. Through the method provided in step S203, the address allocation consistency of the same thread local storage variable required above can be achieved, that is, no matter whether it is a kernel function or a device function, it can find the thread local storage variable b in the thread local memory through address 4.

[0109] It should be noted that, for a target thread local storage variable that is only used by a single first function, it is not necessary to determine the allocation starting point in the thread local memory according to step S203, and it only needs to be allocated in sequence.

[0110] In the code compilation method provided by at least one embodiment of the present disclosure, the attributes of the target thread local storage variable may include: at least one of the first function number corresponding to the target thread local storage variable, the alignment parameter of the target thread local storage variable, the size of the target thread local storage variable and the definition order of the target thread local storage variable.

[0111] For example, the number of first functions corresponding to the target thread local storage variable can be understood as the number of first functions using the target thread local storage variable. For example, assuming that there are n kernel functions that directly or indirectly use the target thread local storage variable, the number of first functions corresponding to the target thread local storage variable is n.

[0112] For example, the alignment parameter of the target thread local storage variable is a parameter used to explicitly specify the alignment of the variable in storage, usually in bytes, such as 4 bytes, 8 bytes, 16 bytes, etc. For example, when the alignment parameter of the target thread local storage variable is 4 bytes, it indicates that the target thread local storage variable needs to be aligned according to 4 bytes when stored.

[0113] For example, the size of the target thread local storage variable can be understood as the size of the storage space that the target thread local storage variable needs to occupy, which is related to the type of the variable and is usually in bytes. For example, the size of the target thread local storage variable of type int is 4 bytes. It should be noted that the size of the target thread local storage variable of the same type may be different in different systems or compilers.

[0114] For example, the definition order of the target thread local storage variables refers to the definition order of each target thread local storage variable in the code, which is determined by the user when writing the code.

[0115] For example, when determining the variable allocation order, one or more attributes of the first number of functions corresponding to the target thread local storage variable, the alignment parameter of the target thread local storage variable, the size of the target thread local storage variable, and the definition order of the target thread local storage variable may be considered. For example, in some examples, the variable allocation order may be determined based on the first number of functions corresponding to at least one target thread local storage variable, the alignment parameter of at least one target thread local storage variable, the size of at least one target thread local storage variable, or the definition order of at least one target thread local storage variable. For example, in other examples, the variable allocation order may be determined based on the first number of functions corresponding to at least one target thread local storage variable, the alignment parameter of at least one target thread local storage variable, the size of at least one target thread local storage variable, and the definition order of at least one target thread local storage variable. For example, two or three of the above four attributes may be selected as the basis for determining the variable allocation order.

[0116] It should be noted that the above three attributes are only examples, and the variable allocation order may also be determined according to other attributes according to actual needs, and the embodiments of the present disclosure are not limited to this.

[0117] In the code compilation method provided in at least one embodiment of the present disclosure, an example of step S201 may include any one of the following steps S301 to S304. For example, if only one attribute is considered when determining the variable allocation order, the corresponding step may be selected from steps S301 to S304 for execution according to the attribute.

[0118] Step S301: Arrange at least one target thread local storage variable in descending order according to the number of first functions corresponding to the at least one target thread local storage variable.

[0119] Step S302: Arrange at least one target thread local storage variable in descending order according to the alignment parameter of at least one target thread local storage variable.

[0120] Step S303: Arrange at least one target thread local storage variable in ascending order according to the size of the at least one target thread local storage variable.

[0121] Step S304: Arrange at least one target thread local storage variable according to the definition order of at least one target thread local storage variable.

[0122] For example, in step S301, the target thread local storage variables may be sorted from large to small according to the first function quantity corresponding to the target thread local storage variables. For example, assuming that there are three target thread local storage variables a, b, and c to be sorted, where the first function quantity corresponding to the target thread local storage variable a is 4, the first function quantity corresponding to the target thread local storage variable b is 5, and the first function quantity corresponding to the target thread local storage variable c is 2, then the sorting result is b, a, c. For example, in some examples, if there are multiple target thread local storage variables with the same corresponding first function quantity, they may be further sorted according to the definition order of the variables.

[0123] For example, in step S302, the target thread local storage variables may be sorted from large to small according to the alignment parameters of the target thread local storage variables. For example, assuming that there are three target thread local storage variables a, b, and c to be sorted, where the alignment parameter of the target thread local storage variable a is 4 bytes, the alignment parameter of the target thread local storage variable b is 8 bytes, and the alignment parameter of the target thread local storage variable c is 16 bytes, then the sorting result is c, b, a. For example, in some examples, if there are multiple target thread local storage variables with the same alignment parameters, they may be further sorted according to the definition order of the variables.

[0124] For example, in step S303, the target thread local storage variables may be sorted from small to large according to their sizes. For example, assuming that there are three target thread local storage variables a, b, and c to be sorted, where the size of the target thread local storage variable a is 8 bytes, the size of the target thread local storage variable b is 1 byte, and the size of the target thread local storage variable c is 4 bytes, then the sorting result is b, c, a. For example, in some examples, if there are multiple target thread local storage variables of the same size, they may be further sorted according to the definition order of the variables.

[0125] For example, in step S304, the target thread local storage variables may be sorted according to the definition order of the target thread local storage variables. For example, assuming there are three target thread local storage variables a, b, c to be sorted, and the definition order is a, b, c, then the sorting result is a, b, c.

[0126] It should be noted that the above sorting method is to make the allocated thread local memory gap as small as possible and improve the utilization rate of the thread local memory. According to actual needs, the sorting can also be performed in other ways, which is not limited in the embodiment of the present disclosure.

[0127] Figure 3 A flowchart of a code compiling method provided for at least one embodiment of the present disclosure.

[0128] like Figure 3 As shown, in the code compilation method provided by at least one embodiment of the present disclosure, if all four attributes mentioned above are considered when determining the variable allocation order, an example of the above step S201 may include the following steps S401 to S404.

[0129] Step S401: Arrange at least one target thread local storage variable in descending order according to the number of first functions corresponding to the at least one target thread local storage variable.

[0130] Step S402: In response to the existence of a plurality of first thread local storage variables having the same number of corresponding first functions in at least one target thread local storage variable, the plurality of first thread local storage variables are arranged in descending order according to alignment parameters of the target thread local storage variables.

[0131] Step S403: In response to the presence of a plurality of second thread local storage variables having the same alignment parameter among the plurality of first thread local storage variables, the plurality of second thread local storage variables are arranged in ascending order according to the size of the target thread local storage variable.

[0132] Step S404: in response to the presence of a plurality of third thread local storage variables of the same size in the plurality of second thread local storage variables, the plurality of third thread local storage variables are arranged according to the definition order of the target thread local storage variable.

[0133] For example, in step S401, the target thread local storage variables may be sorted from large to small according to the number of first functions corresponding to the target thread local storage variables.

[0134] For example, in step S402, according to the sorting result of step S401, if it is determined that there are multiple first thread local storage variables with the same number of first functions in the target thread local storage variables, the multiple first thread local storage variables are arranged from large to small according to the alignment parameters of the target thread local storage variables. It should be noted that step S402 only processes the variables whose order has not been determined in step S401 (i.e., the first thread local storage variables), and the variables whose order has been determined are not considered.

[0135] For example, in step S403, according to the sorting result of step S402, if it is determined that there are multiple second thread local storage variables with the same alignment parameters among the multiple first thread local storage variables, the multiple second thread local storage variables are arranged from small to large according to the size of the target thread local storage variable. It should be noted that step S403 only processes the variables (i.e., the second thread local storage variables) whose order has not been determined in step S402, and the variables whose order has been determined are not considered.

[0136] For example, in step S404, according to the sorting result of step S403, if it is determined that there are multiple third thread local storage variables of the same size among the multiple second thread local storage variables, the multiple third thread local storage variables are arranged according to the definition order of the target thread local storage variable to obtain the final sorting result. The definition order here refers to the order in which the user defines the variables when writing the code. It should be noted that step S404 only processes the variables (i.e., the third thread local storage variables) whose order has not been determined in step S403, and the variables whose order has been determined are not considered.

[0137] It should be noted that the above steps S401 to S404 are only an example, which provides a property priority of: 1. The number of the first function corresponding to the target thread local storage variable; 2. The alignment parameter of the target thread local storage variable; 3. The size of the target thread local storage variable; 4. The definition order of the target thread local storage variable. This property priority is set to make the allocated thread local memory gap as small as possible and improve the utilization rate of the thread local memory. According to actual needs, it can also be sorted according to other property priorities, for example, it can be sorted according to size first, and then further sorted according to alignment parameters, etc., which is not limited by the embodiments of the present disclosure.

[0138] The following is a specific example of the above steps S401 to S404.

[0139] For example, assuming that there are six thread local storage variables a, b, c, d, e, and f to be sorted (all are target thread local storage variables), the number of first functions corresponding to the thread local storage variable a is 4, the number of first functions corresponding to the thread local storage variable b is 5, and the number of first functions corresponding to the thread local storage variables c, d, e, and f are all 2, and step S401 is executed, then the sorting result is b, a, (c, d, e, f), where (c, d, e, f) means that the relative order between variables c, d, e, and f has not yet been determined, and currently it can only be confirmed that these four variables are all located after variable a.

[0140] Since the variables c, d, e, and f have the same number of first functions (i.e., c, d, e, and f are all first thread local storage variables), it is necessary to execute step S402 to determine the relative order between the four variables according to the alignment parameter. For example, further assuming that the alignment parameters of the thread local storage variables c, d, and e are all 4 bytes, and the alignment parameter of the thread local storage variable f is 8 bytes, and executing step S402, the further sorting result is b, a, f, (c, d, e), where (c, d, e) means that the relative order between the variables c, d, and e has not yet been determined, and currently it can only be confirmed that these three variables are all located after the variable f.

[0141] Since the alignment parameters of variables c, d, and e are the same (that is, c, d, and e are all second thread local storage variables), step S403 needs to be executed to determine the relative order of the three variables according to their sizes. For example, further assuming that the size of thread local storage variable c is 2 bytes, and the sizes of thread local storage variables d and e are both 1 byte, step S403 is executed, and the further sorting result is b, a, f, (d, e), c, where (d, e) means that the relative order between variables d and e has not yet been determined, and currently it can only be confirmed that these two variables are both located after variable f and before variable c.

[0142] Since the variables d and e have the same size (i.e., both d and e are third thread local storage variables), step S404 needs to be executed to determine the relative order between the two variables according to the definition order. For example, further assuming that the thread local storage variable d is defined first and the thread local storage variable e is defined later, and step S404 is executed, the further sorting result is b, a, f, d, e, c, which is the final sorting result.

[0143] Figure 4A An exemplary schematic diagram of a code compilation method provided for at least one embodiment of the present disclosure.

[0144] For example, Figure 4A As shown in FIG. 1 , the code includes three kernel functions and six device functions. The three kernel functions are represented by, for example, Kernel 1 to Kernel 3, and the six device functions are represented by, for example, Func 1 to Func 6. Figure 4A As shown in the call graph on the right, kernel function Kernel 1 calls device functions Func 1 and Func 2, kernel function Kernel 2 calls device functions Func3 and Func4, and kernel function kernel3 calls device functions Func5 and Func6. Among them, device function Func1 uses thread local storage variable a, device function Func2 uses thread local storage variable b, device function Func3 uses thread local storage variable b, device function Func4 uses thread local storage variable c, device function Func5 uses thread local storage variable c, and device function Func6 uses thread local storage variable a. The alignment parameters of thread local storage variables a, b, and c are all 4 bytes, and the sizes are also 4 bytes. The above thread local storage variables a, b, and c are all target thread local storage variables.

[0145] Execute the above steps S401 to S404 to sort the thread local storage variables a, b, and c. It can be seen from the above that the number of first functions corresponding to the thread local storage variable a is 2 (kernel functions Kernel 1 and Kernel 3), the number of first functions corresponding to the thread local storage variable b is 2 (kernel functions Kernel 1 and Kernel 2), and the number of first functions corresponding to the thread local storage variable c is 2 (kernel functions Kernel 2 and Kernel 3). Therefore, the number of first functions, alignment parameters, and sizes of the thread local storage variables a, b, and c are the same, and the final sorting result cannot be obtained according to steps S401 to S403. Execute step S404 to sort the thread local storage variables a, b, and c in the order of definition, and the final sorting results are a, b, and c.

[0146] First, the thread local storage variable a is allocated. The thread local storage variable a is used by kernel functions Kernel 1 and Kernel 3, and the allocated offsets corresponding to kernel functions Kernel 1 and Kernel 3 are both 0 bytes, that is, the maximum value of the allocated offset is 0 bytes. Therefore, the allocation start address of the thread local storage variable a in the thread local memory is 0, and 4 bytes of thread local memory space are allocated.

[0147] Next, the thread local storage variable b is allocated. The thread local storage variable b is used by kernel function Kernel 1 and kernel function Kernel 2, and the allocated offset corresponding to kernel function Kernel 1 is 4 bytes, and the allocated offset corresponding to kernel function Kernel 2 is 0 bytes, that is, the maximum value of the allocated offset is 4 bytes. Therefore, the allocation start address of the thread local storage variable b in the thread local memory is 4, and 4 bytes of thread local memory space are allocated.

[0148] Finally, the thread local storage variable c is allocated. The thread local storage variable c is used by kernel functions Kernel 2 and Kernel 3, and the allocated offset corresponding to kernel function Kernel 2 is 8 bytes, and the allocated offset corresponding to kernel function Kernel 3 is 4 bytes, that is, the maximum value of the allocated offset is 8 bytes. Therefore, the allocation start address of the thread local storage variable c in the thread local memory is 8, and 4 bytes of thread local memory space are allocated.

[0149] Figure 4B An exemplary schematic diagram of a code compilation method provided for at least one embodiment of the present disclosure.

[0150] For example, Figure 4B As shown in FIG. 1 , the code includes two kernel functions and four device functions. The two kernel functions are represented by, for example, Kernel 1 to Kernel 2, and the four device functions are represented by, for example, Func 1 to Func 4. Figure 4B As shown in the call graph on the right, kernel function Kernel 1 calls device functions Func 1 and Func 2, and kernel function Kernel 2 calls device functions Func3 and Func4. Among them, device function Func1 uses thread local storage variable a, device function Func2 uses thread local storage variable b, device function Func3 uses thread local storage variable c, and device function Func4 uses thread local storage variable d. The above thread local storage variables a, b, c, and d are all target thread local storage variables. It can be seen that there are no shared thread local storage variables between kernel functions Kernel 1 and Kernel 2, and they can be compactly allocated in sequence. In this case, there is no gap in the thread local memory, and there is no waste of thread local memory space.

[0151] In the code compilation method provided in at least one embodiment of the present disclosure, an independent thread-local memory space is allocated to each kernel function, so the "allocation start address in the thread-local memory" mentioned above refers to the offset relative to the start address of each thread-local memory space. It should be noted that Figure 4A and Figure 4B In the code, the stack is used to indicate the allocation of the target thread local storage variables. The stack here can be regarded as a virtual concept. In LLVM, when the intermediate representation is converted into machine code, the allocation on the stack is mapped to the actual storage space of the target architecture, that is, according to the location of the target thread local storage variable on the stack, it is mapped to the corresponding thread local memory space.

[0152] It should be noted that thread local memory can also be used to store local variables other than the target thread local storage variables. For example, Figure 4A and Figure 4B As shown, the stack corresponding to each kernel function may also include local variables.

[0153] The compiler can automatically manage and optimize memory allocation. By optimizing the frame layout of the target thread's local storage variables, the stack utilization can be effectively improved and the overall space occupied by the thread's local memory can be reduced.

[0154] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S105. Step S105 may be performed before step S102.

[0155] Step S105: inserting instructions for initializing the target thread local storage variables used by each first function into the intermediate representation corresponding to the code.

[0156] For example, in step S105, the code may include multiple first functions and multiple second functions. The target thread local storage variables directly or indirectly used by each first function can be determined according to the call graph. For each first function, an instruction for initializing the target thread local storage variables directly or indirectly used by the first function can be inserted at the start position of its corresponding intermediate representation.

[0157] For example, taking the target thread local storage variable a as an example, an example of an instruction to initialize the target thread local storage variable a is as follows:

[0158] %0 = call ptr @llvm.threadlocal.address.p0(ptr addrspacecast (ptraddrspace(1) @a to ptr))

[0159] store i32 1, ptr %0, align 4

[0160] The above intermediate representation means that the address of the local storage variable a of the target thread is first obtained through the llvm.threadlocal.address.p0 instruction, and then the initial value of the variable a (1 in this example) is stored in the address through the store instruction. It should be noted that the above initialization instruction is only an example, and different initialization instructions can be inserted according to actual needs.

[0161] It should be noted that the above initialization instruction will also be updated in step S102, that is, the llvm.threadlocal.address.p0 instruction will be replaced with the thread local memory address corresponding to the target thread local storage variable a. Therefore, the initialization instruction introduced here will not bring the memory access overhead of reading and writing addresses from registers.

[0162] In the CPU architecture, thread local storage variables are usually initialized by the runtime environment when the program starts. Different from the above initialization method, the code compilation method provided by at least one embodiment of the present disclosure inserts instructions for initializing the target thread local storage variables during compilation, rather than initializing the target thread local storage variables at runtime, which can effectively avoid refreshing the entire thread local memory at runtime and achieve performance optimization. In addition, this method increases the opportunities for compilation optimization, reduces the coupling between compilation and runtime, and reduces the workload of compilation and runtime.

[0163] For example, in at least one embodiment of the present disclosure, before step S105, the usage of each thread local storage variable can be analyzed for each first function, and for thread local storage variables that are not used, there is no need to insert their corresponding initialization instructions. For example, the above analysis operation can be implemented by a compiler. In this way, unnecessary initialization operations can be reduced and performance can be improved.

[0164] For example, in order to further optimize the performance, redundant instruction optimization may be performed on the intermediate representation corresponding to the code before step S101.

[0165] For example, in some examples, for unused thread local storage variables, instructions for obtaining addresses of thread local storage variables may be removed from the intermediate representation corresponding to the code.

[0166] For example, the intermediate representation automatically generated by the Clang front end of LLVM can be analyzed, that is, the usage of each thread local storage variable is analyzed for each first function, and for unused thread local storage variables, the instructions for obtaining the address of the thread local storage variable are removed from the intermediate representation corresponding to the first function.

[0167] For example, in other examples, repeated instructions for obtaining the address of the thread local storage variable can be removed from the intermediate representation corresponding to the code, so that the address is obtained only once for each thread local storage variable in each first function. For example, if there are multiple identical instructions for obtaining the address of the thread local storage variable in the intermediate representation corresponding to a first function, the repeated instructions need to be removed and only the first instruction is retained.

[0168] A special example is as follows. Assume that c is an array type variable. For the kernel function that uses thread local storage variables c[1] and c[2], it is necessary to remove the instructions in the intermediate representation for obtaining the addresses of thread local storage variables c[1] and c[2], and only keep the instructions for obtaining the base address of thread local storage variable c. The addresses of c[1] and c[2] can be calculated by adding the corresponding offset to the base address of c.

[0169] Through the above method, repeated address acquisition operations can be avoided, the operating efficiency can be significantly improved and the overhead of redundant instructions can be reduced.

[0170] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S106. Step S106 may be performed before step S101.

[0171] Step S106: Obtain records of thread local memory sizes required by all target thread local storage variables used by each first function in the code, so as to pre-allocate a fixed thread local memory space for each first function according to the records.

[0172] Frame Lowering is a stage of the LLVM backend, responsible for converting the abstract stack frame operations in the intermediate representation into machine code suitable for the target hardware platform. Its main task is to generate the prologue and epilogue of the function call, that is, the stack frame management code at the function entry and exit. These instructions are mainly used to set up and restore the stack frame so that the function can run correctly and maintain the integrity of the call chain.

[0173] For example, in the compilation stage, step S101 can be used to determine the thread local memory size required for each kernel function. In step S106, the thread local memory size corresponding to all target thread local storage variables used by each kernel function can be recorded through metadata. It should be noted that the thread local memory is also used to store local variables other than the target thread local storage variables. Here, only the thread local memory size required to be occupied by all target thread local storage variables used by the kernel function is recorded, and the thread local memory size occupied by other local variables is not included. For example, when the compiler generates an executable file corresponding to the first function, it usually embeds a piece of metadata. Metadata can be regarded as additional information, which can be used to guide compiler optimization, debugging, code analysis or runtime behavior. These metadata that record the thread local memory size required by each kernel function can be obtained in the early stage of backend code generation, and a fixed thread local memory space is pre-allocated to each kernel function accordingly. Among them, the life cycle of the pre-allocated thread local memory space is consistent with the execution cycle of the corresponding kernel function, that is, the life cycle is until the corresponding kernel function ends. For device functions, the target thread local storage variable can also be obtained from the thread local memory space corresponding to the kernel function. Therefore, through the above-mentioned pre-allocation method, it is not necessary to perform additional processing on the prologue and epilogue of each function to manage the target thread local storage variables in the frame degradation stage, because the space for these variables has been reserved in advance.

[0174] Figure 5 A flowchart of a code compilation method provided in at least one embodiment of the present disclosure is used to compile code including at least one thread local storage variable. Figure 5 As shown, the code compilation method provided by at least one embodiment of the present disclosure includes the following steps S501-S506.

[0175] Step S501: for each target thread local storage variable, determine the first function that uses the target thread local storage variable to construct a variable-function mapping relationship. The description of step S501 can refer to the description of step S104 in the above embodiment, which is not repeated here.

[0176] Step S502: In response to the existence of the target thread local storage variable, at least one target thread local storage variable is allocated to the thread local memory based on the attribute of at least one target thread local storage variable, and the thread local memory address corresponding to each target thread local storage variable is obtained. The description of step S502 can refer to the description of step S101 in the above embodiment, which is not repeated here.

[0177] Step S503: Obtain a record of the thread local memory size required by all target thread local storage variables used by each first function in the code, so as to pre-allocate a fixed thread local memory space for each first function according to the record. For the description of step S503, reference may be made to the description of step S106 in the above embodiment, which will not be repeated here.

[0178] Step S504: inserting instructions for initializing the target thread local storage variables used by each first function into the intermediate representation corresponding to the code. The description of step S504 can refer to the description of step S105 in the above embodiment, which will not be repeated here.

[0179] Step S505: Update the intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread local storage variable. The description of step S505 can refer to the description of step S102 in the above embodiment, which will not be repeated here.

[0180] Step S506: Generate machine code corresponding to the code based on the updated intermediate representation. The description of step S506 can refer to the description of step S103 in the above embodiment, and will not be repeated here.

[0181] It should also be noted that in various embodiments of the present disclosure, the execution order of the various steps of the code compilation method is not limited. Although the execution process of each step is described in a specific order above, this does not constitute a limitation on the embodiments of the present disclosure. The various steps in the code compilation method can be executed serially or in parallel, which can be determined according to actual needs.

[0182] For example, compared with the above description, the code compilation method provided by at least one embodiment of the present disclosure may also include more or fewer steps, and the embodiments of the present disclosure are not limited to this.

[0183] Figure 6 A schematic block diagram of a code compiling device provided in at least one embodiment of the present disclosure. The code compiling device may be, for example, a compiler, configured to compile code including at least one thread local storage variable.

[0184] For example, Figure 6 As shown, the code compiling device provided by at least one embodiment of the present disclosure includes an allocating module 601 , an updating module 602 and a generating module 603 .

[0185] For example, the allocation module 601 is configured to allocate at least one target thread local storage variable to the thread local memory in response to the existence of the target thread local storage variable based on the attribute of at least one target thread local storage variable, and obtain the thread local memory address corresponding to each target thread local storage variable, wherein the target thread local storage variable is a thread local storage variable used by the first function in the at least one thread local storage variable. For the relevant content of the allocation module 601, reference can be made to the relevant description of step S101 in the embodiment of the above code compilation method, which will not be repeated here.

[0186] For example, the update module 602 is configured to update the intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread local storage variable. For the relevant content of the update module 602, reference can be made to the relevant description of step S102 in the embodiment of the code compilation method, which will not be repeated here.

[0187] For example, the generation module 603 is configured to generate machine code corresponding to the code based on the updated intermediate representation. For the relevant content of the generation module 603, reference may be made to the relevant description of step S103 in the embodiment of the code compilation method, which will not be repeated here.

[0188] For example, in at least one embodiment of the present disclosure, the code compiling apparatus 600 may further include an initialization module configured to insert instructions for initializing target thread local storage variables used by each first function into the intermediate representation corresponding to the code.

[0189] For example, in at least one embodiment of the present disclosure, the code compiling apparatus 600 may further include a construction module configured to determine, for each target thread local storage variable, a first function that uses the target thread local storage variable to construct a variable-function mapping relationship.

[0190] For example, in at least one embodiment of the present disclosure, the first function directly uses the target thread local storage variable, or uses the target thread local storage variable by calling the second function.

[0191] For example, in at least one embodiment of the present disclosure, the first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0192] For example, in at least one embodiment of the present disclosure, the allocation module includes a determination unit and an allocation unit. The determination unit is configured to determine a variable allocation order based on an attribute of at least one target thread local storage variable. The allocation unit is configured to allocate at least one target thread local storage variable to the thread local memory in sequence according to the variable allocation order.

[0193] For example, in at least one embodiment of the present disclosure, the attributes of the target thread local storage variable include: at least one of the first function number corresponding to the target thread local storage variable, the alignment parameter of the target thread local storage variable, the size of the target thread local storage variable, and the definition order of the target thread local storage variable.

[0194] For example, in at least one embodiment of the present disclosure, the determination unit is further configured to: arrange at least one target thread local storage variable in descending order according to the first function quantity corresponding to the at least one target thread local storage variable; arrange at least one target thread local storage variable in descending order according to the alignment parameter of the at least one target thread local storage variable; arrange at least one target thread local storage variable in ascending order according to the size of the at least one target thread local storage variable; or arrange at least one target thread local storage variable in the definition order of the at least one target thread local storage variable.

[0195] For example, in at least one embodiment of the present disclosure, the determination unit is further configured to: arrange at least one target thread local storage variable in descending order according to the number of first functions corresponding to at least one target thread local storage variable; in response to the presence of multiple first thread local storage variables with the same number of corresponding first functions in one target thread local storage variable, arrange the multiple first thread local storage variables in descending order according to the alignment parameters of the target thread local storage variables; in response to the presence of multiple second thread local storage variables with the same alignment parameters among the multiple first thread local storage variables, arrange the multiple second thread local storage variables in ascending order according to the size of the target thread local storage variable; in response to the presence of multiple third thread local storage variables with the same size among the multiple second thread local storage variables, arrange the multiple third thread local storage variables according to the definition order of the target thread local storage variable.

[0196] For example, in at least one embodiment of the present disclosure, the allocation module may also include a starting address determination unit, which is configured to use, for target thread local storage variables corresponding to multiple first functions, a maximum value among the allocated offsets corresponding to the multiple first functions as the allocation starting point of the target thread local storage variables in the thread local memory.

[0197] For example, in at least one embodiment of the present disclosure, the intermediate representation corresponding to the code includes instructions for obtaining the address of the target thread local storage variable, and the update module is further configured to: for each target thread local storage variable, replace the instruction for obtaining the address of the target thread local storage variable with the thread local memory address corresponding to the target thread local storage variable.

[0198] For example, in at least one embodiment of the present disclosure, the code compiling device 600 may further include a record acquisition module, which is configured to obtain records of the thread local memory sizes required for all target thread local storage variables used by each first function in the code, so as to pre-allocate a fixed thread local memory space for each first function according to the records.

[0199] For example, in at least one embodiment of the present disclosure, the life cycle of the pre-allocated thread local memory space is consistent with the execution cycle of the corresponding first function.

[0200] For example, in at least one embodiment of the present disclosure, the thread local memory is located in a graphics processor.

[0201] It should be noted that the above-mentioned various modules and units can be implemented by software, hardware, firmware or any combination thereof. For example, the allocation module, update module and generation module can be implemented as a distribution circuit, an update circuit and a generation circuit respectively. The embodiments of the present disclosure do not limit their specific implementation methods.

[0202] It should be understood that the code compiling device 600 provided in at least one embodiment of the present disclosure can be used to implement the aforementioned code compiling method, and can also achieve technical effects similar to those of the aforementioned code compiling method, which will not be elaborated herein.

[0203] It should be noted that in the embodiment of the present disclosure, the code compiling device 600 may include more or fewer modules or units, and the connection relationship between the modules or units is not limited and can be determined according to actual needs. The specific configuration of each module or unit is not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or in other applicable ways.

[0204] Figure 7 A schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.

[0205] For example, Figure 7 As shown, the electronic device 700 includes at least one processor 701 and at least one memory 702. The at least one memory 702 includes one or more computer program modules. The one or more computer program modules are stored in the memory 702 and are configured to be executed by the at least one processor 701. The one or more computer program modules include instructions for executing the above-mentioned code compilation method. When the instructions are executed by the at least one processor 701, one or more steps in the code compilation method provided in at least one embodiment of the present disclosure can be executed. The memory 702 and the processor 701 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0206] For example, the processor 701 may be a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), a general purpose graphics processing unit (GPGPU), an artificial intelligence (AI) accelerator, or other forms of processing units with data processing capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) may be an X86, ARM, RISC-V architecture, etc. The processor 701 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 700 to perform desired functions.

[0207] For example, the memory 702 may include any combination of one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (Cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disk read-only memory (CD-ROM), USB memory, flash memory, etc.

[0208] Figure 8 A schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.

[0209] The electronic devices in at least one embodiment of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0210] The electronic device includes at least one processor and a memory. The processor here may be referred to as the processing device 801 described below, and the memory may include at least one of the read-only memory (ROM), random access memory (RAM) and storage device 808 described below. The memory is used to store programs for executing the methods described in the above-mentioned various method embodiments; the processor is configured to execute the programs stored in the memory. The processor may include a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0211] like Figure 8As shown, the electronic device 800 may include a processing device 801 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device 808 into a random access memory (RAM). In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface is also connected to the bus 804.

[0212] Typically, the following devices may be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a display, a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0213] In particular, according to at least one embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, at least one embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of at least one embodiment of the present disclosure are executed.

[0214] It should be noted that the computer-readable medium of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In at least one embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In at least one embodiment of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0215] The computer-readable medium may be included in the electronic device 800 ; or may exist independently without being incorporated into the electronic device 800 .

[0216] Figure 9 A schematic block diagram of a non-transitory computer-readable storage medium is provided for at least one embodiment of the present disclosure.

[0217] For example, Figure 9 As shown, a non-transitory computer-readable storage medium 900 stores computer-readable instructions 901 , and when the computer-readable instructions 901 are executed by at least one processor, one or more steps of the above-mentioned code compilation method are performed.

[0218] For example, the storage medium may include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, or other applicable storage media. For example, the readable storage medium may also be Figure 7 For the memory 702 in, the related description can be referred to the aforementioned content and will not be repeated here.

[0219] Although the disclosure has been described in detail above with general descriptions and specific implementation methods, it is obvious to those skilled in the art that some modifications or improvements may be made to the embodiments of the disclosure. Therefore, these modifications or improvements made without departing from the spirit of the disclosure are within the scope of protection claimed by the disclosure.

[0220] There are a few points to note about this disclosure:

[0221] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to the general design.

[0222] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual scale.

[0223] (3) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to obtain new embodiments.

[0224] The above description is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.

Claims

1. A code compilation method, wherein: The code includes at least one thread local storage variable, and the code compiling method includes: In response to the existence of a target thread local storage variable, based on the attribute of at least one target thread local storage variable, the at least one target thread local storage variable is allocated to the thread local memory, and a thread local memory address corresponding to each target thread local storage variable is obtained, wherein the target thread local storage variable is a thread local storage variable used by the first function among the at least one thread local storage variable; Update the intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread local storage variable; Generate machine code corresponding to the code based on the updated intermediate representation.

2. The code compiling method according to claim 1, further comprising: Inserting instructions for initializing target thread local storage variables used by each first function into the intermediate representation corresponding to the code.

3. The code compiling method according to claim 1, further comprising: For each target thread local storage variable, a first function that uses the target thread local storage variable is determined to construct a variable-function mapping relationship.

4. The code compiling method according to claim 3, wherein: The first function directly uses the target thread local storage variable, or calls a second function to use the target thread local storage variable.

5. The code compiling method according to claim 4, wherein: The first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

6. The code compiling method according to claim 1, wherein: The allocating the at least one target thread local storage variable to the thread local memory based on the attribute of the at least one target thread local storage variable comprises: Determining a variable allocation order based on properties of the at least one target thread local storage variable; The at least one target thread local storage variable is sequentially allocated to the thread local memory according to the variable allocation order.

7. The code compiling method according to claim 6, wherein: The attributes of the target thread local storage variable include: at least one of the first function quantity corresponding to the target thread local storage variable, the alignment parameter of the target thread local storage variable, the size of the target thread local storage variable and the definition order of the target thread local storage variable.

8. The code compiling method according to claim 7, wherein: The determining the variable allocation order based on the attribute of the at least one target thread local storage variable comprises: Arrange the at least one target thread local storage variable in descending order according to the number of first functions corresponding to the at least one target thread local storage variable; Arrange the at least one target thread local storage variable in descending order according to an alignment parameter of the at least one target thread local storage variable; Arrange the at least one target thread local storage variable in ascending order according to the size of the at least one target thread local storage variable; or The at least one target thread local storage variable is arranged according to a definition order of the at least one target thread local storage variable.

9. The code compiling method according to claim 7, wherein: The determining the variable allocation order based on the attribute of the at least one target thread local storage variable comprises: Arrange the at least one target thread local storage variable in descending order according to the number of first functions corresponding to the at least one target thread local storage variable; In response to the existence of a plurality of first thread local storage variables having the same number of corresponding first functions in the at least one target thread local storage variable, arranging the plurality of first thread local storage variables in descending order according to alignment parameters of the target thread local storage variables; In response to the presence of a plurality of second thread local storage variables with the same alignment parameter among the plurality of first thread local storage variables, arranging the plurality of second thread local storage variables in ascending order according to the size of the target thread local storage variable; In response to the presence of a plurality of third thread local storage variables of the same size among the plurality of second thread local storage variables, the plurality of third thread local storage variables are arranged according to a definition order of the target thread local storage variable.

10. The code compiling method according to claim 6, wherein: The allocating the at least one target thread local storage variable to the thread local memory based on the attribute of the at least one target thread local storage variable further includes: For target thread local storage variables corresponding to multiple first functions, a maximum value among the allocated offsets corresponding to the multiple first functions is used as an allocation starting point for the target thread local storage variable in the thread local memory.

11. The code compiling method according to claim 1, wherein: The updating of the intermediate representation corresponding to the code based on the thread local memory address corresponding to each target thread local storage variable includes: For each target thread local storage variable, replace the instruction for obtaining the address of the target thread local storage variable with the thread local memory address corresponding to the target thread local storage variable, The intermediate representation corresponding to the code includes the instruction for obtaining the address of the local storage variable of the target thread.

12. The code compiling method according to claim 1, further comprising: A record of thread local memory size required by all target thread local storage variables used by each first function in the code is obtained, so as to pre-allocate a fixed thread local memory space for each first function according to the record.

13. The code compiling method according to claim 12, wherein: The life cycle of the pre-allocated thread local memory space is consistent with the execution cycle of the corresponding first function.

14. An electronic device comprising: at least one processor; at least one memory including one or more computer program modules; The one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the code compiling method according to any one of claims 1 to 13.

15. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed by at least one processor, the code compiling method according to any one of claims 1 to 13 is performed.

Citation Information

Patent Citations

  • Method and device for transplanting code from process model to thread model

    CN105786525A

  • Method and device for realizing thread local storage

    CN106445656A

  • Data processing method and device, storage medium and electronic equipment

    CN119127345A

  • Compiler optimization of coroutines

    WO2016176058A1

Cited By

  • Swan gap containerization TLS compatible method based on context awareness

    CN121387766A

  • A context-aware based harmonious containerized tls compatible method

    CN121387766B