Code compiling method, electronic equipment and storage medium

By dynamically allocating thread-local memory and generating machine code containing loading instructions, the problem that traditional thread-local storage implementation cannot be applied to GPU architecture is solved, and the utilization rate of thread-local memory is improved.

CN120215957AActive Publication Date: 2025-06-27SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510534628.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional thread-local storage implementations cannot be directly applied to GPU architecture, resulting in the difficulty in meeting the needs of creating thread-private data and supporting related data management and optimization in application scenarios such as parallel computing and machine learning.

Method used

A code compilation method is adopted to obtain records of thread local storage variables used by each function in the code, and pass them to the drive to dynamically allocate thread local memory, and generate machine code containing loading instructions to obtain the memory address of thread local storage variables.

Benefits of technology

It effectively avoids the waste of thread-local memory space, improves the utilization rate of thread-local memory space, and supports the thread-local storage model under the GPU architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215957A_ABST
    Figure CN120215957A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a code compiling method, electronic equipment and a storage medium. The code compiling method is used for compiling a code comprising at least one thread local storage variable, and comprises the following steps: acquiring a record of the thread local storage variable used by each first function in the code, and transmitting the record to a driver, enabling the driver to distribute each thread local storage variable to a thread local memory according to the record during operation; a machine code corresponding to the code is generated according to the intermediate representation corresponding to the code, the machine code comprises a loading instruction, and the loading instruction is used for obtaining a thread local memory address corresponding to each thread local storage variable based on the allocation result of the driver when being executed. According to the code compiling method, a mechanism of dynamically allocating the address space during operation is adopted, so that the waste of the local memory space of the thread is avoided, and the utilization rate of the local memory space of the thread is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a code compilation method, an electronic device, and a storage medium. Background Art

[0002] The Low-Level Virtual Machine (LLVM) is an open-source compiler framework, which is designed as a collection of modular and reusable compiler and toolchain technologies. LLVM supports the Thread Local Storage (TLS) model and can define thread-local storage variables, enabling each thread to have an independent variable copy, thereby avoiding data competition problems between threads. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a code compilation method, where the code includes at least one thread-local storage variable, and the code compilation method includes: obtaining a record of the thread-local storage variables used by each first function in the code and passing the record to a driver, so that the driver allocates each thread-local storage variable to thread-local memory according to the record during runtime; generating machine code corresponding to the code according to the intermediate representation corresponding to the code, where the machine code includes a load instruction, and the load instruction is used to obtain the thread-local memory address corresponding to each thread-local storage variable based on the allocation result of the driver when being executed.

[0004] In the code compilation method provided by at least one embodiment of the present disclosure, generating the machine code corresponding to the code according to the intermediate representation corresponding to the code includes: converting an instruction for obtaining the address of a thread-local storage variable in the intermediate representation corresponding to the code into the load instruction.

[0005] The code compilation method provided by at least one embodiment of the present disclosure further includes: inserting an instruction for initializing the thread-local storage variables used by each first function in the intermediate representation corresponding to the code.

[0006] In the code compilation method provided by at least one embodiment of the present disclosure, obtaining a record of the thread-local storage variables used by each first function in the code includes: for each thread-local storage variable, determining the first function using the thread-local storage variable to construct a thread-local storage variable - first function mapping relationship; performing an inverse mapping on the mapping relationship to obtain a record of the thread-local storage variables used by each first function in the code.

[0007] In the code compilation method provided by at least one embodiment of the present disclosure, the first function directly uses the thread local storage variable or calls a second function to use the thread local storage variable.

[0008] In the code compilation method provided by at least one embodiment of the present disclosure, the first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0009] The code compilation method provided by at least one embodiment of the present disclosure further includes: for an unused thread local storage variable, removing from the intermediate representation corresponding to the code an instruction for obtaining the address of the unused thread local storage variable.

[0010] The code compilation method provided by at least one embodiment of the present disclosure further includes: removing from the intermediate representation corresponding to the code duplicate instructions for obtaining the addresses of thread local storage variables, so that the address of each thread local storage variable in each first function is obtained only once.

[0011] In the code compilation method provided by at least one embodiment of the present disclosure, the load instruction is further used for: for each first function, reading from the register corresponding to the first function an address table of the thread local storage variables used by the first function; obtaining, based on the address table, the thread local memory address corresponding to each thread local storage variable used by the first function, where the address table is created by the driver in the register corresponding to the first function during runtime according to the address allocation result.

[0012] In the code compilation method provided by at least one embodiment of the present disclosure, the obtaining, based on the address table, the thread local memory address corresponding to each thread local storage variable used by the first function includes: calculating, according to the order of each thread local storage variable in the record corresponding to the first function and the address occupation size of each thread local storage variable used by the first function, an offset corresponding to the thread local memory address of each thread local storage variable used by the first function in the address table; and determining, based on the offset, the thread local memory address corresponding to each thread local storage variable used by the first function.

[0013] In the code compilation method provided by at least one embodiment of the present disclosure, the register includes a constant scalar register.

[0014] At least one embodiment of the present disclosure provides a code compilation device. The code includes at least one thread local storage variable. The code compilation device includes: an acquisition module configured to acquire records of thread local storage variables used by each first function in the code and pass the records to a driver, so that the driver allocates each thread local storage variable to thread local memory according to the records during runtime; a generation module configured to generate machine code corresponding to the code according to an intermediate representation corresponding to the code, where the machine code includes a load instruction, and the load instruction is used to obtain a thread local memory address corresponding to each thread local storage variable based on the allocation result of the driver when being executed.

[0015] At least one embodiment of the present disclosure provides an electronic device. The electronic device includes: at least one processor; at least one memory including one or more computer program modules; where the one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the code compilation method provided by at least one embodiment of the present disclosure.

[0016] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium having computer instructions stored thereon. The computer instructions, when executed by at least one processor, execute the code compilation method provided by at least one embodiment of the present disclosure.

[0017] The code compilation method, code compilation device, electronic device, and non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure adopt a mechanism of dynamically allocating an address space during runtime. For however many thread local storage variables the kernel function uses, the driver will correspondingly allocate that many thread local memory spaces for it, without causing waste of thread local memory space and improving the utilization rate of thread local memory space. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0019] Figure 1 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0020] Figure 2 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0021] Figure 3 It is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure;

[0022] Figure 4 An exemplary schematic diagram of a scheme for statically allocating an address space at compile time;

[0023] Figure 5 An exemplary schematic diagram of a scheme for dynamically allocating an address space at runtime provided by at least one embodiment of the present disclosure;

[0024] Figure 6 A schematic block diagram of a code compilation device provided by at least one embodiment of the present disclosure;

[0025] Figure 7 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure;

[0026] Figure 8 A schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure; and

[0027] Figure 9 A schematic block diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0028] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0029] Flowcharts are used in the present disclosure to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily have to be executed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0030] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are merely used to distinguish different components. Words such as "comprising" or "including" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0031] The following illustrates this disclosure through several specific embodiments. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of known functions and known components (elements) may be omitted. When any component (element) of an embodiment of this disclosure appears in more than one drawing, the component (element) is denoted by the same or similar reference numerals in each drawing.

[0032] LLVM is an open-source compiler framework that is designed as a collection of modular and reusable compiler and toolchain technologies. LLVM includes a front end, a middle end, and a back end.

[0033] The front end of LLVM is responsible for converting the source code of different programming languages or programming models into a general intermediate representation (IR). For example, Clang is a compiler front end natively supported by LLVM for C, C++ and Objective-C. When a user declares a variable using thread_local (C++11) or __thread (GCC extension), Clang can mark the variable as a thread-local storage variable (TLS variable), for example, it can be marked through the thread_local attribute and a specific linkage type (such as dso_local) in the generated intermediate representation.

[0034] The middle end of LLVM is responsible for optimizing the intermediate representation generated by the front end to improve the performance of the finally generated code. For example, it can simplify the thread-local storage access pattern and combine multiple accesses into one address calculation. For example, it can also implicitly select the thread-local storage model according to the scope of the variable, or explicitly specify the thread-local storage model in the code.

[0035] The backend of LLVM is responsible for converting the optimized intermediate representation into machine code or assembly code according to the characteristics of the target hardware platform. The specific tasks of the backend can include aspects such as instruction selection, register allocation, instruction scheduling, and code generation. For example, instruction selection refers to mapping the intermediate representation to the instruction set of the target hardware platform; register allocation refers to mapping the virtual registers in the intermediate representation to the actual registers of the target hardware; instruction scheduling refers to rearranging the instruction order to maximize the utilization of hardware resources; and code generation refers to generating the machine code or assembly code corresponding to the target hardware platform.

[0036] LLVM supports the thread-local storage model. Through the thread_local keyword, thread-local storage variables can be defined, enabling each thread to have an independent copy of the variable, thereby avoiding data competition issues between threads. In other words, thread-local storage variables are not shared between threads, which means that even if multiple threads access the same variable name, they are actually operating on their respective independent memory regions.

[0037] However, the inventors of the present disclosure have noticed that the implementation of thread-local storage depends on the support of the target platform. Currently, in traditional Central Processing Unit (CPU) architectures (such as x86 and ARM), the operating system and the Application Binary Interface (ABI) provide support for thread-local storage. For example, it can support the implementation of thread-local storage through specific segments in the Executable and Linkable Format (ELF) file (such as the.tdata segment for storing uninitialized thread-local storage variables and the.tbss segment for storing initialized thread-local storage variables). However, in backends related to Graphics Processing Unit (GPU), General-Purpose Graphics Processing Unit (GPGPU), or other parallel computing architectures (such as CUDA and OpenCL), there is no similar mechanism to support the implementation of thread-local storage, which poses challenges for application scenarios that require large-scale parallel processing such as high-performance computing and machine learning.

[0038] The parallel thread model of the GPU (such as CUDA and OpenCL) usually requires each thread to access its private data independently, but the traditional thread local storage implementation method (such as the CPU's __thread keyword) cannot be directly mapped to the GPU architecture, and it is difficult to meet the requirements of creating thread private data and supporting related data management and optimization. In addition, the GPU lacks hardware-level thread local storage support, which requires developers to manually manage thread local data (such as calculating offsets through thread indexes), increasing development complexity and making it prone to errors. In existing GPU programming practices, developers need to explicitly pass thread indexes (such as threadIdx.x in CUDA) and manually allocate global memory to simulate the storage of thread private data, which leads to code redundancy and performance loss. In addition, the existing LLVM compiler does not provide sufficient support for GPU thread local storage and cannot automatically generate efficient memory allocation and access code.

[0039] In response to the above problems, the inventors of the present disclosure have proposed a code compilation method that provides support for a thread local storage model under a parallel computing architecture, solving the problem that traditional thread local storage implementations cannot be directly applied to GPU architectures. The method provides a solution for statically allocating address space during compilation. Through static allocation and address solidification, it is not necessary to load the thread local memory address corresponding to each thread local storage variable from a register, effectively reducing register and move instruction overhead. However, the inventors of the present disclosure noticed that in this method, if a thread local storage variable is shared by multiple kernel functions, then even a kernel function that does not use the thread local storage variable will be forced to allocate the thread local memory space corresponding to the thread local storage variable, which results in a waste of thread local memory space in some scenarios.

[0040] At least one embodiment of the present disclosure provides a code compilation method for code including at least one thread local storage variable, the code compilation method comprising: obtaining a record of the thread local storage variables used by each first function in the code, and passing the record to a driver, so that the driver allocates each thread local storage variable to a thread local memory according to the record at runtime; generating a machine code corresponding to the code according to an intermediate representation corresponding to the code, wherein the machine code includes a load instruction, and the load instruction is used to obtain a thread local memory address corresponding to each thread local storage variable based on an allocation result of the driver when executed.

[0041] The code compiling method provided by at least one embodiment of the present disclosure may be executed by a compiler. For example, the compiler may be located at a host end, and the code compiling method may be executed by the host end.

[0042] In the code compilation method provided by at least one embodiment of the present disclosure, a mechanism of dynamically allocating address space at runtime is adopted. For the number of thread local storage variables used by the kernel function, the driver will allocate the corresponding amount of thread local memory space for it, which will not cause waste of thread local memory space and improves the utilization rate of thread local memory space.

[0043] Figure 1 The flowchart of a code compilation method provided by at least one embodiment of the present disclosure is used to compile code including at least one thread local storage variable. For example, as Figure 1 shown, the code compilation method provided by at least one embodiment of the present disclosure includes the following steps S101 to S102.

[0044] Step S101: Obtain the records of the thread local storage variables used by each first function in the code, and pass the records to the driver so that the driver can allocate each thread local storage variable to the thread local memory according to the records at runtime.

[0045] For example, the above-mentioned "at least one thread local storage variable" may include the thread local storage variables used by the first function, or may include the thread local storage variables not used by any first function. Each of the "each thread local storage variable" mentioned later refers to the thread local storage variables that are used, and the thread local storage variables that are not used will not be allocated to the thread local memory.

[0046] Thread Local Memory (TLM) refers to the memory area private to each thread, which is mainly used to store the data private to the thread and cannot be directly accessed by other threads. For example, the thread local memory can be the thread local memory in the GPU. By storing the thread local storage variables on the thread local memory, the burden on the registers in the GPU can be reduced. For example, the thread local memory can also be the thread local memory in other parallel computing architectures, and the embodiments of the present disclosure do not limit this.

[0047] It should be noted that thread local memory and the thread local storage introduced above are two different concepts. Thread local memory refers to a memory area, while thread local storage refers to a storage method of variables.

[0048] In the code compilation method provided by at least one embodiment of the present disclosure, the code may include multiple first functions and multiple second functions. The first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0049] For example, in at least one embodiment of the present disclosure, the host side may include a central processing unit (CPU), and the device side may include a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), etc.

[0050] For example, in at least one embodiment of the present disclosure, the first function may be a kernel function, and the second function may be a device function. A kernel function refers to a function called by the host side and executed by the device side, and a device function refers to a function called by the device side and executed by the device side. For example, a kernel function may be called by a host function (such as a host function) running on the CPU, while a device function may be called by a kernel function executed at the GPU or other device functions. The code usually includes multiple kernel functions and device functions. Each kernel function may call multiple device functions, and each device function may also be called by multiple kernel functions.

[0051] The "thread local storage variable used by the first function" refers to the thread local storage variable directly used by the first function, as well as the thread local storage variable used by the first function by calling the second function.

[0052] For example, when the compiler generates the executable file corresponding to the first function, it usually embeds a piece of metadata. Metadata can be regarded as a kind of additional information, which can be used to guide compiler optimization, debugging, code analysis, or runtime behavior. Metadata can provide necessary information to the runtime environment to guide the driver to perform corresponding processing, such as allocating thread local storage variables to thread local memory.

[0053] The thread local storage variables used by each first function in the code can be recorded in the metadata and passed to the driver, so that the driver can allocate each thread local storage variable to the thread local memory according to the records in the metadata at runtime.

[0054] The driver, for example, is a driver program, which is software running on the host side (such as the CPU) and is used to interact with the hardware and manage the hardware physical resources.

[0055] Step S102: Generate the machine code corresponding to the code according to the intermediate representation of the code, where the machine code includes a load instruction, and the load instruction is used to obtain the thread local memory address corresponding to each thread local storage variable based on the allocation result of the driver when being executed.

[0056] For example, in step S102, the intermediate representation corresponding to the code can be converted into machine code applicable to the target hardware platform through steps such as instruction selection, register allocation, instruction scheduling, and code generation introduced above. The intermediate representation corresponding to the code here can be the initial intermediate representation automatically generated by the Clang front end of LLVM, or the intermediate representation obtained after optimizing the initial intermediate representation.

[0057] For example, in the instruction selection phase, the instruction for obtaining the address of a thread-local storage variable in the intermediate representation corresponding to the code is converted into a load instruction, and the load instruction will be embedded in the finally generated machine code. The load instruction is used to obtain the thread-local memory address corresponding to each thread-local storage variable based on the allocation result of the driver when it is executed. When the machine code is executed, the included load instruction will also be executed, so as to obtain the thread-local memory address corresponding to each thread-local storage variable based on the allocation result of the driver.

[0058] An example of an instruction for obtaining the address of a thread-local storage variable in an intermediate representation is as follows:

[0059] @llvm.threadlocal.address.p0(ptr addrspacecast (ptr addrspace(1) @ato ptr)

[0060] In the above intermediate representation, the llvm.threadlocal.address.p0 instruction is used to obtain the address of the thread-local storage variable a.

[0061] In the code compilation method provided in at least one of the above embodiments of the present disclosure, during compilation, the compiler obtains a record of the thread-local storage variables used by each first function in the code and passes the record to the driver, so that the driver can allocate each thread-local storage variable to the thread-local memory according to the record at runtime; and, the machine code generated by the compiler includes a load instruction, and the load instruction can obtain the thread-local memory address corresponding to each thread-local storage variable based on the allocation result of the driver when it is executed. Therefore, the record obtained in step S101 is the basis for the address allocation by the driver at runtime, and the address allocation result of the driver is the basis for the load instruction generated in step S102 to achieve address acquisition.

[0062] In the code compilation method provided in at least one of the above embodiments of the present disclosure, through the mechanism of dynamic allocation at runtime, the driver will allocate as much thread-local memory space as the number of thread-local storage variables used by the first function, without causing waste of thread-local memory space.

[0063] At runtime, the driver allocates each thread-local storage variable to the thread-local memory according to the record of the thread-local storage variables used by each first function, and then creates an address table for each first function according to the result of the address allocation. The address table records the thread-local memory addresses corresponding to each thread-local storage variable in the first function. For example, the order in which the thread-local memory addresses corresponding to each thread-local storage variable are recorded in the address table can be determined by the order of the thread-local storage variables recorded in the metadata.

[0064] It should be noted that the occupied size of the thread-local memory address corresponding to the thread-local storage variable and the size of the thread-local storage variable are two different concepts.

[0065] For example, the occupied size of the thread-local memory address corresponding to the thread-local storage variable (hereinafter also referred to as the address occupied size) refers to the size of the space occupied by the address in the address table, rather than the size of the thread-local storage variable. For example, different hardware unit addresses can correspond to different occupied sizes, and the occupied size of the thread-local memory address can be 4 bytes, for example. For example, according to the address occupied size, the offset of the thread-local memory address corresponding to the thread-local storage variable in the address table can be determined.

[0066] For example, the size of the thread-local storage variable can be understood as the storage space size that the thread-local storage variable needs to occupy, which is related to the type of the variable and is usually in bytes. For example, the size of a thread-local storage variable of the short type is 2 bytes, and the size of a thread-local storage variable of the int type is 4 bytes. It should be noted that in different systems or compilers, the size of the same type of thread-local storage variable may be different. For example, the size of the corresponding thread-local storage variable can be determined by the size field in the metadata, and the driver can allocate each thread-local storage variable to the thread-local memory according to this information at runtime.

[0067] For example, a corresponding register can be set for each first function, and the driver can create an address table in the register. An example of the register can be a constant scalar register (CSR).

[0068] Correspondingly, in the code compilation method provided in at least one embodiment of the present disclosure, the load instruction is further used to execute the following steps S110 to step S120.

[0069] Step S110: For each first function, read the address table of the thread-local storage variables used by the first function from the register corresponding to the first function.

[0070] For example, in step S110, the address table is created by the driver during runtime in the register corresponding to the first function according to the address allocation result.

[0071] Step S120: Obtain the thread-local memory address corresponding to each thread-local storage variable used by the first function based on the address table.

[0072] For example, in step S120, the thread-local memory address corresponding to each thread-local storage variable used by the first function can be read from the address table.

[0073] In the code compilation method provided by at least one embodiment of the present disclosure, an example of step S120 may include steps S121 to S122.

[0074] Step S121: Calculate the offset corresponding to the thread-local memory address of each thread-local storage variable used by the first function in the address table according to the order of each thread-local storage variable in the record corresponding to the first function and the address occupation size of each thread-local storage variable used by the first function.

[0075] Since the driver determines the order of the thread-local memory addresses corresponding to each thread-local storage variable recorded in the address table according to the order of the thread-local storage variables recorded in the metadata when creating the address table, in step S121, it is also necessary to find the storage position of the thread-local memory address corresponding to each thread-local storage variable in the address table according to this order. For example, for each thread-local storage variable, the offset relative to the starting position of the address table can be calculated in turn according to its order recorded in the metadata and the occupation size of its corresponding thread-local memory address.

[0076] In some examples, the relevant information of thread-local storage variables a, b, and c is recorded in the metadata in sequence. Assuming that the address occupation size of each thread-local storage variable is 4 bytes, then the offset corresponding to the thread-local memory address of thread-local storage variable a in the address table is 0, the offset corresponding to the thread-local memory address of thread-local storage variable b in the address table is 4, and the offset corresponding to the thread-local memory address of thread-local storage variable c in the address table is 8.

[0077] Step S122: Determine the thread-local memory address corresponding to each thread-local storage variable used by the first function based on the offset.

[0078] For example, in step S122, for each thread-local storage variable, the thread-local memory address corresponding to the thread-local storage variable can be found in the address table according to the calculated offset.

[0079] In the above example, the thread-local memory address corresponding to the thread-local storage variable a can be found at the starting position of the address table, the thread-local memory address corresponding to the thread-local storage variable b can be found at the position of the starting position of the address table + 4, and the thread-local memory address corresponding to the thread-local storage variable c can be found at the position of the starting position of the address table + 8.

[0080] Figure 2 The flowchart of a code compilation method provided by at least one embodiment of the present disclosure.

[0081] As Figure 2 shown, in the code compilation method provided by at least one embodiment of the present disclosure, an example of "obtaining records of thread-local storage variables used by each first function in the code" in step S101 may include the following steps S201 to step S202.

[0082] Step S201: For each thread-local storage variable, determine the first function that uses the thread-local storage variable to construct a thread-local storage variable - first function mapping relationship.

[0083] In step S201, the transfer path of the variable in the function can be traced according to the call graph (also known as the call relationship graph) to determine which functions use the thread-local storage variable. The call graph is a directed graph that can represent the call relationship between functions in a program. The compiler can generate the call graph by means of, for example, static analysis. For example, the call graph can be generated by parsing the call relationship between functions through a call graph builder integrated in the compiler. For example, the thread-local storage variable - first function mapping relationship records which first functions use the thread-local storage variable. The thread-local storage variable - first function mapping relationship can be presented in the form of a mapping graph.

[0084] It should be noted that since some thread-local storage variables in the code may only be defined but not used by any first function, not every thread-local storage variable can find a corresponding first function in the thread-local storage variable - first function mapping relationship.

[0085] In step S201, the "first function that uses the thread-local storage variable" refers to the first function that directly or indirectly uses the thread-local storage variable. For example, "directly uses" means that the first function directly uses the thread-local storage variable, and "indirectly uses" means that the first function uses the thread-local storage variable by calling a second function.

[0086] For example, an example of step S201 may be: for each thread local storage variable, determine the kernel functions that directly or indirectly use the thread local storage variable, so as to build a thread local storage variable - first function mapping relationship.

[0087] Step S202: Perform an inverse mapping on the mapping relationship to obtain a record of the thread local storage variables used by each first function in the code.

[0088] For example, the mapping relationship obtained in step S201 is a mapping relationship of thread local storage variable → first function, that is, which first functions use each thread local storage variable. In step S202, perform an inverse mapping on it to obtain an inverse mapping relationship of first function → thread local storage variable, that is, which thread local storage variables each first function uses.

[0089] In step S202, the inverse mapping relationship obtained through the inverse mapping can be recorded in the metadata.

[0090] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S103. Step S103 may be executed before step S101.

[0091] Step S103: For unused thread local storage variables, remove the instructions for obtaining the addresses of the unused thread local storage variables from the intermediate representation corresponding to the code.

[0092] For example, the initial intermediate representation automatically generated by the Clang front end of LLVM can be analyzed, that is, the usage of each thread local storage variable by each first function is analyzed. For unused thread local storage variables, remove the instructions for obtaining the addresses of the thread local storage variables from the intermediate representation corresponding to the first function.

[0093] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S104. Step S104 may be executed before step S101.

[0094] Step S104: Remove duplicate instructions for obtaining the addresses of thread local storage variables from the intermediate representation corresponding to the code, so that the address of each thread local storage variable in each first function is obtained only once.

[0095] For example, the instructions for obtaining the addresses of thread local storage variables in the intermediate representation corresponding to the code (such as llvm.threadlocal.address.p0) can be optimized, and duplicate instructions are removed, so that the address of each thread local storage variable in each first function (such as a kernel function) is obtained only once.

[0096] For example, if there are multiple identical instructions for obtaining the address of a thread-local storage variable in the intermediate representation corresponding to a first function, the duplicate instructions need to be removed, and only the first instruction is retained.

[0097] A special example is as follows. Suppose c is a variable of array type. For a kernel function that uses thread-local storage variables c[1] and c[2], the instructions for obtaining the addresses of thread-local storage variables c[1] and c[2] in the intermediate representation need to be removed, and only the instruction for obtaining the base address of thread-local storage variable c is retained. The addresses of c[1] and c[2] can be calculated by adding the corresponding offsets to the base address of c.

[0098] For example, the code compilation method provided by at least one embodiment of the present disclosure may include steps S103, S101, and S102 that are executed in sequence. Alternatively, the code compilation method provided by at least one embodiment of the present disclosure may include steps S104, S101, and S102 that are executed in sequence.

[0099] In other examples, the code compilation method provided by at least one embodiment of the present disclosure may include steps S103, S104, S101, and S102 that are executed in sequence. Alternatively, the code compilation method provided by at least one embodiment of the present disclosure may include steps S104, S103, S101, and S102 that are executed in sequence.

[0100] Through the above method, duplicate address acquisition operations can be avoided, significantly improving the running efficiency and reducing the overhead of redundant instructions.

[0101] The code compilation method provided by at least one embodiment of the present disclosure may further include the following step S105. Step S105 may be executed before step S103 and step S104.

[0102] Step S105: Insert instructions for initializing the thread-local storage variables used by each first function into the intermediate representation corresponding to the code.

[0103] For example, in step S105, the code may include multiple first functions and multiple second functions. According to the call graph, the thread-local storage variables directly or indirectly used by each first function can be determined. For each first function, instructions for initializing the thread-local storage variables directly or indirectly used by it can be inserted at the starting position of its corresponding intermediate representation.

[0104] For example, taking the thread-local storage variable a as an example, an example of an instruction for initializing the thread-local storage variable a is as follows:

[0105] %0 = call ptr @llvm.threadlocal.address.p0(ptr addrspacecast (ptr addrspace(1) @a to ptr))

[0106] store i32 1, ptr %0, align 4

[0107] The meaning of the above intermediate representation is that first, the llvm.threadlocal.address.p0 instruction is used to obtain the address of the thread-local storage variable a, and then the store instruction is used to store the initial value of variable a (which is 1 in this example) to this address. It should be noted that the above initialization instructions are only examples, and different initialization instructions can be inserted according to actual needs.

[0108] It should be noted that the above initialization instructions will also be optimized together in steps S103 and S104, that is, the repeated llvm.threadlocal.address.p0 instructions will be removed.

[0109] In the CPU architecture, thread-local storage variables are usually initialized by the runtime environment when the program starts. Different from the above initialization method, the code compilation method provided by at least one embodiment of the present disclosure inserts instructions for initializing thread-local storage variables during compilation, rather than initializing thread-local storage variables at runtime, which can effectively avoid refreshing the entire thread-local memory at runtime and achieve performance optimization. In addition, this method increases more opportunities for compilation optimization, reduces the coupling between compilation and runtime, and reduces the workload during compilation and runtime.

[0110] Figure 3 The figure is a flowchart of a code compilation method provided by at least one embodiment of the present disclosure for compiling code including at least one thread-local storage variable. For example, as Figure 3 shown, the code compilation method provided by at least one embodiment of the present disclosure includes the following steps S301 to S305.

[0111] Step S301: Insert instructions for initializing the thread-local storage variables used by each first function into the intermediate representation corresponding to the code. The description of step S301 can refer to the description of step S105 in the above embodiment and will not be elaborated here.

[0112] Step S302: For unused thread-local storage variables, remove the instructions for obtaining the addresses of the unused thread-local storage variables from the intermediate representation corresponding to the code. The description of step S302 can refer to the description of step S103 in the above embodiment and will not be elaborated here.

[0113] Step S303: Remove the duplicate instructions for obtaining the addresses of thread-local storage variables from the intermediate representation corresponding to the code, so that the address of each thread-local storage variable in each first function is obtained only once. For the description of step S303, reference can be made to the description of step S104 in the above embodiments, which will not be elaborated here.

[0114] Step S304: Obtain the records of the thread-local storage variables used by each first function in the code, and pass the records to the driver, so that the driver allocates each thread-local storage variable to the thread-local memory according to the records during runtime. For the description of step S304, reference can be made to the description of step S101 in the above embodiments, which will not be elaborated here.

[0115] Step S305: Generate the machine code corresponding to the code according to the intermediate representation corresponding to the code, where the machine code includes load instructions, and the load instructions are used to obtain the thread-local memory addresses corresponding to each thread-local storage variable based on the allocation result of the driver when being executed. For the description of step S305, reference can be made to the description of step S102 in the above embodiments, which will not be elaborated here.

[0116] Next, a set of specific examples will be used to illustrate the differences between the scheme of statically allocating address space during compilation and the scheme of dynamically allocating address space during runtime proposed by at least one embodiment of the present disclosure.

[0117] Figure 4 It is an exemplary schematic diagram of a scheme for statically allocating address space during compilation.

[0118] For example, as Figure 4 shown, the code includes three kernel functions and six device functions. The three kernel functions are respectively represented as Kernel 1 to Kernel 3, and the six device functions are respectively represented as Func 1 to Func 6. As Figure 4 shown in the right call graph, the kernel function Kernel 1 calls the device functions Func 1 and Func 2, the kernel function Kernel 2 calls the device functions Func3 and Func4, and the kernel function kernel3 calls the device functions Func5 and Func6. Among them, the device function Func1 uses the thread-local storage variable a, the device function Func2 uses the thread-local storage variable b, the device function Func3 uses the thread-local storage variable b, the device function Func4 uses the thread-local storage variable c, the device function Func5 uses the thread-local storage variable c, and the device function Func6 uses the thread-local storage variable a.

[0119] First, allocate the thread-local storage variable a. The thread-local storage variable a is used by kernel function Kernel 1 and kernel function Kernel 3, and the allocated offsets corresponding to kernel function Kernel 1 and Kernel 3 are both 0 bytes, that is, the maximum value of the allocated offset is 0 bytes. Therefore, the starting address of the allocation of the thread-local storage variable a in the thread-local memory is 0, and 4 bytes of thread-local memory space are allocated.

[0120] Next, allocate the thread-local storage variable b. The thread-local storage variable b is used by kernel function Kernel 1 and kernel function Kernel 2, and the allocated offset corresponding to kernel function Kernel 1 is 4 bytes, and the allocated offset corresponding to kernel function Kernel 2 is 0 bytes, that is, the maximum value of the allocated offset is 4 bytes. Therefore, the starting address of the allocation of the thread-local storage variable b in the thread-local memory is 4, and 4 bytes of thread-local memory space are allocated.

[0121] Finally, allocate the thread-local storage variable c. The thread-local storage variable c is used by kernel function Kernel 2 and kernel function Kernel 3, and the allocated offset corresponding to kernel function Kernel 2 is 8 bytes, and the allocated offset corresponding to kernel function Kernel 3 is 4 bytes, that is, the maximum value of the allocated offset is 8 bytes. Therefore, the starting address of the allocation of the thread-local storage variable c in the thread-local memory is 8, and 4 bytes of thread-local memory space are allocated.

[0122] After the address allocation is completed, this solution optimizes the intermediate representation (directly replaces the instruction for obtaining the address of the thread-local storage variable with the thread-local memory address corresponding to the thread-local storage variable), directly fixes the corresponding thread-local memory address during compilation, and there is no need to load the thread-local memory address corresponding to each thread-local storage variable from the register, effectively reducing the register and move instruction overhead.

[0123] In this solution, an independent thread-local memory space is allocated for each kernel function. Therefore, the "starting address of the allocation in the thread-local memory" in the above text refers to the offset relative to the starting address of each thread-local memory space. It should be noted that Figure 4 uses a stack to illustrate the allocation of thread-local storage variables. The stack here can be regarded as a virtual concept. In LLVM, when the intermediate representation is converted into machine code, the allocation on the stack will be mapped to the actual storage space of the target architecture, that is, it is mapped to the corresponding thread-local memory space according to the position of the thread-local storage variable on the stack.

[0124] In the case of statically allocating address space at compile time, since multiple kernel functions and multiple device functions may use the same thread-local storage variable, and the device function cannot determine which kernel functions will call it, it is required that the addresses allocated to the same thread-local storage variable in all kernel functions and device functions are consistent, that is, no matter which kernel function or device function it is, it can find the thread-local storage variable in the thread-local memory through the same address. By Figure 4 's example, the consistency of the address allocation of the same thread-local storage variable required above can be achieved.

[0125] However, through Figure 4 it can be seen that in the static allocation scheme, if a thread-local storage variable is shared by multiple kernel functions, then even a kernel function that does not use the thread-local storage variable will be forced to allocate the thread-local memory space corresponding to the thread-local storage variable. For example, kernel function Kernel 2 does not use thread-local storage variable a, and kernel function Kernel 3 does not use thread-local storage variable b, but the corresponding thread-local memory space is still allocated, which causes waste of thread-local memory space.

[0126] The code compilation method provided by at least one embodiment of the present disclosure gives a scheme for dynamically allocating address space at runtime. Compared with the above scheme of statically allocating address space at compile time, this scheme can effectively reduce the waste of thread-local memory space.

[0127] Figure 5 It is an exemplary schematic diagram of a scheme for dynamically allocating address space at runtime provided by at least one embodiment of the present disclosure.

[0128] For example, as Figure 5 shown, the code includes three kernel functions and six device functions. The three kernel functions are respectively denoted as Kernel 1~Kernel 3, and the six device functions are respectively denoted as Func 1~Func 6. As Figure 5 shown in the right call graph, kernel function Kernel 1 calls device functions Func 1 and Func 2, kernel function Kernel 2 calls device functions Func3 and Func4, and kernel function kernel3 calls device functions Func5 and Func6. Among them, device function Func1 uses thread-local storage variable a, device function Func2 uses thread-local storage variable b, device function Func3 uses thread-local storage variable b, device function Func4 uses thread-local storage variable c, device function Func5 uses thread-local storage variable c, and device function Func6 uses thread-local storage variable a.

[0129] For example, as Figure 5As shown in the figure, the constant scalar register (CSR) corresponding to the kernel function Kernel 1 records the address table of the thread-local storage variables used by the kernel function Kernel 1, which stores the thread-local memory addresses corresponding to the thread-local storage variables a and b respectively; the constant scalar register corresponding to the kernel function Kernel 2 records the address table of the thread-local storage variables used by the kernel function Kernel 2, which stores the thread-local memory addresses corresponding to the thread-local storage variables b and c respectively; the constant scalar register corresponding to the kernel function Kernel 3 records the address table of the thread-local storage variables used by the kernel function Kernel 3, which stores the thread-local memory addresses corresponding to the thread-local storage variables a and c respectively. These address tables are stored in the corresponding constant scalar registers by the driver at runtime, that is, the thread-local memory addresses corresponding to each thread-local storage variable are allocated by the driver at runtime, and the compiler only needs to obtain the thread-local memory addresses corresponding to each thread-local storage variable from the constant scalar register. Through the mechanism of dynamic allocation at runtime, the driver will allocate the corresponding thread-local memory space according to the number of thread-local storage variables used by the kernel function, without wasting the thread-local memory space.

[0130] It should be noted that Figure 5 uses a stack to illustrate the allocation of thread-local storage variables, and the stack here can be regarded as a virtual concept. In LLVM, when the intermediate representation is converted into machine code, the allocation on the stack will be mapped to the actual storage space of the target architecture, that is, it is mapped to the corresponding thread-local memory space according to the position of the thread-local storage variables on the stack.

[0131] It should be noted that thread-local memory can also be used to store local variables other than thread-local storage variables. For example, as Figure 4 and Figure 5 shown, the stack corresponding to each kernel function can also include local variables.

[0132] Compared with the waste of thread-local memory space caused by the scheme of statically allocating address space at compile time, the scheme of dynamically allocating address space at runtime provided by at least one embodiment of the present disclosure effectively avoids the waste of thread-local memory space.

[0133] It should also be noted that in various embodiments of the present disclosure, the execution order of the steps of the code compilation method is not limited. Although the execution processes of the steps are described in a specific order above, this does not constitute a limitation on the embodiments of the present disclosure. The steps in the code compilation method can be executed serially or in parallel, which can be determined according to actual needs.

[0134] For example, compared with the above description, the code compilation method provided by at least one embodiment of the present disclosure may further include more or fewer steps, and the embodiments of the present disclosure do not limit this.

[0135] Figure 6 It is a schematic block diagram of a code compilation device provided by at least one embodiment of the present disclosure. The code compilation device may be, for example, a compiler, and is used to compile code including at least one thread-local storage variable.

[0136] For example, as Figure 6 shown, the code compilation device provided by at least one embodiment of the present disclosure includes an acquisition module 601 and a generation module 602.

[0137] For example, the acquisition module 601 is configured to acquire records of thread-local storage variables used by each first function in the code, and pass the records to the driver, so that the driver allocates each thread-local storage variable to the thread-local memory according to the records during runtime. For the relevant content of the acquisition module 601, reference may be made to the relevant description of step S101 in the embodiment of the above code compilation method, which will not be elaborated here.

[0138] For example, the generation module 602 is configured to generate machine code corresponding to the code according to the intermediate representation corresponding to the code, where the machine code includes a load instruction, and the load instruction is used to obtain the thread-local memory address corresponding to each thread-local storage variable based on the allocation result of the driver when being executed. For the relevant content of the generation module 602, reference may be made to the relevant description of step S102 in the embodiment of the above code compilation method, which will not be elaborated here.

[0139] For example, in at least one embodiment of the present disclosure, the generation module 602 is further configured to convert an instruction for obtaining the address of a thread-local storage variable in the intermediate representation corresponding to the code into a load instruction.

[0140] For example, in at least one embodiment of the present disclosure, the code compilation device 600 may further include an initialization module, which is configured to insert an instruction for initializing the thread-local storage variables used by each first function in the intermediate representation corresponding to the code.

[0141] For example, in at least one embodiment of the present disclosure, the acquisition module 601 is further configured to, for each thread-local storage variable, determine the first function using the thread-local storage variable to construct a thread-local storage variable - first function mapping relationship; perform an inverse mapping on the mapping relationship to obtain records of the thread-local storage variables used by each first function in the code.

[0142] For example, in at least one embodiment of the present disclosure, the first function directly uses the thread-local storage variable, or uses the thread-local storage variable by calling a second function.

[0143] For example, in at least one embodiment of the present disclosure, the first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

[0144] For example, in at least one embodiment of the present disclosure, the code compilation device 600 may further include a first removal module configured to remove, from the intermediate representation corresponding to the code, an instruction for obtaining the address of an unused thread-local storage variable for an unused thread-local storage variable.

[0145] For example, in at least one embodiment of the present disclosure, the code compilation device 600 may further include a second removal module configured to remove, from the intermediate representation corresponding to the code, duplicate instructions for obtaining the address of a thread-local storage variable, such that the address of each thread-local storage variable within each first function is obtained only once.

[0146] For example, in at least one embodiment of the present disclosure, the load instruction is further configured to: for each first function, read, from the register corresponding to the first function, an address table of thread-local storage variables used by the first function; and obtain, based on the address table, the thread-local memory address corresponding to each thread-local storage variable used by the first function, where the address table is created by the driver at runtime in the register corresponding to the first function according to the address allocation result.

[0147] For example, in at least one embodiment of the present disclosure, obtaining, based on the address table, the thread-local memory address corresponding to each thread-local storage variable used by the first function includes: calculating, according to the order of each thread-local storage variable in the record corresponding to the first function and the address occupation size of each thread-local storage variable used by the first function, an offset corresponding to the thread-local memory address of each thread-local storage variable used by the first function in the address table; and determining, based on the offset, the thread-local memory address corresponding to each thread-local storage variable used by the first function.

[0148] For example, in at least one embodiment of the present disclosure, the register includes a constant scalar register.

[0149] For example, in at least one embodiment of the present disclosure, the thread-local memory is located in the graphics processing unit.

[0150] It should be noted that the above various modules and units can be implemented by software, hardware, firmware, or any combination thereof. For example, the acquisition module and the generation module can be respectively implemented as an acquisition circuit and a generation circuit, and the embodiments of the present disclosure do not limit their specific implementation manners.

[0151] It should be understood that the code compilation device 600 provided in at least one embodiment of the present disclosure can be used to implement the foregoing code compilation method, and can also achieve technical effects similar to those of the foregoing code compilation method, which will not be elaborated herein.

[0152] It should be noted that in the embodiments of the present disclosure, the code compilation device 600 may include more or fewer modules or units, and the connection relationships between the various modules or units are not limited and can be determined according to actual needs. The specific composition manners of the various modules or units are not limited and can be constituted by analog devices according to circuit principles, or can be constituted by digital chips, or in other applicable manners.

[0153] Figure 7 It is a schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure.

[0154] For example, as Figure 7 shown, the electronic device 700 includes at least one processor 701 and at least one memory 702. The at least one memory 702 includes one or more computer program modules. The one or more computer program modules are stored in the memory 702 and are configured to be executed by the at least one processor 701. The one or more computer program modules include instructions for executing the foregoing code compilation method. When executed by the at least one processor 701, one or more steps in the code compilation method provided in at least one embodiment of the present disclosure can be executed. The memory 702 and the processor 701 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0155] For example, the processor 701 can be a central processing unit (CPU), a digital signal processor (DSP), an image processor (GPU), a general-purpose graphics processing unit (GPGPU), an artificial intelligence (AI) accelerator, or other forms of processing units with data processing capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) can be of X86, ARM, RISC-V architectures, etc. The processor 701 can be a general-purpose processor or a dedicated processor and can control other components in the electronic device 700 to execute desired functions.

[0156] For example, the memory 702 may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.

[0157] Figure 8 Schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure.

[0158] The electronic device in at least one embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle terminals (such as in-vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0159] The electronic device includes at least one processor and a memory. The processor here may be referred to as the processing device 801 described below, and the memory may include at least one of the read-only memory (ROM), random access memory (RAM), and storage device 808 described below. The memory is used to store programs for executing the methods described in the above respective method embodiments; the processor is configured to execute the programs stored in the memory. The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0160] As Figure 8 shown, the electronic device 800 may include a processing device 801 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) or programs loaded from the storage device 808 into the random access memory (RAM). In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface is also connected to the bus 804.

[0161] Typically, the following devices can be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a display, a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 can allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0162] In particular, according to at least one embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, at least one embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the method of at least one embodiment of the present disclosure are executed.

[0163] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In at least one embodiment of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In at least one embodiment of the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0164] The above computer-readable medium can be included in the above electronic device 800; or it can exist separately and not be assembled into the electronic device 800.

[0165] Figure 9 It is a schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure.

[0166] For example, as Figure 9 shown, computer-readable instructions 901 are stored on a non-transitory computer-readable storage medium 900, and when the computer-readable instructions 901 are executed by at least one processor, one or more steps of the above code compilation method are executed.

[0167] For example, the storage medium may include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), a portable compact disc read only memory (CD-ROM), a flash memory, or any combination of the above storage media, and may also be other applicable storage media. For example, the readable storage medium may also be the Figure 7 memory 702 in

[0168] Although the present disclosure has been described in detail with general descriptions and specific embodiments above, based on the embodiments of the present disclosure, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure all fall within the scope of protection required by the present disclosure.

[0169] The following points need to be noted regarding the present disclosure:

[0170] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0171] (2) For the sake of clarity, in the accompanying drawings used to describe the embodiments of the present disclosure, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual scale.

[0172] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0173] The above is only the specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A code compilation method, wherein: The code includes at least one thread local storage variable, and the code compiling method includes: Obtaining a record of thread local storage variables used by each first function in the code, and passing the record to a driver, so that the driver allocates each thread local storage variable to a thread local memory according to the record at runtime; Generate machine code corresponding to the code according to the intermediate representation corresponding to the code, wherein the machine code includes a load instruction, and the load instruction is used to obtain the thread local memory address corresponding to each thread local storage variable based on the allocation result of the driver when executed.

2. The code compiling method according to claim 1, wherein: Generating a machine code corresponding to the code according to the intermediate representation corresponding to the code includes: An instruction for obtaining an address of a thread local storage variable in an intermediate representation corresponding to the code is converted into the load instruction.

3. The code compiling method according to claim 1, further comprising: Inserting instructions for initializing thread local storage variables used by each first function into the intermediate representation corresponding to the code.

4. The code compiling method according to claim 1, wherein: The obtaining of a record of thread local storage variables used by each first function in the code includes: For each thread local storage variable, determining a first function that uses the thread local storage variable to construct a thread local storage variable-first function mapping relationship; The mapping relationship is reversely mapped to obtain a record of the thread local storage variable used by each first function in the code.

5. The code compiling method according to claim 4, wherein: The first function uses the thread local storage variable directly or by calling a second function to use the thread local storage variable.

6. The code compiling method according to claim 5, wherein: The first function is a function called by the host side and executed by the device side, and the second function is a function called by the device side and executed by the device side.

7. The code compiling method according to claim 1, further comprising: For unused thread local storage variables, instructions for obtaining addresses of the unused thread local storage variables are removed from the intermediate representation corresponding to the code.

8. The code compiling method according to claim 1, further comprising: Repeated instructions for obtaining addresses of thread local storage variables are removed from the intermediate representation corresponding to the code, so that the address of each thread local storage variable in each first function is obtained only once.

9. The code compiling method according to claim 1, wherein: The load instruction is further used to: For each first function, Reading an address table of thread local storage variables used by the first function from a register corresponding to the first function; Based on the address table, obtain the thread local memory address corresponding to each thread local storage variable used by the first function, The address table is created by the driver in a register corresponding to the first function according to an address allocation result when the driver is running.

10. The code compiling method according to claim 9, wherein: The acquiring, based on the address table, a thread local memory address corresponding to each thread local storage variable used by the first function includes: Calculate the offset of the thread local memory address corresponding to each thread local storage variable used by the first function in the address table according to the order of each thread local storage variable in the record corresponding to the first function and the address occupied by each thread local storage variable used by the first function; A thread local memory address corresponding to each thread local storage variable used by the first function is determined based on the offset.

11. The code compiling method according to claim 9, wherein: The registers include constant scalar registers.

12. An electronic device comprising: at least one processor; at least one memory including one or more computer program modules; The one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the code compiling method according to any one of claims 1 to 11.

13. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed by at least one processor, the code compiling method according to any one of claims 1 to 11 is performed.

Citation Information

Patent Citations

  • Memory allocation method and device, computer equipment and storage medium

    CN111324461A

  • Memory monitoring method and device, medium and electronic equipment

    CN112084024A

  • General register dynamic allocation method and device, computer equipment and storage medium

    CN115934102A

  • Industrial program compiling method, industrial program running method and related devices

    CN118585198A

  • Method and device for determining program memory allocation based on large model, equipment and medium

    CN119597296A

Cited By

  • NPU instruction optimization system and optimization method for OpenCL heterogeneous computing

    CN121541929A