Memory access security check method for GPU compiler

By performing memory access security checks on low-level IR instructions in the backend of the GPU compiler, the problem of memory access security vulnerabilities in domestic GPU compilers is solved, the detection of memory out-of-bounds access and buffer overflow is realized, and the system stability and security are improved.

CN120610908AActive Publication Date: 2025-09-09WUHAN LINGJIU MICROELECTRONICS CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202511121712.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-09
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Domestic GPU compilers are prone to security vulnerabilities such as buffer overflow and illegal address access under complex thread scheduling and address calculation logic, which may lead to system crashes or security risks.

Method used

Memory access security checks are performed on low-level IR instructions in the GPU compiler backend. Memory allocation boundaries and runtime information are extracted through metadata information. The range analysis method of transfer functions is used to perform data flow analysis, mark potential memory access risks, and generate a security analysis report.

Benefits of technology

It improves the efficiency of checking memory access security, can discover problems caused by the compiler's back-end optimization stage, is versatile and efficient, avoids path explosion, and improves operational efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610908A_ABST
    Figure CN120610908A_ABST
Patent Text Reader

Abstract

The invention provides a memory access security checking method for a GPU compiler. The memory access security checking method comprises the following steps: converting a GPU code into a high-level IR instruction; according to metadata information in the high-level IR instruction, extracting a boundary allocated by each memory, and converting the high-level IR instruction into a low-level IR instruction; and based on a range analysis method of a transfer function, in the rear end of the compiler, performing data stream analysis on the low-level IR instruction, calculating a symbolized address range of each memory access, and checking the security of each memory access based on the symbolized range of each memory access and the boundary of each memory allocation. The method for performing memory access security check on the low-level IR language level in the GPU compiler is used for checking the problems of memory cross-border access or buffer overflow and the like which possibly occur after the rear end of the compiler is subjected to multiple optimizations, does not depend on a specific hardware architecture or a specific IR language, and has certain universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of GPU compilers, and more particularly to a memory access security checking method for GPU compilers. Background Art

[0002] In recent years, with continued investment in high-performance computing hardware, domestically produced GPU chips have made significant progress. These products have been widely used in scenarios such as AI training, graphics rendering, and scientific computing. Driven by large language models and intelligent computing platforms like DeepSeek, demand for parallel computing power has skyrocketed. GPUs, with their thousands of computing cores and exceptional throughput, have become the core support for the domestic AI computing ecosystem.

[0003] However, the high degree of parallelism in GPU architectures also presents significant challenges in software development and operational security. Memory security is an increasingly prominent issue. GPU kernel programs often need to be executed concurrently across tens of thousands of thread instances, which frequently access global memory, local shared memory, and private thread memory. Complex thread scheduling and address calculation logic can easily lead to security vulnerabilities such as buffer overflows and illegal address accesses, which can lead to unstable model training results, system crashes, and even security risks. Summary of the Invention

[0004] In response to the technical problems existing in the prior art, the present invention provides a memory access security checking method for a GPU compiler, which overcomes the problem that security vulnerabilities such as buffer overflow and illegal address access are very likely to occur under complex thread scheduling and address calculation logic, thereby leading to system crashes and even security risks.

[0005] The present invention provides a memory access security checking method for a GPU compiler, comprising: Step 1: Convert the GPU code into high-level IR instructions using a standard tool chain, wherein the high-level IR instructions include metadata information; Step 2: extract and symbolize the boundary and runtime information of each memory allocation based on the metadata information in the high-level IR instruction, and convert the high-level IR instruction into a low-level IR instruction; Step 3: Based on the range analysis method of the transfer function, in the compiler backend, data flow analysis is performed on the low-level IR instructions, the symbolic address range of each memory access is calculated, and the security of each memory access is checked based on the boundaries of each memory allocation and runtime information in the metadata information; Step 4: Mark memory accesses with security risks and generate a security analysis report.

[0006] The present invention provides a memory access security checking method for a GPU compiler. The method converts GPU code into high-level IR instructions; extracts the boundaries of each memory allocation based on metadata information in the high-level IR instructions, and converts the high-level IR instructions into low-level IR instructions; and performs data flow analysis on the low-level IR instructions in the compiler backend based on a range analysis method based on a transfer function, calculates the symbolic address range of each memory access, and checks the security of each memory access based on the symbolic range of each memory access and the boundaries of each memory allocation. The present invention performs a memory access security checking method at the low-level IR language level in a GPU compiler, and is used to check for problems such as memory out-of-bounds access or buffer overflow that may occur after multiple optimizations in the compiler backend. The method is independent of specific hardware architecture and specific IR language, and has a certain degree of versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 A flow chart of a memory access security checking method for a GPU compiler provided in one embodiment of the present invention; Figure 2 A flow chart showing register status information update according to an embodiment of the present invention; Figure 3 This is a flowchart of memory access out-of-bounds checking according to an embodiment of the present invention. DETAILED DESCRIPTION

[0008] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention can be arbitrarily combined with each other to form a feasible technical solution. This combination is not restricted by the sequence of steps and / or structural composition mode, but must be based on the ability of ordinary technicians in this field to implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0009] A GPU compiler is a special software program that converts high-level code (such as CUDA C++, OpenCL, HLSL) written by programmers to run on a GPU (graphics processing unit) into low-level instructions (machine code) that the GPU hardware can directly understand and execute.

[0010] In detail: The GPU compiler pipeline is a multi-stage process that converts high-level parallel code (such as CUDA C++ or GLSL) into machine code executable by specific GPU hardware. This process begins with the front-end, which parses the source code and generates a universal, hardware-independent intermediate representation (IR), such as LLVM IR or SPIR-V. This IR then enters the optimization phase, where a series of hardware-independent optimization passes (such as function inlining and dead code elimination) are executed to improve the general efficiency of the code. Next, on the back-end, this universal IR is translated into a vendor-specific IR (such as NVIDIA's PTX). At this point, a series of highly hardware-dependent back-end passes come into play, performing instruction scheduling to hide memory latency, performing sophisticated register allocation to improve GPU utilization, coalescing memory accesses to maximize bandwidth, and handling branch divergence. Finally, the fully optimized vendor IR is compiled into final, GPU-specific native machine code (such as SASS). This final compilation step is typically performed by the GPU driver during runtime (Just-in-Time) to ensure optimal performance on the user's specific hardware.

[0011] Why do we need a dedicated GPU compiler? This stems from the fundamental architectural differences between CPUs and GPUs: CPU (Central Processing Unit): Low latency and high single-core performance. Excels at handling complex logic, branching, and serial tasks. They have a small number of cores (from a few to dozens), but each core is very powerful, with a complex control unit and large cache.

[0012] GPU (Graphics Processing Unit): High-throughput, massively parallel computing. It excels at processing large numbers of similarly structured, simple computing tasks that can be performed simultaneously. It has hundreds or thousands of simple computing cores, but a relatively simple control unit and cache.

[0013] To improve program efficiency, domestic GPU compilation toolchains perform numerous optimization passes on the backend of the GPU compiler to optimize the IR code. Problems can arise during the compiler's source code parsing, front-end code conversion, and back-end optimization passes, leading to errors in the compiled program. Problems caused by memory access are particularly serious, easily causing system instability and making it difficult to trace the root cause.

[0014] In order to analyze and check the problems caused by memory access, the present invention proposes a method for performing memory security analysis on low-level IR in the back end of a GPU compiler.

[0015] See also Figure 1The present invention provides a memory access security checking method for a GPU compiler, comprising the following steps: Step 1: Use a standard tool chain to convert GPU code into high-level IR instructions, where the high-level IR instructions include metadata information.

[0016] An intermediate representation (IR) is a data structure or code used by compilers to represent source code. It represents the program between the source language and the target language during the compilation process. Almost all compilers require some form of IR to model the code being analyzed, transformed, and optimized. During the compilation process, the IR must be sufficiently expressive to accurately represent the source code without losing information, and must fully consider the completeness of the compilation from source to target code, the ease of use of compilation optimizations, and the performance.

[0017] The step 1, converting the GPU code into high-level IR instructions using a standard tool chain, includes: Step 11: Use a standard compiler front-end like Clang to convert the GPU code into high-level IR instructions (such as SPIR-V). The high-level IR instructions already contain the core metadata required for subsequent processing. Step 12: Create a mapping from high-level IR instructions to original source code files and line numbers as debugging information to be used later.

[0018] Step 2: Extract and symbolize the boundary and runtime information of each memory allocation based on the metadata information in the high-level IR instruction, and convert the high-level IR instruction into a low-level IR instruction.

[0019] It is understandable that after the GPU source code is converted into high-level IR instructions, metadata is extracted from the high-level IR instructions, including memory object metadata extraction and runtime information metadata extraction.

[0020] Wherein, step 2 includes: Step 21: Extracting Memory Object Metadata: Based on the metadata information in the high-level IR instructions, a unique identifier is created for each memory allocation in the code, recording its boundary size. Different identifiers are used to record the boundary size for different memory space types. For pointers to the unknown global address space, a symbol (such as size(@buffer_A)) is used to represent the boundary. For pointers to the known local address space and private address space, the exact byte size (such as 1024) is recorded.

[0021] Among them, different types of memory spaces are introduced below.

[0022] Global Memory: Global memory holds a large amount of memory. Its access speed is slower than other memories such as shared memory and registers. All running threads can read and write to global memory. The CPU can also read and write to global memory. Global memory is allocated and released by the host.

[0023] Constant Memory: This memory is also part of the GPU main memory, it has its own cache and is independent of the L1 and L2 of the global memory. All threads can access the same constant memory, but they can only read, not write. The CPU sets the value in the constant memory before launching the kernel.

[0024] Shared Memory: Shared memory is used to implement fast communication between threads in a thread block. Shared memory exists only during the lifetime of the block.

[0025] Local Memory: Local memory is part of the GPU's main memory (like global memory) and is therefore typically slower. When registers are exhausted or unavailable, threads automatically use local memory. This is called register spilling. This can occur if there are too many variables per thread to fit in registers or if the kernel uses structures. Arrays not indexed with constants also use local memory, as registers have no addresses and must use addressable memory. Local memory is scoped to each thread.

[0026] Step 22: Runtime metadata extraction: Based on the metadata in the high-level IR instructions, an initial range is assigned to built-in functions such as get_global_id(). Typically, this range is symbolic, such as get_local_id(0) , which is in the range [0, get_local_size(0) - 1]. When additional constraints are imposed, such as required_work_group_size(128, ...), this information is used to refine the range to a constant, such as getting_local_id(0) to a more precise range of [0, 127].

[0027] Step 23: After extracting metadata information from the high-level IR instruction, convert the high-level IR instruction into a low-level IR instruction.

[0028] Step 3: Based on the range analysis method of the transfer function, in the compiler backend, data flow analysis is performed on the low-level IR instructions, the symbolic address range of each memory access is calculated, and the security of each memory access is checked based on the boundaries of each memory allocation and runtime information in the metadata information.

[0029] In the compiler world, a pass is an independent, modular processing unit that traverses the intermediate representation (IR) of a program and performs a specific task, which can be: (1) Analysis Pass: Checks the code and collects information without modifying the code. For example, a pass might be used to analyze which variables in a loop are no longer used outside the loop.

[0030] (2) Transform Pass: Modify the code based on analysis results or specific rules. For example, a pass may be used to remove the calculation of variables that are no longer used outside the loop (dead code elimination).

[0031] In an embodiment of the present invention, a data flow analysis pass is added to the compiler backend process. This pass performs data flow analysis on the low-level IR in step S2, and its core concept is range-based pointer analysis.

[0032] In one embodiment of the present invention, step 3 specifically includes the following steps: Step 31: define register status information.

[0033] Registers are the fastest memory on the GPU. Variables declared in a kernel will use registers unless they run out or they cannot be stored in registers, then local memory will be used.

[0034] In order to complete the analysis of data flow, the register status information is defined as follows: State(%reg) = { BasePointerID, AddressSpace, OffsetRange} in: State(%reg): indicates the status information of register %reg BasePointerID: represents a unique identifier, indicating the pointer base address.

[0035] AddressSpace: Indicates the memory space type, including private, global, local and NonPointer.

[0036] OffsetRange: Expressed as a pointer offset [min_offset, max_offset], which is the minimum and maximum byte offset of the pointer relative to BasePointerID.

[0037] Step 32: Obtain the metadata information, the low-level IR instructions of the current state in the compiler backend, and the corresponding control flow graph CFG, perform data flow analysis based on the control flow graph CFG, and update the register state information.

[0038] In one embodiment, step 32 specifically includes the following steps: Step 321 , for each basic block B in the control flow graph CFG, allocate storage space to record its entry state IN[B] and exit state OUT[B].

[0039] For each basic block in the control flow graph (CFG), storage space is allocated to record its entry state (IN[B]) and exit state (OUT[B].) For the entry block of the control flow graph (CFG), its IN state is set to an initial condition, which represents prior knowledge before kernel or function execution. For all other basic blocks, their IN and OUT states are initialized to the "top" state (Top Element), representing the state with the least information or the most uncertainty.

[0040] Step 322: Create a work list, where the work list is used to store the data structure of the basic block to be processed.

[0041] Step 323: After initialization, the entry basic block in the control flow graph CFG is put into the work list.

[0042] It is understandable that a worklist is created. The worklist is a data structure for storing basic blocks to be processed. Initially, the entry basic block in the control flow graph CFG is added to the worklist to start the iteration process.

[0043] Step 324: Take out a basic block currently to be processed from the work list, and record it as B_current.

[0044] Step 325 , calculate a new entry state IN[B_current] for B_current.

[0045] The entry state of the basic block is calculated, which is the collection of the exit state information of all its predecessor nodes. For the selected basic block B_current, its entry state IN[B_current] is calculated as follows: In the control flow graph (CFG), identify all basic blocks that directly point to B_current, i.e., its predecessor node set Predecessors(B_current). Apply a predefined meet / join operator to the exit states OUT[P] of all predecessor nodes P ∈ Predecessors(B_current). This operator merges data flow information from multiple input paths into a single state. Its mathematical representation is: IN[B_current] = Meet(OUT[P_1], OUT[P_2], ..., OUT[P_i], ..., OUT[P_n]); Where P_i ∈ Predecessors(B_current), n is the number of predecessor basic blocks of B_current.

[0046] It's important to note that, for example, when analyzing an if-else branch structure, traditional methods split the analysis into two independent paths to explore separately. If the branches are nested, the number of paths increases exponentially. However, the present invention performs a state merge (Join) operation at the branch's junction. If the range of a variable x after the if branch is [10, 10] and after the else branch is [100, 100], the merged state is [10, 100]. Subsequent analysis can then continue based solely on the more universal range of [10, 100], eliminating the need to track two paths. This approach, which covers multiple possibilities with a single abstract state, fundamentally avoids the exponential growth of the number of paths. This ensures that even with complex control flows (such as multiple branches and loops), analysis can be completed efficiently and within a predictable timeframe, avoiding the path explosion associated with traditional methods and improving operational efficiency.

[0047] Step 326 , based on the new inlet state IN[B_current] of B_current, calculate the new outlet state OUT[B_current] for B_current, and update the register state information.

[0048] It is understood that the exit state of a basic block is obtained by processing the instructions in the basic block according to the entry state of the basic block. During the execution of the instructions in the basic block, the register state is updated.

[0049] Specifically, all instructions in the basic block B_current are traversed and processed, and the exit state OUT[B_current] of B_current is calculated. For each instruction type in the basic block B_current, a transfer function is defined to complete the register state update. The workflow of defining a transfer function to complete the register state update according to each instruction type in the basic block can be found in Figure 2 .

[0050] In addition to memory access instructions, each instruction needs to define a corresponding transfer function to complete the update of the register status. The pseudo instructions and status update methods of typical instructions are shown in Table 1.

[0051] Table 1 Typical instructions and status update methods

[0052] Table 1 shows different types of instructions and how they update register status.

[0053] Step 327 , check whether the calculated new exit state of B_current causes a state change of a subsequent basic block. If so, store the data structure of the subsequent basic block of B_current into the work list, and remove B_current from the work list.

[0054] It is understood that after calculating the new exit state of the current B_current block, it is determined whether the new exit state of B_current causes a state change for subsequent basic blocks. If so, the data structure of the subsequent basic blocks of B_current is stored in the worklist to facilitate the subsequent calculation of the new entry state and new exit state of the subsequent basic blocks in the worklist. After calculating the new entry state and new exit state of the current B_current basic block, the B_current basic block is removed from the worklist.

[0055] Step 33: When a memory access instruction is encountered, current register status information is obtained.

[0056] Step 34, obtaining the corresponding memory allocation boundary according to the base address and memory space type in the current register state information; Step 35: Determine whether the pointer offset is out of bounds based on the pointer offset in the current register state information and the corresponding memory allocation boundary.

[0057] It is understandable that when a memory access instruction (such as load / store) is encountered during the execution of each instruction of a basic block, a security check is performed. The process of security check for memory access is as follows: Figure 3shown.

[0058] Specifically, the current status information of the register is obtained to obtain {BasePointerID, AddressSpace, OffsetRange}, where OffsetRange = [min_offset, max_offset].

[0059] According to the base address BasePointerID and memory space type, find the corresponding memory allocation boundary information [0, alloc_size] from the metadata information. Perform checks based on the information obtained at this time: (1) Check the lower bound of the pointer offset: whether min_offset >= 0 is satisfied. If so, the lower bound is within the range.

[0060] If min_offset < 0 and min_offset is a negative constant (such as -1), an out-of-bounds error is reported; if min_offset depends on the symbol (such as sym_id_10), a potential out-of-bounds warning is reported.

[0061] (2) Check the upper bound of the pointer offset: whether max_offset + access_size <= alloc_size is satisfied. If so, the upper bound is not exceeded.

[0062] For private and local memory, the memory bounds (alloc_size) are compile-time constants. You can directly calculate and compare them: for example, if max_offset + access_size > alloc_size, an explicit out-of-bounds error will be reported.

[0063] If global memory is used, alloc_size is a symbol, such as size(@global_buffer). The check is a proof of a symbolic inequality. For example, max_offset might be (global_size_0 - 1) × 4. Check that (global_size_0 - 1) × 4 + 4 <= size(@buffer_A), that is, global_size_0 × 4 <= size(@buffer_A). If the compiler can prove that this inequality always holds (which it may in some simple cases), it is safe. Otherwise, a conditional warning is issued: "Out-of-bounds access may occur when startup parameters do not satisfy global_work_size[0] × 4 <= size_of_buffer_A." If the lower and upper bounds are not exceeded after checking, the address space is checked. The memory access instruction itself usually specifies the address space, and the address range is calculated based on the pointer base address and pointer offset in the register state. For example, the current register state information is {@local_buf_a, local_address, [0, 128]}, where @local_buf_a is the pointer base address and [0, 128] is the pointer offset. The address range is [local_buf_a, local_buf_a + 128].

[0064] The calculated address range is matched with the specified address space range. If it does not match, it indicates that the pointer is misreferenced. The address space ranges of different types are determined at compile time (the ranges of private and local type address spaces are determined at compile time).

[0065] For example, if the local address space is specified by the compiler as [2048, 4096], all local type pointers must be within this range. If the calculated address range of the pointer is [2048, 6144], it may exceed the specified address space range of the local type, which is not a complete match and poses a security risk.

[0066] During the data flow analysis process, it is determined whether the data flow analysis has reached convergence, that is, a stable state. The specific determination method is: For each basic block in the control flow graph CFG, the calculated new exit state OUT_new[B_current] is compared with the previously stored old exit state OUT_old[B_current]; If OUT_new is different from OUT_old, it indicates that the data flow information has changed and the data flow analysis has not yet converged; Update the exit state of B_current to OUT_new, and add the data structures of all successor nodes of B_current in the control flow graph CFG to the work list; If OUT_new is the same as OUT_old, no action is performed; When the work list is empty, it means that the states of all basic blocks have stabilized and no longer change, the system has reached the fixed point, and the iterative process terminates.

[0067] Step 4: Mark memory accesses with security risks and generate a security analysis report.

[0068] According to the security check results of the memory access in the above steps, a security analysis report is generated. Specifically, the steps include: In step 41 , all instances marked as risks and their related context data are extracted from the analyzed code.

[0069] In step 42, each individual risk instance is assigned a stable and unique ID for continued tracking, and is classified into high (deterministic error), medium (potential risk), or low (information) levels based on the certainty of the risk.

[0070] In step 43, the risk points in the low-level IR are accurately mapped back to their file names and line numbers in the high-level source code, and a concise and clear text is programmatically generated to describe the nature of the risk and potential triggering conditions.

[0071] In step 44, all generated risk records are compiled into a structured file (such as JSON or SARIF format) and output.

[0072] The memory access security checking method for a GPU compiler provided by the present invention has the following advantages: (1) The analysis process of the present invention directly acts on the private low-level IR after being processed by the back-end of the domestic GPU compiler. The object of analysis is the final code that is closer to the actual execution of the hardware. Compared with all existing technologies that perform analysis at the high-level IR level (such as LLVM IR), the present invention can discover and report memory security issues that are introduced or exposed at the back-end optimization stage of the domestic GPU manufacturer's compiler, eliminating the analysis blind spots of the existing technology.

[0073] (2) This invention captures the symbolic boundaries of memory objects and key variables of the execution environment through a systematic metadata extraction and symbolization mechanism before analysis. During security checks, an inequality consisting of symbols is constructed based on the precise context to make judgments, improving the accuracy of judgments.

[0074] (3) The core algorithm of this invention uses symbolic scope analysis rather than heavyweight, path-by-path exploration. When encountering branches or loops, the states of multiple paths are merged into a unified abstract scope, avoiding the exponential growth of the number of paths. This enables the analysis of complex control flows to be completed in a predictable and reasonable time, with high operational efficiency.

[0075] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0076] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0077] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0078] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0080] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0081] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A memory access security checking method for a GPU compiler, characterized in that: include: Step 1: Convert the GPU code into high-level IR instructions using a standard tool chain, wherein the high-level IR instructions include metadata information; Step 2: extract and symbolize the boundary and runtime information of each memory allocation based on the metadata information in the high-level IR instruction, and convert the high-level IR instruction into a low-level IR instruction; Step 3: Based on the range analysis method of the transfer function, in the compiler backend, data flow analysis is performed on the low-level IR instructions, the symbolic address range of each memory access is calculated, and the security of each memory access is checked based on the boundaries of each memory allocation and runtime information in the metadata information; Step 4: Mark memory accesses with security risks and generate a security analysis report.

2. The memory access security checking method according to claim 1, wherein: Step 1, converting GPU code into high-level IR instructions using a standard tool chain, includes: Step 11: Convert the GPU code into high-level IR instructions based on the standard compiler front end, wherein the high-level IR instructions include core metadata; Step 12: Establish a mapping from the high-level IR instructions to GPU source code files and line numbers.

3. The memory access security checking method according to claim 1, wherein: The step 2, extracting and symbolizing the boundaries and runtime information of each memory allocation based on the core metadata in the high-level IR instruction, includes: Step 21, memory object metadata extraction: Based on the metadata information in the high-level IR instructions, a unique identifier is created for each memory allocation in the code. The unique identifier is used to record the boundary size of each memory allocation. For pointers to the global address space of unknown size, a symbol is used to represent its boundary; for pointers to the local address space and private address space of known size, the boundary is represented by the exact byte size. Step 22, runtime information metadata extraction: Based on the metadata information in the high-level IR instruction, an initial value range is assigned to each built-in function, where the instruction range is symbolic. When there are additional constraints, the value range is refined into a constant based on the constraint information; Step 23: After completing the metadata information extraction, convert the high-level IR instruction into a low-level IR instruction.

4. The memory access security checking method according to claim 1, wherein: The step 3, based on the range analysis method of the transfer function, performs data flow analysis on the low-level IR instructions in the compiler backend, calculates the symbolic address range of each memory access, and checks the security of each memory access based on the boundaries of each memory allocation and runtime information in the metadata information, including: Step 31, define register state information, which is expressed as State(%reg) ={BasePointerID, AddressSpace, OffsetRange}, where State(%reg) represents the state information of register %reg; BasePointerID represents a unique identifier, indicating the pointer base address; AddressSpace represents the memory space type, which has private, global, local, and NonPointer types; OffsetRange represents the pointer offset [min_offset, max_offset], which is the minimum and maximum byte offset of the pointer relative to BasePointerID; Step 32: Obtain the metadata information, the low-level IR instructions of the current state in the compiler backend, and the corresponding control flow graph CFG, perform data flow analysis based on the control flow graph CFG, and update the register state information; Step 33, when encountering a memory access instruction, obtaining current register state information; Step 34, obtaining the corresponding memory allocation boundary according to the base address and memory space type in the current register state information; Step 35: Determine whether the pointer offset is out of bounds based on the pointer offset in the current register state information and the corresponding memory allocation boundary.

5. The memory access security checking method according to claim 4, characterized in that: The step 32, obtaining the metadata information, the low-level IR instructions of the current state in the compiler backend and the corresponding control flow graph CFG, performing data flow analysis based on the control flow graph CFG, and updating the register state information, includes: Step 321 , for each basic block B in the control flow graph CFG, allocate storage space to record its entry state IN[B] and exit state OUT[B]; Step 322: Create a work list, wherein the work list is used to store the data structure of the basic block to be processed. Step 323: After initialization, the entry basic block in the control flow graph CFG is put into the work list; Step 324: Take out a basic block currently to be processed from the work list, denoted as B_current; Step 325, calculate a new entry state IN[B_current] for B_current; Step 326 , based on the new inlet state IN[B_current] of B_current, calculate the new outlet state OUT[B_current] for B_current, and update the register state information; Step 327 , check whether the calculated new exit state of B_current causes a state change of a subsequent basic block. If so, store the data structure of the subsequent basic block of B_current into the work list, and remove B_current from the work list.

6. The memory access security checking method according to claim 5, characterized in that: The step 325, calculating a new entry state IN[B_current] for B_current, includes: In the control flow graph CFG, all basic blocks directly pointing to B_current are identified, that is, the predecessor basic block set of B_current, Predecessors(B_current); For the exit state OUT[P] of all predecessor basic blocks P∈Predecessors(B_current), a predefined confluence operator Meet / Join Operator is applied. This operator merges the data flow information on multiple input paths into a single state. Its mathematical representation is: IN[B_current] = Meet(OUT[P_1], OUT[P_2], ..., OUT[P_i], ..., OUT[P_n]); Where P_i ∈ Predecessors(B_current), n is the number of predecessor basic blocks of B_current.

7. The memory access security checking method according to claim 5, characterized in that: Step 326, based on the new inlet state IN[B_current] of B_current, calculate the new outlet state OUT[B_current] of B_current and update the register state information, including: Traverse and process all instructions in B_current, calculate the new exit state OUT[B_current] of B_current, and for each instruction type in B_current, complete the update of the register state by defining the transfer function.

8. The memory access security checking method according to claim 4, wherein: The step 35, judging whether the pointer offset is out of bounds based on the pointer offset in the current register state information and the corresponding memory allocation boundary, includes: Get the pointer offset [min_offset, max_offset] in the current register state information, where the corresponding memory allocation boundary obtained from the metadata information is recorded as [0, alloc_size]; Check the lower bound min_offset and the upper bound max_offset of the pointer offset respectively; Among them, the lower bound min_offset is checked, including: If min_offset >= 0, the lower bound is not exceeded; If min_offset < 0, and min_offset is a negative constant, it is a clear lower bound error; If it is not entirely certain that min_offset >= 0, and min_offset depends on the sign, it is a potential lower bound error; Check the upper bound max_offset, including: If max_offset + access_size <= alloc_size, then perform address space check; If max_offset + access_size > alloc_size, and max_offset + access_size is a constant, then it is a clear upper bound out of bounds error; If max_offset + access_size > alloc_size, and max_offset + access_size depends on the sign, it is a potential upper bound error; The address space check includes: The address range is calculated based on the pointer base address and pointer offset in the current register state information, and the address range is matched with the specified address space range. If there is a mismatch, it indicates that the pointer is incorrectly referenced, where the specified address space range of different types of memory space is determined at compile time.

9. The memory access security checking method according to claim 5, wherein: Step 327 checks whether the calculated new exit state of B_current causes a state change of a subsequent basic block. If so, stores the data structure of the subsequent basic block of B_current in the work list and removes B_current from the work list. The step also includes determining whether the data flow analysis has reached convergence. If so, the data flow analysis ends. Determining whether the data flow analysis has reached convergence includes: For each basic block in the control flow graph CFG, the calculated new exit state OUT_new[B_current] is compared with the previously stored old exit state OUT_old[B_current]; If OUT_new is different from OUT_old, it indicates that the data flow information has changed and the data flow analysis has not yet converged; Update the exit state of B_current to OUT_new, and add the data structures of all successor nodes of B_current in the control flow graph CFG to the work list; If OUT_new is the same as OUT_old, no action is performed; When the work list is empty, the states of all basic blocks are stable and the data flow analysis converges.

10. The memory access security checking method according to claim 5, wherein: Step 4, marking memory accesses with security risks and generating a security analysis report, includes: Step 41 , extracting all memory access instances marked as risky and their related context data from the analyzed low-level IR instructions; Step 42 , assigning a stable and unique ID to each independent risky memory access instance, and classifying it as high risk, medium risk, or low risk level based on the certainty of the risk; Step 43 , mapping the risk points in the low-level IR instructions back to their file names and line numbers in the high-level GPU source code, and programmatically generating description text to explain the nature of the risk and potential triggering conditions; In step 44, all generated risk records are compiled into a structured file and output.

Citation Information

Patent Citations

  • Method for dynamically detecting memory overflow on GPU based on address compression technology

    CN107908954A

  • Method for detecting memory boundary overflow errors

    CN108197035A

  • Method and device for detecting unsecure direct memory access in driver

    CN112925524A

  • Symbol analysis method and device, equipment and storage medium

    CN113157731A

  • Nonvolatile memory check point generation method and device and electronic equipment

    CN113515412A