GPGPU hidden instruction mining method based on semi-supervised adaptive walk

Through the semi-supervised adaptive wandering method, the GPGPU instruction space is parsed and compressed, the key fields are extracted and the instructions to be tested is assembled, and the existing technology is difficult to explore GPGPU hidden instructions is realized, and the GPGPU security situation is effectively monitored and controlled, providing reliable security guarantees for high-performance parallel computing.

CN120088118APending Publication Date: 2025-06-03UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161528.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to quickly, efficiently and reliably mine hidden instructions in general graphics processors (GPGPUs), resulting in the security and trusted computing of high-performance parallel computing being compromised.

Method used

The semi-supervised adaptive wandering method is adopted to analyze the encoding format and composition structure of GPGPU instructions by flipping the instruction bits and reversing the bits; the semi-supervised dimensionality reduction algorithm is used to compress the instruction redundant space, extract key fields and assemble the instructions to be tested, and search and identify hidden instructions through adaptive wandering.

Benefits of technology

It realizes rapid and effective mining of hidden GPGPU instructions, improves the observability and controllability of the GPGPU security situation, and provides reliable security guarantees for high-performance parallel computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088118A_ABST
    Figure CN120088118A_ABST
Patent Text Reader

Abstract

The invention discloses a GPGPU hidden instruction mining method based on semi-supervised adaptive walk, and relates to the field of processor security. According to the method, the problem that hidden instructions cannot be quickly, effectively and reliably mined due to the fact that a GPGPU instruction rule structure is complex in the prior art is solved. Compared with the prior art, the method starts from the two aspects of semi-supervised compression instruction space and instruction self-adaptive variable-step walk search, deep analysis and research are conducted on the instruction set micro-architecture of the GPGPU, the hidden instruction vulnerabilities of the GPGPU are rapidly and effectively mined, GPGPU processor manufacturers can be helped to repair the hidden instruction vulnerabilities, and the hidden instruction vulnerabilities of the GPGPU are effectively found. And meanwhile, a reliable technical means is provided for a user to detect whether the application program has instruction vulnerabilities or not and for GPGPU high-performance parallel operation analysis and security evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of processor security. Aiming at the problem of hidden instructions in general-purpose graphics processors, an effective semi-supervised adaptive walk search and recognition method is proposed, providing a reliable and secure technical means for the application of general-purpose processors in high-performance parallel computing. Background Art

[0002] A processor is the arithmetic control core for information processing and program execution in a computer system. When designing a processor, processor manufacturers may add some instructions for testing for the convenience of testing and debugging. These instructions are hidden and not publicly available to users and developers, and these instructions pose a backdoor risk. At the same time, the incompleteness in the design and implementation of the processor instruction decoding module may also lead to the introduction of hidden instructions. This type of instruction is uncertain in function and behavior, and may cause processor anomalies such as calculation errors, privilege escalation, denial of service, and data leakage. For example: In 1994, the FDIV instruction vulnerability in the Pentium processor caused errors in floating-point division; in 1997, the F00F exception instruction in the Pentium processor caused the CPU to deny service; in 2018, at the BlackHat security conference, there was a backdoor hardware problem in the form of an undisclosed instruction in the Taiwan VIA C3 processor, enabling ordinary users to obtain superuser privileges. Therefore, the trusted computing of processors caused by hidden instructions is a hardware security issue that the industry pays increasing attention to.

[0003] With the continuous development of emerging fields such as video image processing, autonomous driving, and artificial intelligence, general-purpose graphics processors (GPGPUs) are widely used due to their advantages such as high data throughput and high-speed parallel computing. However, current existing hidden instruction analysis technologies such as Sandsifter, CPU Security Benchmark, iScanU, Armshaker, DSPUIM, etc. mainly target X86, ARM, and DSP processors. In the analysis of hidden instructions in GPGPUs for high-performance parallel computing, there are still huge technical challenges such as complex instruction rule structures, high instruction execution overheads, and insufficient observability of the GPGPU security situation. There is an urgent need to develop an effective, reliable, and adaptive search and recognition analysis method for GPGPU hidden instructions. Summary of the Invention

[0004] The present invention proposes a method for mining hidden instructions of a general - purpose graphics processing unit based on semi - supervised adaptive walk. This method solves the problem that the prior art cannot quickly, effectively, and reliably mine hidden instructions due to the complex rule structure of GPGPU instructions. Compared with the prior art, this method starts from two aspects: semi - supervised compressed instruction space and instruction adaptive variable - step walk search, deeply analyzes and studies the instruction set micro - architecture of GPGPU, realizes the rapid and effective mining of GPGPU hidden instruction vulnerabilities, helps GPGPU processor manufacturers repair hidden instruction vulnerabilities, and at the same time detects whether there are instruction vulnerabilities in application programs for users, providing a reliable technical means for GPGPU high - performance parallel computing analysis and security evaluation.

[0005] Therefore, the technical solution of the present invention is a method for mining GPGPU hidden instructions based on semi - supervised adaptive walk, and this method includes:

[0006] Step 1: Instruction bit flipping and parsing the composition structure;

[0007] Flip the bit positions of the GPGPU instruction from high to low in turn, and then parse the encoding format of the instruction to obtain the instruction bit - field of each component part; The GPGPU instruction is composed of an opcode, an operand, a reserved bit, and a predicate, and its bit length is L; Denote the GPGPU instruction as I, and the bit - fields of the opcode, operand, reserved bit, and predicate are I op 、I or 、I de and I pr , as shown in (1):

[0008] I = [I op ,I or ,I de ,I pr (1)

[0009] Set a binary one - hot mask M with the same bit length, where the first bit is 1 and the rest of the bits are 0, and the initial value of M is Perform an exclusive - OR operation on the instruction I and the mask M bit - by - bit to obtain the flipped instruction As shown in formula (2)

[0010]

[0011] where, represents the exclusive - OR operation symbol;

[0012] Step 2: Reverse - identify bits and update and shift the mask;

[0013] Assume that the regular assembly instruction text corresponding to the instruction I is T, which is composed of an opcode, an operand, and a predicate, as shown in formula (3):

[0014] T = [op, or, pr] (3)

[0015] Then the flip instruction The corresponding instruction text is As shown in Equation (4):

[0016]

[0017] Compare the instruction text T and When the opcode operand or predicate changes, then this bit belongs to the bit field of the corresponding opcode, operand or predicate; if there is no change, then this bit belongs to the reserved bit; shift the mask M one bit to the right logically and perform the exclusive OR operation with the instruction I again, for a total of L loops;

[0018] Step 3: Extract the key fields and compress the redundant space;

[0019] Use semi-supervised dimensionality reduction to compress the instruction redundant space and separate the high-value search space; assume that I op has a bit width of L op , then the semi-supervised compressed instruction space S reduced is:

[0020]

[0021] Step 4: Assemble the instruction to be tested and analyze the disassembly status;

[0022] At this time, the length of the compressed instruction space is only related to I op , and the remaining instruction bit fields I or , I de and I pr are all set to 0; according to the instruction encoding rule of Equation (1), assemble them into an instruction to be tested I T , as shown in Equation (6);

[0023] I T = [I op , 0, 0, 0] (6)

[0024] Perform disassembly analysis on the instruction to be tested I T and set the disassembly status code Q to judge whether it is a disassemblable instruction; if the disassembly status code Q = 0 and the instruction to be tested I T is a disassemblable instruction, then the instruction adaptively varies its step and walks; if the disassembly status code Q = 1 and the instruction to be tested I T is a non-disassemblable instruction, then save it as an instruction to be executed;

[0025] Step 5: Instruction Adaptive Walk, Saving Instructions to be Executed;

[0026] Increment the instruction counter k by 1. When it reaches the maximum count value k cnt_max reset the instruction counter k to 0, as shown in Equation (7);

[0027]

[0028] The walk step size λ is adaptively incremented by 1, as shown in Equation (8);

[0029] λ = λ + 1 λ ≤ λ step_max (8) When λ grows to the maximum step size λ step_max reset the instruction counter k, set λ adaptively to 1, and set an empty set E to save the instructions to be executed, as shown in Equation (9);

[0030]

[0031] Subsequently, update the walk position x. If x exceeds the walk boundary value, end the walk process, as shown in Equation (10);

[0032]

[0033] Step 6: Generate an Execution File, Identifying Hidden Instructions;

[0034] The computer host loads the instructions and data into its GPGPU global memory and starts the GPGPU to execute. The source code is compiled to generate an executable file, which is the medium for storing the instructions and data. From the set E of instructions to be executed generated in the previous step, sequentially take out the i-th instruction I i for the GPGPU to generate an executable file. Determine whether the instruction I i is a hidden instruction according to whether the executable file triggers an illegal instruction exception state.

[0035] Furthermore, the starting position of the walk in Step 5 is 0, and the walk boundary is

[0036] Furthermore, in Step 5, set the initial value of the walk step size λ to 1, and the maximum step size λ step_max to 5; set the instruction counter k to 0, and the maximum count value k cnt_max to 10.

[0037] Furthermore, the specific method of Step 6 is:

[0038] The executable file F consists of four parts: the header ehead, the section section, the section header table shead, and the program header table phead; among them, the section header table offset e_shoff in the header ehead indicates the position of the section header table; the section header table entry size e shentsize and the number of section header table entries e shnum , indicating the position of the table entry storing the instructions; the program segment offset sh_offset in the table entry is the final position addr where the instructions are stored, and the instruction I i is embedded at this position, as shown in Equation (11):

[0039]

[0040] If the GPGPU triggers an illegal instruction exception and returns a status code E equal to 715, then it is determined that the instruction I i is an illegal instruction and is discarded; if the returned status code E is not equal to 715, then it is determined that the instruction I i is a hidden instruction and is saved to the hidden instruction table H, as shown in Equation (12);

[0041] I i ∈H E≠715 (12)

[0042] Judge whether the instructions to be executed in the set E have been completed. If not, update and execute the next instruction to be executed; if completed, exit.

[0043] The present invention proposes a method for mining hidden instructions of a general-purpose graphics processing unit based on semi-supervised adaptive random walk. Through reverse-engineering the GPGPU instruction encoding rules, the present invention proposes a semi-supervised instruction space compression algorithm, which compresses the hidden instruction search space, effectively removes the redundant information in the instruction space, and improves the mining efficiency of hidden instructions; according to the characteristics of the hidden instruction distribution, the search efficiency of hidden instructions is improved through the instruction space adaptive variable-step random walk algorithm; a simple and feasible method for using the GPGPU to execute the instructions to be tested is used to reduce the instruction execution complexity and overhead; the invention effectively reveals the security vulnerabilities of GPGPU hidden instructions, improves the observability and controllability of the GPGPU security situation, and provides reliable technical support for high-performance parallel secure and trustworthy computing. Description of the Drawings

[0044] Figure 1 is the overall flowchart of the method for mining hidden instructions of GPGPU based on semi-supervised adaptive random walk.

[0045] Figure 2 is the flowchart of the instruction adaptive variable-step random walk algorithm.

[0046] Figure 3It is a flowchart for the execution and recognition of hidden instructions. Specific implementation mode

[0047] The present invention solves the problem that the existing methods cannot directly and effectively analyze hidden instructions for GPGPU. First, the bit flipping algorithm is used to reverse the instruction encoding rule of GPGPU to analyze the instruction composition structure of GPGPU; on this basis, the bit fields of non-reserved bits in the disassembly text are reversely identified and circularly logically shifted until the end bit of the instruction; subsequently, under the guidance of the GPGPU instruction rule experience, the key opcode fields of the instruction are extracted to semi-supervised compress the redundant space of the instruction; then, the instruction to be tested is assembled and its disassembly state is analyzed; further, the instruction space instruction adaptively wanders after compression to search for and identify potential hidden instructions as the instruction to be tested; finally, the instruction to be tested is injected into GPGPU for execution, and it is judged whether it is a hidden instruction according to the processor state. This method effectively solves the problems of low efficiency and high complexity in searching and identifying hidden instructions caused by the high redundant instruction space of GPGPU, improves the accuracy and efficiency of searching and identifying hidden instructions of GPGPU, and provides security guarantee for GPGPU parallel computing applications.

[0048] The specific method is as follows:

[0049] Step 1: Instruction bit flipping and composition structure analysis;

[0050] First, flip the bit of the GPGPU instruction from high to low in turn, and then analyze the encoding format of the instruction to obtain the instruction bit field of each component; the GPGPU instruction consists of an opcode, an operand, a reserved bit, and a predicate, and its bit length is L; denote the GPGPU instruction as I, and the bit fields of the opcode, operand, reserved bit, and predicate are I op , I or , I de and I pr , as shown in (1):

[0051] I = [I op , I or , I de , I pr (1)

[0052] Set a binary one-hot mask M with the same bit length, where the l-th bit is 1 and the rest of the bits are 0, and the initial value of M is Perform an exclusive OR operation on the instruction I and the mask M bit by bit to obtain the flipped instruction As shown in formula (2)

[0053]

[0054] Among them, represents the exclusive OR operation symbol;

[0055] Step 2: Reverse-identify bits, update and shift the mask;

[0056] The assembly instruction text corresponding to the GPGPU instruction's corresponding rule explains its function; assume the rule assembly instruction text corresponding to instruction I is T, which consists of an opcode, operands, and predicates, as shown in Equation (3):

[0057] T = [op, or, pr] (3)

[0058] Then the flip instruction The corresponding instruction text is As shown in Equation (4):

[0059]

[0060] Compare the instruction text T and When the opcode operands or predicates change, then the bit belongs to the bit field of the corresponding opcode, operands, or predicates; if there is no change, then the bit belongs to the reserved bit; logically shift the mask M one bit to the right and perform an exclusive OR operation with instruction I again, for a total of L loops;

[0061] Step 3: Extract key fields and compress redundant space;

[0062] In GPGPU instructions, the opcode determines the behavior and function of the instruction, which is the key information of the instruction and belongs to the high-value search space; the remaining bit fields have nothing to do with the function of the instruction and are the redundant information of the instruction, belonging to the low-value search space. Therefore, in this step, use semi-supervised dimensionality reduction to compress the instruction redundant space and separate the high-value search space. Assume the bit width of I op is L op , then the semi-supervised compressed instruction space S reduced is:

[0063]

[0064] Step 4: Assemble the instruction to be tested and analyze the disassembly status

[0065] At this time, the length of the compressed instruction space is only related to I op , and the remaining instruction bit fields I or , I de and I pr can all be set to the fixed value 0. According to the instruction encoding rule of Equation (1), assemble it into an instruction to be tested I T , as shown in Equation (6).

[0066] I T= [I op , 0, 0, 0] (6)

[0067] For the instruction I to be tested T perform disassembly analysis and set the disassembly status code Q to determine whether it is a disassembly - enabled instruction. If the disassembly status code Q = 0, the instruction I to be tested T is a disassembly - enabled instruction, then the instruction adaptively varies the step for walking; if the disassembly status code Q = 1, the instruction I to be tested T is a non - disassembly - enabled instruction, then save it as an instruction to be executed.

[0068] Step 5: Instruction adaptive walking, save the instruction to be executed

[0069] For the instruction I to be tested with the disassembly status code Q being 0 T perform adaptive variable - step walking, and the flowchart is as Figure 2 shown.

[0070] As can be seen from Step 4, the instruction I to be tested T is only related to I op , then I op is the walking instruction bit field, denoted by x. The starting position of walking is 0, and the walking boundary is Set the initial value of the walking step size λ to 1, and the maximum step size λ step_max is 5; set the instruction counter k to 0, and the maximum count value k cnt_max is 10.

[0071] Increment the instruction counter k by 1. When it reaches the maximum count value k cnt_max , reset the instruction counter k to 0, as shown in Equation (7).

[0072]

[0073] The walking step size λ adaptively increases by 1, as shown in Equation (8).

[0074] λ = λ + 1 s.t. λ ≤ λ step_max (8) When λ grows to the maximum step size λ step_max , reset the instruction counter k, adaptively set λ to 1, and set an empty set E to save the instruction to be executed, as shown in Equation (9).

[0075]

[0076] Subsequently, update the walking position x. If x exceeds the walking boundary value, then end the walking process, as shown in Equation (10).

[0077]

[0078] Step 6: Generate an executable file and identify hidden instructions

[0079] The computer host loads the instructions and data into its GPGPU global memory and starts the GPGPU to execute. The source code is compiled to generate an executable file, which is the medium for storing the instructions and data. From the set of instructions to be executed E generated in the previous step, the i-th instruction I is sequentially taken out i Used for the GPGPU to generate an executable file, and determine the instruction I according to whether the executable file triggers an illegal instruction exception status i Whether it is a hidden instruction, the flowchart is as Figure 3 shown

[0080] The executable file F consists of four parts: a header ehead, a section section, a section header table shead, and a program header table phead. Among them, the section header table offset e_shoff in the header ehead indicates the position of the section header table. The section header table entry size e_shentsize and the number of section header table entries e_shnum in the section header table shead indicate the position of the table entry storing the instructions. The program segment offset sh_offset in the table entry is the final position addr for storing the instructions. Embed the instruction I i into this position, as shown in Equation (11):

[0081]

[0082] If the GPGPU triggers an illegal instruction exception and returns a status code E equal to 715, then it is determined that the instruction I i is an illegal instruction and the instruction is discarded; if the returned status code E is not equal to 715, then it is determined that the instruction I i is a hidden instruction and is saved to the hidden instruction table H, as shown in Equation (12).

[0083] I i ∈H s.t. E≠715 (12)

[0084] Determine whether the instructions to be executed in the set E have been completed. If not, update and execute the next instruction to be executed; if completed, exit

[0085] This embodiment provides a method for mining GPGPU hidden instructions based on semi-supervised adaptive walk. In this embodiment, the specific model of the GPGPU is NVIDIA GP107, the experimental device is a heterogeneous computing platform Intel Core i5-12400F + GTX1050Ti, the operating system is Windows 10, the graphics card driver version is 535.146.02, and the CUDA version is 12.2. The following table shows the results of GPGPU hidden instruction mining

[0086] Table 1 Results of GPGPU Hidden Instruction Mining

[0087]

[0088]

[0089]

Claims

1. A GPGPU hidden instruction mining method based on semi-supervised adaptive walking, the method comprising: Step 1: Instruction bit flipping and analysis of the structure; Flip the bits of the GPGPU instruction from high to low, and then parse the encoding format of the instruction to obtain the instruction bit field of each component; the GPGPU instruction consists of an opcode, an operand, a reserved bit, and a predicate, and its bit length is L; the GPGPU instruction is denoted as I, and the bit field of the opcode, operand, reserved bit, and predicate is denoted as I op ,I or ,I de and I pr , as shown in (1): I=[I op ,I or ,I de ,I pr ] (1) Set a binary one-hot mask M with the same bit length, where the first bit is 1 and the rest of the bits are 0. The initial value of M is Perform bitwise XOR operation on instruction I and mask M to obtain the flipped instruction As shown in formula (2) in, Represents the exclusive OR operator symbol; Step 2: Reverse identification bit, mask update shift; Assume that the regular assembly instruction text corresponding to instruction I is T, which consists of an opcode, an operand, and a predicate, as shown in formula (3): T=[op,or,pr] (3) Then flip the instruction The corresponding instruction text is As shown in formula (4): Compare the instruction text T and When the opcode operand or predicate When it changes, the bit belongs to the bit field of the corresponding opcode, operand or predicate; if it does not change, the bit belongs to the reserved bit; the mask M is logically shifted right by one bit, and the XOR operation is performed again with the instruction I, and the total loop is L times; Step 3: Extract key fields and compress redundant space; Using semi-supervised dimensionality reduction to compress the redundant instruction space and separate the high-value search space; Assumption I op The bit width is L op , then the semi-supervised compressed instruction space S reduced Then: Step 4: Assemble the instructions to be tested and analyze the disassembly status; At this time, the length of the compressed instruction space is only equal to I op For the remaining instruction bits, I or ,I de and I pr are all set to 0; according to the instruction encoding rule of formula (1), assemble into a test instruction I T , as shown in formula (6); I T =[I op ,0,0,0] (6) For the instruction to be tested I T Perform disassembly analysis and set the disassembly status code Q to determine whether it is a disassemblyable instruction; if the disassembly status code Q = 0, the instruction to be tested I T If the disassembly status code Q=1, the instruction to be tested I T If it is a non-disassembled instruction, it is saved as an instruction to be executed; Step 5: Instructions are adaptively moved and instructions to be executed are saved; The instruction counter k increases by 1. When the maximum count value k is reached, cnt_max When , reset the instruction counter k to 0, as shown in formula (7); The walking step length λ is adaptively increased by 1, as shown in formula (8); λ=λ+1 λ≤λ step_max (8) When λ increases to the maximum step size λ step_max When , reset the instruction counter k, λ is adaptively set to 1, and set an empty set E to store the instructions to be executed, as shown in formula (9); Then the wandering position x is updated. If x exceeds the wandering boundary value, the wandering process ends, as shown in formula (10); Step 6: Generate an executable file and identify hidden instructions; The host computer loads the instructions and data into its GPGPU global memory and starts the GPGPU to execute. The source code is compiled to generate an executable file, which is the medium for storing instructions and data. From the set of instructions to be executed E generated in the previous step, the i-th instruction I is taken out in turn. i Used for GPGPU to generate executable files and judge instruction I based on whether the executable file triggers an illegal instruction exception state i Whether it is a hidden command.

2. A GPGPU hidden instruction mining method based on semi-supervised adaptive walking as claimed in claim 1, characterized in that: In step 5, the starting position of the walk is 0 and the walk boundary is 3. A GPGPU hidden instruction mining method based on semi-supervised adaptive walking as claimed in claim 1, characterized in that: In step 5, the initial value of the walking step length λ is set to 1, and the maximum step length λ step_max is 5; set instruction counter k to 0, maximum count value k cnt_max is 10.

4. A GPGPU hidden instruction mining method based on semi-supervised adaptive walking as claimed in claim 1, characterized in that: The specific method of step 6 is: The executable file F consists of four parts: the header ehead, section section, section header table shead, and program header table phead. The section header table offset e_shoff in the header ehead indicates the location of the section header table. The section header table entry size e in the section header table shead indicates the location of the section header table. shentsize and the number of section header entries e shnum , indicating the table entry location where the instruction is stored; the program segment offset sh_offset in the table entry is the location addr where the instruction is finally stored. i Embedded into this position, as shown in formula (11): If GPGPU triggers an illegal instruction exception and returns status code E equal to 715, then the instruction I i It is an illegal instruction and is discarded; if the return status code E is not equal to 715, the instruction I is determined to be i is a hidden instruction and is saved in the hidden instruction table H, as shown in formula (12); I i ∈H E≠715 (12) Determine whether the pending instructions in set E have been executed. If not, update and execute the next pending instruction; if completed, exit.