Return-Oriented Code Automatic Generation Method and Device for Handling Side Effects of Code Fragments

By segmenting and processing the side effects of code snippets under embedded architecture, more efficient ROP automatic generation is achieved, solving the problem of low success rate in embedded architecture in the existing technology, and improving the success rate of ROP generation and the availability of code snippets.

CN115437622BActive Publication Date: 2025-06-24INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210932704.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-06-24
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

The existing ROP automatic generation technology has a low success rate under embedded architecture, mainly due to the small number of code snippets and the small optional range of single-function code snippets, making it difficult to build an effective ROP chain.

Method used

A return-oriented code automatic generation method for dealing with side effects of code snippets is proposed. By segmenting the multifunction code snippet into a single-function subcode snippet, combining symbol execution and SMT solver, the side effects of code snippets are identified and processed, thereby improving the success rate of ROP generation.

Benefits of technology

By dealing with the side effects of code snippets, the success rate of generating return-oriented utilization code in binary files is increased, the number of available code snippets under the embedded architecture is expanded, and the application and expansion capabilities of ROP technology in embedded software are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115437622B_ABST
    Figure CN115437622B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for automatically generating return-oriented code for handling side effects of code fragments, which relates to the technology of automatically generating vulnerability exploitation. The present invention takes into account the side effects of code fragments and sets corresponding processing strategies; uses the method of splitting code fragments to describe multifunctional code fragments as single-functional sub-code fragments; uses symbolic execution to implement an automatic ROP generation framework, and through designing the functional description of code fragments, identifies and specifically processes the side effects in the code fragments, thereby enabling the use of code fragments with side effects and increasing the success rate of generating return-oriented exploitation code in binary files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of automatically generating exploit for vulnerabilities, and specifically to a method and device for automatically generating for Return-Oriented Programming (ROP). Background Art

[0002] Memory corruption vulnerabilities, as a type of high-risk vulnerabilities, allow attackers to hijack the expected control flow of a program by corrupting memory and then execute the behaviors desired by the attackers. The corruption behaviors can be injecting a piece of shellcode to control the target system program, installing malware, or leaking sensitive information stored in the program. To defend against attacks caused by memory corruption vulnerabilities (such as code injection attacks), modern operating systems have all enabled the W⊕X protection. Attackers then use the executable code already in the memory space to execute attacks. For example, Glibc is the core runtime library in the GNU system, and it provides many available functions, such as execve and system. Once an overflow vulnerability is found, attackers can construct the call parameters of the function and redirect the return address pointer to the starting address of these Glibc functions, and then start the process that the attackers hope to execute under the permissions of the current program. In 2007, Shachamet et al. proposed the Return-Oriented Programming (ROP) technology. Its core idea is as follows: Programs usually contain a large number of return instructions (ret), which are usually located at the end of a function or where a return is needed in the middle of a function. And the instruction sequence from any instruction to the ret instruction is called a code fragment, which makes it possible to combine operations such as reading and writing memory, arithmetic and logical operations, control flow jumps, and function calls into a whole. Therefore, executing the various code fragments in the memory space in a certain order can achieve the purpose of performing arbitrary operations. And in order to "piece together" the various code fragments, attackers need to construct a special sequence to achieve the desired effect. However, the control environments caused by different vulnerabilities are also different. For example, a stack overflow allows attackers to directly control the return address and the subsequent execution sequence. In the face of heap-related vulnerabilities, attackers need to hijack the control flow of the program through methods such as heap feng shui, and depending on different situations, the ROP starting point controlled registers are different, which requires attackers to perform stack migration according to specific situations. Because ROP allows attackers to use any non-randomized code to conduct attacks. The ROP technology is widely used to construct vulnerabilities. Especially in modern operating systems, non-randomized code may not directly contain functions useful to attackers. And previous research has also shown that it is entirely possible to find a Turing-complete set of code fragments in open-source and commercial projects, such as Glibc, the Windows kernel, etc.

[0003] Evaluating vulnerabilities can not only provide sufficient security risk assessment information for software, but also be used to add basis for formulating protection rules for intrusion detection systems. Generating exploit is an important means to evaluate the threat of vulnerabilities. The traditional manual method of writing exploits is time-consuming and laborious, and has high requirements for the ability of the writer, which is not conducive to the rapid and effective development of the evaluation work. Therefore, researching the automatic exploit generation technology (Automatic Exploit Generation, AEG) is an important method and way to improve the vulnerability evaluation ability and the software security evaluation ability. Among them, how to automatically generate ROP has always been a popular research topic. The existing ROP generation work is mainly based on the functions of the exploitable code fragments in the software. The code function information of the exploitable code fragments is predefined in advance, and then the code fragments with corresponding functional characteristics are found in the program. Usually, the code fragments used by the traditional ROP generation method only contain a specific function (referred to as single-functional code fragments). However, single-functional code fragments are only a small part of the available code fragments in the program. In traditional architectures (such as x86, x64, etc.), due to the sufficient number of code fragments, it is easy to select available single-functional code fragments from these code fragments to construct ROP. However, with the booming development of embedded software, in embedded architectures (such as MIPS, PowerPC, etc.), the number of code fragments in the software is relatively small, and the optional range of single-functional code fragments becomes smaller. This seriously restricts the application and expansion of ROP technology in the exploitation of embedded software vulnerabilities. In the traditional x86 architecture, a variable-length instruction set is adopted, and memory alignment is not required. If the attacker controls the ip register to jump to the middle of some instructions, unexpected instructions will be generated. And the instruction set of x86 is very large and its encoding is very dense. Therefore, even in relatively small programs, there are various instructions available, enriching the number of code fragments. While in embedded systems, architectures such as ARM and PowerPC are usually adopted, which use a fixed-length instruction set design and enforce alignment during instruction fetching, resulting in fewer available code fragments than x86. With the increasing popularity of the Internet of Things and other embedded devices, such as ARM and MIPS are widely used in small routers, smart devices and other systems, while PowerPC is widely used in network devices and industrial control systems. The security issues on these systems have also attracted more and more attention from attackers and hackers. And the code fragment function description strategy in the existing ROP generation technology is no longer sufficient to find available code fragments in the software of the above architectures.

[0004] To sum up, the applicability of the existing ROP automatic generation technology is not high and the success rate is relatively low, and this problem needs to be solved. Summary of the Invention

[0005] In order to solve the problem of inaccurate and incomplete functional description of code snippets, the present invention proposes a method and device for automatically generating return-oriented code for processing side effects of code snippets based on the functional description of Q's code snippets. The present invention takes into account the side effects of code snippets and sets corresponding processing strategies; uses the method of splitting code snippets to split multifunctional code snippets into single-function sub-code snippets for description; uses symbolic execution to implement an automatic ROP generation framework, and identifies and specifically processes the side effects in the code snippets by designing the functional description of the code snippets, thereby realizing the use of code snippets containing side effects and increasing the success rate of generating return-oriented exploit code in binary files.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for automatically generating return-oriented code for processing side effects of code snippets comprises the following steps:

[0008] For the input binary file, collect code snippets from it;

[0009] Filter code snippets and remove useless code snippets;

[0010] Perform semantic analysis on the filtered code snippets to obtain the functional information of each code snippet, split the multi-functional code snippets into multiple sub-code snippets with single functions, and construct a code snippet collection;

[0011] Acquire the utilization restriction information input by the user, and split the utilization restriction information into function information;

[0012] According to the split functional information, the code snippet that best meets the functional requirements is found from the code snippet set, and the constraints and other side effects that meet the utilization constraint information are extracted from it;

[0013] Process the side effects in the code snippet, convert them into constraints, combine them with the constraints that meet the utilization constraint information, and solve them. If the solution fails, find and perform the same process from the remaining code snippets. If the solution succeeds, obtain the code snippet that meets the utilization constraint information.

[0014] The code fragments that meet the utilization constraint information are combined to form a ROP chain, i.e., a return-oriented code.

[0015] Furthermore, the ropper and ropgadget tools are used to collect code snippets from the input binary files.

[0016] Further, the code snippets are screened according to the following 4 rules: screen out the code snippets with illegal memory access, screen out the code snippets with instructions outside the privilege level range of program operation, screen out the code snippets that cannot control the jump address, and screen out the code snippets with excessive stack memory offset.

[0017] Further, whether a code snippet can control the jump address is determined by judging whether the corresponding register can be controlled. If the corresponding register can be controlled, it is judged that the code snippet can control the jump address; otherwise, it is judged that the code snippet cannot control the jump address.

[0018] Further, the method for semantic analysis of the screened code snippets is as follows: pass the screened code snippets to the code snippet function analyzer. The code snippet function analyzer first initializes the state of symbolic execution and defines event breakpoints of angr to listen for corresponding events in the predefined function information set, and then performs symbolic execution on each code snippet. When angr stops symbolic execution when triggering a specific event, record the event breakpoint and analyze the function information of the code snippet.

[0019] Further, in addition to including the regular function information of the code snippet, the function information set also includes the following function information:

[0020] JUMPADDRG function, representing the address pointed to by a direct jump instruction;

[0021] SYSCALLG function, representing a system call;

[0022] STACKPIVOTG function, representing stack migration;

[0023] WRITESTACKG function, representing analyzing the impact of the analysis operation on the use of subsequent code snippet linkage and specifically eliminating the impact during the subsequent combination process;

[0024] CRASHG function, representing that the current code snippet will cause the program to crash;

[0025] IFG function, representing that the current code snippet will generate two branches, and different branches correspond to different conditions.

[0026] Further, the "most in line with the functional requirements" means: the number of functions meeting the requirements is more than a threshold; the number of side effects is less than a threshold; the stack frame change value is less than a threshold, and the above three thresholds are set according to actual requirements.

[0027] Further, the z3-solver is selected as the solution algorithm.

[0028] Further, the side effects of code snippets include three categories: register modification side effects, memory access side effects, and stack impact side effects.

[0029] Further, the processing of side effects includes the following three types:

[0030] For register modification side effects, for register modification side effects with known values, add a code snippet with the LOADCONSTG function after using the code snippet to restore the original value of the register; for register modification side effects with unknown values, select a code snippet with the MOVEREGG function to restore; if there is no available MOVEREGG function code snippet, use the code snippets with the STOREMEMG and LOADMEMG functions respectively to save a copy in memory to restore the original value of the register;

[0031] For memory access side effects, obtain its target memory address Des_addr, and set the register or memory through the code snippet with the corresponding function to ensure that the Des_addr is a legal address;

[0032] For stack impact side effects, the reason for generating this side effect is that the data referenced by the two combined code snippets conflicts. By adding a padding code snippet between the two code snippets, the conflicting areas are staggered.

[0033] A return-oriented code automatic generation device for processing the side effects of code snippets includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the steps of the above method are implemented.

[0034] In the method of the present invention, a functional description scheme for code snippets is proposed. In practice, most code snippets contain more than one functional description. If the original method is used and only a single function definition is used, it is difficult to accurately describe these code snippets. Therefore, this scheme splits these multi-functional code snippets into multiple single-functional sub-code snippets, and by combining these single-functional sub-code snippets, the code snippet can be more accurately described. At the same time, a code snippet contains multiple functions, but when selecting a code snippet, sometimes only one or several of the functions are needed. The remaining functions may cause interference to the utilization. These functions that may cause interference to the utilization generation are called the side effects of the code snippet. In order to make the success rate of utilization as high as possible, the present invention sets corresponding processing methods for the side effects of code snippets. First, the side effects of code snippets are classified into three categories: (1) register modification side effects, (2) memory access side effects, (3) stack impact side effects. And different processing methods are constructed for different side effect types to achieve the purpose of eliminating side effects.

[0035] The present invention has the following advantages:

[0036] 1. The present invention expands the functional semantic description of code snippets, sets corresponding processing strategies for the side effects of code snippets, constructs ROPs by fixing the side effects of these categories of code snippets. Compared with existing tools, when generating ROPs, the available code snippets are increased, and the number of available code snippets for programs with architectures such as MIPS, ARM, and PowerPC is expanded and increased, improving the possibility of successfully constructing ROPs.

[0037] 2. The present invention expands on the function set of Q, and describes the functions of code snippets more comprehensively and accurately. A new code snippet function description scheme is proposed, which splits multifunctional code snippets into multiple single-functional sub-code snippets for description.

[0038] 3. Based on technical foundations such as symbolic execution and SMT solvers, the present invention implements an easily extensible automatic ROP generation framework, which can generate ROP attack chains for programs on different platforms and architectures, facilitating the use by people with different needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of an automatic return-oriented code generation method for processing the side effects of code snippets according to an embodiment of the present invention.

[0040] Figure 2 is a list of function descriptions of code snippets.

[0041] Figure 3 is a schematic diagram of the processing methods for different types of side effects in code snippets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the above features and advantages of the present invention more obvious and understandable, specific embodiments are given below and described in detail in conjunction with the accompanying drawings as follows.

[0043] This embodiment provides an automatic return-oriented code generation method for processing the side effects of code snippets. For a given binary file, the method is processed according to the Figure 1 flow shown, and the specific processing steps are as follows:

[0044] 1. Use the ropper and ropgadget tools to collect code snippets from the input binary file.

[0045] 2. Screen the code snippets according to preset rules to prevent useless code snippets from interfering with subsequent analysis. The specific implementation of these rules depends on different systems and architectures.

[0046] The preset rules include the following 4:

[0047] Rule 1: Screen out code segments with illegal memory accesses;

[0048] Rule 2: Screen out code segments of instructions outside the privilege level range of program operation;

[0049] Rule 3: Screen out code segments that cannot control the jump address;

[0050] Rule 4: Screen out code segments with excessive stack memory offsets.

[0051] In PowerPC, the jumps of code segments are all determined by control registers. For example, blr will jump to the address stored in the LR register. Therefore, the way to determine whether a code segment can jump smoothly is to simply determine whether the corresponding register can be controlled. For privileged instructions, only some specific registers in PowerPC can be accessed in privileged mode. If it is a user-mode PowerPC program, it only needs to determine whether there are access operations to these registers. The framework of this embodiment incorporates screening rules for PowerPC, MIPS, ARM, and X86 architectures. For other architectures, users can also expand according to their own needs.

[0052] 3. Pass the screened code segments to the code segment function analyzer. At the same time, the code segment function analyzer initializes the state of symbolic execution and defines angr event breakpoints to listen for corresponding events in the pre-defined function information set. In angr, the smallest step unit of symbolic execution is a basic block. However, a basic block may contain multiple functions. If analyzed after the execution of the basic block, a lot of function information may be lost. Therefore, this method chooses to let angr stop symbolic execution when a specific event is triggered, which is convenient for analyzing its state. Therefore, this method uses angr event breakpoints. For example, when the program executes to the instruction stw r28,8(r1), angr will consider it as a mem_read(BP_BEFORE) event and stop just before the read. At this time, the target address r1 + 8 and the value r28 to be written can be obtained. If r1 + 8 at this time is a symbolic value, a new area of symbolic variables will be applied, r1 will be pointed to the target area, and this memory read will be recorded. In this way, event breakpoints can be used to analyze the functions in the code segments.

[0053] 4. Symbolically execute each code snippet using the symbolically executed state after initialization to analyze the functional information of each code snippet. After each code snippet is symbolically executed, the functional information of each modified register or read / written memory can be obtained. Split the multifunctional code snippets into multiple single-functional sub-code snippets. The code snippet set consists of the filtered single-functional code snippets and the split single-functional snippets. Each functional information corresponds one-to-one with the split single-functional sub-code snippets, and the functional information list of the multifunctional code snippet is composed of the functional information combinations of these sub-code snippets. Among them, the set of functional information used in this method is as Figure 2 shown, where InReg, OutReg, etc. represent general-purpose registers other than the stack address register, M[addr] represents a memory access operation to the address addr outside the stack space, Stack[Offset] represents an index to a specific offset on the stack, < >b represents any arithmetic operation, OutReg←InReg represents assigning the value of InReg to OutReg, Figure 2 The semantics marked with * in it are the semantics extended by the present invention. This method adds the following functions compared with the existing methods:

[0054] JUMPADDRG represents the address pointed to by a direct jump instruction. With the continuous maturity of vulnerability exploitation techniques, techniques such as ret2csu have been widely applied to the writing of exploitation code. The key instruction used in it is qword ptr[r12 + rbx * 8]. When using this code snippet to write ROP, generally, the value of r12 + rbx * 8 is made to point to the data in the GOT table to achieve the exploitation effect of ret2libc. To better describe such code snippets and achieve a similar exploitation effect, the function of JUMPADDRG needs to be added.

[0055] SYSCALLG represents a system call. Although the existing work has considered the function of system calls in the design, it has not analyzed the code snippets containing system call instructions. And many current exploitation techniques will use code snippets containing system calls. To solve the problems related to the use of system calls, this method processes the parameters and return values of system calls and adds the analysis of system calls.

[0056] STACKPIVOTG represents the function of stack pivoting. When exploiting heap vulnerabilities, it is difficult for attackers to control the data on the stack. In this case, attackers tend to use stack pivoting techniques to point the stack pointer to a memory area that the attacker can control to implement ROP. Existing work cannot handle code snippets related to stack pivoting, resulting in its inapplicability to some scenarios such as the exploitation writing of kernel vulnerabilities. To implement the stack pivoting function in such exploitation processes, this method adds the analysis of such code snippets to better adapt to the exploitation generation in different scenarios.

[0057] WRITESTACKG is a very special code snippet function. Many existing exploitation techniques, such as ret2csu and ret2dl_resolve, will widely use code snippets ending with call. The difference between call and ret is that the call instruction will change the content on the stack, affecting the combination of subsequent code snippets. Moreover, in PowerPC, most code snippets will affect the subsequent stack frames. To accurately describe such an impact, this method introduces the functional description of WRITESTACKG to analyze the impact of these operations on the use of subsequent code snippet linking and to specifically eliminate such an impact during the subsequent combination process.

[0058] CRASHG represents that the current code snippet will cause the program to crash. During the process of processing code snippets, it is impossible to predict whether the current code snippet can be successfully executed because tools like ropgadget and ropper for obtaining code snippets do not perform such checks. Therefore, whether the current code snippet can be successfully executed should be used as the basis for evaluating whether the current code snippet is usable. If a code snippet ends with CRASHG, it means that this code snippet will cause the exploitation to not continue, and such a code snippet will not be considered for use during the combination of code snippets.

[0059] IFG represents that the current code snippet will generate two branches, and different branches correspond to different conditions. In the actual exploitation writing process, the function of one of the branches may be used. For example, in the instruction sequence existing in the code snippet Gadget used by ret2csu:

[0060] add rbx,1

[0061] cmp rbp,rbx

[0062] jnz $-0x14

[0063] In general, attackers will set the value of rbx to 0 and rbp to 1 to prevent the jump from happening. This method also hopes to consider using such code snippets as much as possible to increase the success rate of the exploit, so the functional analysis of such code snippets is added.

[0064] 5. Get the exploit constraint information entered by the user and split the exploit constraint information into functional information. For example, if the current program is running under the Linux system and the user's desired exploit effect is to enable a shell, this method will split it into:

[0065] (1) Write the string “ / bin / sh” to the writable area in the memory.

[0066] (2) Set the r3 register (the register where the first parameter of a function is stored in the PowerPC architecture) to the string address. (3) Jump to the system function to complete the exploit.

[0067] 6. According to the functional information split by the utilization constraints, select the code snippet that best meets the functional requirements from the code snippet set, and extract the constraints and other side effects that meet the utilization constraint information.

[0068] The most suitable functional requirements include: 1) as many functions as possible that meet the target exploitation effect; 2) as few side effects as possible other than these functions that meet the target exploitation effect; 3) as small a stack frame change value as possible to make the final ROP chain as short as possible; the above three items can set thresholds according to requirements.

[0069] 7. Use the preset strategy to process the side effects in the code snippet, convert them into constraints, combine them with the constraints that meet the constraint information, and use z3-solver to try to solve them. If the solution fails, return to step 6 and continue to search from the remaining code snippets. Specifically, in the code snippet analysis stage, this embodiment can obtain the stack frame change value O of each Gadget. For each position i (0≤i<O) starting from the top of the stack, this embodiment will set corresponding constraints for the corresponding index Si on the stack S. For example, this embodiment uses the code snippet: lwz r0,0x8(r1); mtlr r0; lwz r3,(r1); blr; to set the value of the r3 register to X and jump to the Gadget with address Y, which can be converted into the following conditions: (The lwz instruction will index 4 bytes of data.) Assuming the previous step used the code snippet: lwz r0,(r1); mtlr r0; lwz r4,0x4(r1); blr; (with the WRITESTACKG side effect) setting r4 to Z will add restrictions to the current code snippet. On this basis, if the overflow is caused by improper use of the strcpy function, additional restrictions will be added. Finally, the above conditions are combined and handed over to the SMT solver for solving. If it can be solved successfully, it means that the current code snippet meets the requirements. Otherwise, it will continue to try with the next alternative code snippet.

[0070] The treatment of three side effects is aimed at eliminating side effects:

[0071] Register modification side effects: For register modification side effects with known values, add a gadget with LOADCONSTG function to restore the original value of the register after using such a gadget. For register modification side effects with unknown values, a gadget with MOVEREGG function will be selected for recovery; if there is no available MOVEREGG function gadget, a copy will be saved in memory using STOREMEMG and LOADMEMG function gadgets respectively to restore the original value of the register;

[0072] Memory access side effects: For this type of side effects, the target memory address Des_addr will be obtained, where Des_addr may be a single register, or it may be the result of the operation of multiple registers or data in memory, which is uniformly represented as S1◇bS2◇b...Sn (where ◇b represents any operation). Therefore, for Des_addr, the register or memory can be set through the corresponding function gadgets such as LOADCONSTG to ensure that Des_addr is a legal address;

[0073] Stack-affecting side effects: This type of side effect is caused by a conflict in the data referenced by the two combined gadgets. A filler gadget is added between the two gadgets to offset the conflicting areas, allowing the ROP to be successfully constructed.

[0074] The schematic diagram of the treatment methods for the above three types of side effects is as follows Figure 3 After completing the processing of the gadget side effects, the corresponding ROP can be generated according to the specific utilization constraints.

[0075] 8. Combine the code fragments that meet the exploit constraint information (divided into code fragments that do not need to be split and sub-code fragments of the code fragments that have been split) to form a ROP chain, i.e., return-oriented code. The ROP chain and the exploit constraint information (i.e., the exploit constraint information that can correctly jump to the ROP execution code when exploiting the vulnerability) are combined to output a complete exploit script, which can correctly execute the ROP chain script.

[0076] Although the present invention has been disclosed above by way of examples, it is not intended to limit the present invention. Any appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to that defined by the claims.

Claims

1. A method for automatically generating return-oriented code to handle the side effects of code fragments, characterized in that The following steps are involved: For the input binary file, collect code snippets from it; Filter code snippets and remove useless code snippets; Perform semantic analysis on the filtered code snippets to obtain the functional information of each code snippet, and split the multifunctional code snippets into multiple sub-code snippets with single functions to construct a code snippet set. The method for performing semantic analysis on the filtered code snippets is as follows: the filtered code snippets are passed to a code snippet functional analyzer, which first initializes the state of symbolic execution and defines angr's event breakpoints to monitor the corresponding events in the pre-defined functional information set, and then performs symbolic execution on each code snippet. When angr triggers a specific event, it stops the symbolic execution, records the event breakpoints, and analyzes the functional information of the code snippet. Acquire the utilization restriction information input by the user, and split the utilization restriction information into function information; According to the split functional information, the code snippet that best meets the functional requirements is found from the code snippet set, and the constraints and other side effects that meet the utilization constraint information are extracted from it; Process the side effects in the code snippet, convert them into constraints, combine them with the constraints that meet the utilization constraint information, and solve them. If the solution fails, find and perform the same process from the remaining code snippets. If the solution succeeds, obtain the code snippet that meets the utilization constraint information. The code fragments that meet the utilization constraint information are combined to form a ROP chain, i.e., a return-oriented code.

2. The method according to claim 1, characterized in that Use the ropper and ropgadget tools to collect code snippets from input binaries.

3. The method according to claim 1, characterized in that The code snippets are filtered according to the following four rules: filter out code snippets with illegal memory access, filter out code snippets with instructions that are not within the privilege level range of the program, filter out code snippets that cannot control the jump address, and filter out code snippets with excessive stack memory offset.

4. The method according to claim 3, wherein Whether the code snippet can control the jump address is determined by judging whether the corresponding register can be controlled. If the corresponding register can be controlled, it is judged that the code snippet can control the jump address, otherwise it is judged that the code snippet cannot control the jump address.

5. The method according to claim 1, wherein In addition to the general function information of the code snippet, the function information set also includes the following function information: The JUMPADDRG function indicates the address pointed to by a direct jump instruction; SYSCALLG function, indicating a system call; STACKPIVOTG function, indicating stack migration; WRITESTACKG function, which indicates the impact of the analysis operation on the use of subsequent code fragment links, and specifically eliminates the impact in the subsequent combination process; CRASHG function, indicating that the current code snippet will cause the program to crash; The IFG function indicates that the current code snippet will generate two branches, and different branches correspond to different conditions.

6. The method according to claim 1, wherein The most consistent functional requirement means: the number of functions that meet the requirement is greater than a threshold; the number of side effects is less than a threshold; the stack frame change value is less than a threshold, and the above three thresholds are set according to actual requirements.

7. The method according to claim 1, characterized in that, The solving algorithm selects z3-solver.

8. The method according to claim 1, characterized in that, The side effects of the code snippet include three categories: register modification side effects, memory access side effects, and stack impact side effects; handling the side effects includes: For register modification side effects, for those with known values, after using the code snippet, add a code snippet with the LOADCONSTG function to restore the original value of the register; for those with unknown values, select a code snippet with the MOVEREGG function to restore; if there is no available MOVEREGG function code snippet, use the code snippets with the STOREMEMG and LOADMEMG functions respectively to save a copy in memory to restore the original value of the register; For memory access side effects, obtain the target memory address Des_addr, and set the register or memory through the code snippet with the corresponding function to ensure that the Des_addr is a legal address; For stack impact side effects, the reason for this side effect is that the data referenced by the two combined code snippets conflicts. By adding a padding code snippet between the two code snippets, the conflicting areas are staggered.

9. An automatic return-oriented code generation device for handling side effects of code snippets, characterized in that, It includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the steps of the method described in any one of claims 1-8 are implemented.

Citation Information

Patent Citations

  • Software security detection method for code reuse programming

    CN105138914A

  • ROP chain detection method and device and medium

    CN114826793A