Binary program processing method, device and equipment
By converting binary programs into low-level intermediate representations, identifying library function call sites and anchor variables, and using the known behavior of library functions to infer source code-level information, a high-level intermediate representation is generated, which solves the problem of high cost and low efficiency of binary-level analysis and achieves more efficient analysis.
Patent Information
- Application Number
- CN202510756309.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
Existing binary-level analysis is costly and inefficient, has difficulty effectively handling behavioral changes caused by compiler-induced transformations and optimizations, and lacks source-level information.
By converting the binary program into a low-level intermediate representation, identifying the call sites of library functions and determining anchor variables, obtaining the context of anchor variables, and using the known behavior of library functions to infer missing source code-level information, a high-level intermediate representation is generated.
It reduces the cost of binary-level analysis, improves analysis efficiency, and enables better utilization of existing source code analysis algorithms for further analysis.
Smart Images

Figure CN120670023A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of program analysis technology, and in particular to a binary program processing method, apparatus, and device. Background Art
[0002] Binary-level analysis is an essential and crucial aspect of program analysis.
[0003] First, while source code analysis is advantageous, it's not always feasible. Due to increasing emphasis on intellectual property and security protection, many programs (e.g., commercial software and firmware) typically don't disclose their source code. Consequently, there's a growing interest in analyzing and testing third-party binary files.
[0004] Second, even if source code is available, modern build systems are becoming increasingly complex and diverse due to the proliferation of compiler types, versions, and configurations. Therefore, it is difficult for source code analysis tools to ensure compatibility with a wide variety of build systems. Performing binary-level analysis can help circumvent compatibility issues.
[0005] Third, source code does not take into account the transformations, reorderings, and optimizations caused by the compiler. These processes can cause unexpected execution or even security vulnerabilities in the behavior of the source code. Analyzing the actual binary program files helps to carefully examine the actual implemented behavior.
[0006] Current binary-level analysis focuses on developing dedicated binary analysis algorithms based on disassembled code or decompiled pseudocode, which is costly and inefficient.
[0007] Based on this, a binary-level analysis solution is needed that can help reduce costs and improve efficiency. Summary of the Invention
[0008] One or more embodiments of this specification provide a binary program processing method, apparatus, device, and storage medium to solve the following technical problem: a binary-level analysis solution is needed that helps reduce costs and improve efficiency.
[0009] To solve the above technical problems, one or more embodiments of this specification are implemented as follows:
[0010] One or more embodiments of this specification provide a binary program processing method, including:
[0011] Obtaining a low-level intermediate representation obtained by disassembling the target binary program;
[0012] In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site;
[0013] In the low-level intermediate representation, obtaining a context related to the anchor variable;
[0014] Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions;
[0015] The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.
[0016] One or more embodiments of this specification provide a binary program processing device, including:
[0017] A low-level intermediate representation acquisition module obtains a low-level intermediate representation obtained by disassembling the target binary program;
[0018] an anchor variable determination module, which identifies call sites of well-defined library functions in the low-level intermediate representation and determines anchor variables in the low-level intermediate representation based on variables associated with the call sites;
[0019] A context acquisition module, which acquires context related to the anchor variable in the low-level intermediate representation;
[0020] a missing information inference module, which infers the missing source code level information of the low-level intermediate representation based on the context related to the anchor variable by matching it with the known behavior of the library function;
[0021] The high-level intermediate representation generation module amends the low-level intermediate representation according to the missing source code level information to generate a high-level intermediate representation.
[0022] One or more embodiments of this specification provide a binary program processing device, including:
[0023] at least one processor; and,
[0024] a memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor to enable the at least one processor to perform:
[0026] Obtaining a low-level intermediate representation obtained by disassembling the target binary program;
[0027] In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site;
[0028] In the low-level intermediate representation, obtaining a context related to the anchor variable;
[0029] Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions;
[0030] The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.
[0031] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to:
[0032] Obtaining a low-level intermediate representation obtained by disassembling the target binary program;
[0033] In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site;
[0034] In the low-level intermediate representation, obtaining a context related to the anchor variable;
[0035] Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions;
[0036] The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.
[0037] At least one of the above-mentioned technical solutions adopted in one or more embodiments of this specification can achieve the following beneficial effects: first, the binary program is converted into an easily obtainable low-level intermediate representation, and then, based on a clearly defined library function, in the low-level intermediate, anchor variables that can help infer missing source code-level information are more accurately found, and then based on the anchor variables, reasoning is carried out in the low-level intermediate representation, and the known behavior of the library function is used as a reference for matching, so as to infer more missing source code-level information, and use this information to correct the low-level intermediate representation to generate a high-level intermediate representation. Since the high-level intermediate representation has more complete source code-level information, it helps to further utilize the existing source code analysis algorithm for further analysis without the need to implement redundant and specialized binary analysis algorithms, thereby helping to reduce costs and improve efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0039] Figure 1 A flowchart of a binary program processing method provided for one or more embodiments of this specification;
[0040] Figure 2 A schematic flow chart of a solution for determining an anchor variable according to one or more embodiments of this specification;
[0041] Figure 3 A flowchart of a missing source code level information inference solution based on matching known behavior of library functions provided in one or more embodiments of this specification;
[0042] Figure 4 One or more embodiments of this specification provide Figure 1 A schematic flow chart of a specific embodiment of the method;
[0043] Figure 5 A schematic diagram of the structure of a binary program processing device provided for one or more embodiments of this specification;
[0044] Figure 6 A schematic diagram of the structure of a binary program processing device provided in one or more embodiments of this specification. DETAILED DESCRIPTION
[0045] The embodiments of this specification provide a binary program processing method, apparatus, device, and storage medium.
[0046] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0047] A binary program is a program expressed in machine code. It's typically generated by compiling high-level language source code into an instruction format directly understood by the processor. These instructions, consisting of a series of binary bits (0s and 1s), are the closest form of program expression to the hardware level, capable of running directly on a specific hardware platform. For example, common executable files like ".exe" files in Windows and "ELF" files in Linux are binary programs.
[0048] As mentioned in the background art, current binary-level analysis focuses on developing dedicated binary analysis algorithms based on disassembled code or decompiled pseudocode.
[0049] The basic principles of binary analysis algorithms are generally consistent with those of source code analysis algorithms. Binary analysis algorithms must address special cases that arise during compilation, such as the lack of explicit type information and changes in code structure. However, the source code analysis community has extensively developed and maintained a large number of high-quality analysis tools and algorithms.
[0050] Based on this, the present application considers that if the binary program can be converted into an intermediate representation for source code analysis, and the necessary information required by the source code analysis algorithm can be inferred and recovered (which does not exist in the binary program), a more advanced intermediate representation can be obtained. In this way, it can be more adaptable to some existing source code analysis algorithms, making it easier to directly use these source code analysis algorithms to implement further binary program analysis. In this way, the need for repeated development of binary analysis algorithms can be eliminated, and program analysis can unify different scenarios into a single analysis process, thereby improving overall efficiency, and there is no need to spend a high cost to develop the above-mentioned dedicated binary analysis algorithm.
[0051] Based on this general idea, the solution of this application will be further described below.
[0052] Figure 1 A flowchart of a binary program processing method provided in one or more embodiments of this specification.
[0053] Figure 1 The process in includes the following steps:
[0054] S102: Obtain a low-level intermediate representation obtained by disassembling the target binary program.
[0055] The intermediate representation is an abstract representation between high-level language source code and machine language. It plays a key role in binary analysis, acting as an abstraction layer between machine code and high-level programming languages. It is the core data structure in the compilation process. The compiler performs various optimization operations on it, such as eliminating redundant code and constant folding, to convert the intermediate representation into machine code for the specific target hardware.
[0056] Using an intermediate representation, analysts can derive actionable logical expressions from machine code and assembly language, thus avoiding the challenges of directly analyzing low-level assembly code. The choice of intermediate representation depends on its level of abstraction and the specific analysis goals, and each intermediate representation type has different advantages and limitations.
[0057] The intermediate representation (IR) currently seen is usually what this application considers to be a low-level intermediate representation. For example, the intermediate representation of the low-level virtual machine (LLVM, also known as the low-level virtual machine) is a low-level intermediate representation that is still independent of the specific hardware.
[0058] Low-level intermediate representations (LIRs), such as those described in intermediate languages like P-Code, MicroCode, and Vex, are specifically designed to emulate the behavior of a program during execution on a processor, making them particularly well-suited for analyzing the core logic of binaries. However, due to their low-level nature, they require the analyst to have a deep understanding of assembly and machine code, which can complicate analysis, especially for complex binaries. Furthermore, LILs often require the creation of custom algorithms and techniques tailored for specific types of analysis, which is both time-consuming and resource-intensive. Furthermore, LILs often lack high-level constructs, including variable and function names, and data types, that are often essential for certain types of analysis.
[0059] In one or more embodiments of the present specification, the machine code in the binary program is first reversely converted back to a more general low-level intermediate representation. Since binary programs generally lack the rich information found in high-level languages (e.g., variable names, function structures, etc.), the corresponding low-level intermediate representation often also lacks this information. It is for this reason that it is considered to be relatively "low-level" and this information belongs to the source code level.
[0060] For ease of description, some of the following embodiments are mainly described using the intermediate representation of the underlying virtual machine as an example.
[0061] The underlying virtual machine is the basic framework for a modular, reusable compiler toolchain. It does not specifically refer to a complete compiler, but rather a compilation infrastructure that supports multiple languages and can be used to build and optimize language compilers. Its features and uses include: supporting the compilation of high-level languages into an intermediate representation; providing powerful optimization capabilities (for example, code execution speed, size optimization, etc.); supporting front-end (converting source code into an intermediate representation), mid-end (optimizing the intermediate representation), and back-end (converting the intermediate representation into machine code on the target architecture) processes; and being widely used in modern compiler toolchains. For example, the Clang compiler (C, C++, and Objective-C compilers) is based on the underlying virtual machine. The advantage of the underlying virtual machine lies in its cross-platform capabilities and the flexibility to support multiple languages. Many languages (such as Rust and Swift) develop their compilers based on the underlying virtual machine framework.
[0062] The intermediate representation of the underlying virtual machine is a static single-assignment representation, where each variable is assigned a value only once, facilitating analysis and optimization. It exists in three forms: text (human-readable), binary (efficient machine storage), and in-memory graphical representation (used internally by the compiler). Current intermediate representations of the underlying virtual machine are generally considered low-level.
[0063] To address the issues mentioned in the background, we consider converting a low-level intermediate representation (LIR) into a high-level intermediate representation (HIR) by inferring or correcting missing or erroneous source-level information in the LIR. This can be achieved by first disassembling the target binary program using existing tools or methods.
[0064] S104: In the low-level intermediate representation, identifying call sites with clearly defined library functions, and determining anchor variables in the low-level intermediate representation based on variables associated with the call sites.
[0065] For the low-level intermediate representation, we attempt to identify key reference points, including variables that help infer types and semantics. These are called anchor variables. To improve reliability, we select anchor variables based on well-defined library functions, especially standard library functions. Library functions have relatively standard documentation for their functions. Therefore, at the call point of the library function, the return value type and parameter types of the function are determined. Using this as the anchor point for reasoning (specifically represented by the anchor variable), we conduct reverse data dependency tracing and perform type corrections on associated variables and the functions that process them.
[0066] For example, in the low-level intermediate representation, we focus on library functions with clear types and semantics. We scan the obtained low-level intermediate representation files, locate the call sites of these library functions, and extract related variables (preferably those with accurate types and operational semantics), especially pointer-type variables with relatively large impact ranges, as anchor variables. We can record the location of the anchor variable, the variable name, and its related function. In this way, the anchor variable is closely associated with the library function and can be identified by its accurate type and operational semantics, thus providing a reliable basis for recovering missing source-level information.
[0067] For library functions, you can use functions in the standard library (for example, libc, glibc). Of course, some third-party link libraries with clear definitions can also be used, such as opencv, boost and other libraries.
[0068] Based on this approach, one or more standard library functions with clear types and semantics can be identified as target standard library functions. In the low-level intermediate representation, the call site of the target standard library function is identified. One or more pointer-type variables associated with the call site (e.g., the associated return value or variables contained in the function parameters) are extracted as anchor variables in the low-level intermediate representation. The target standard library function can be, for example, fopen, malloc, or read.
[0069] S106: In the low-level intermediate representation, obtain the context related to the anchor variable.
[0070] Anchor variables may affect multiple different code branches through calls between functions or variable references. We should try to fully explore the code branches affected by the anchor variables and obtain the context related to the anchor variables. For example, we can describe this context in the form of function call chains or variable association relationships.
[0071] In one or more embodiments of the present specification, the function parameters or return values that the anchor variable depends on in the low-level intermediate representation are determined; for the call point of the caller function of the function to which the function parameter or return value belongs, the variables corresponding to the function parameter or return value in the caller function are identified as transition variables; based on the transition variables, iterative dependency analysis is performed to find the functions that directly or indirectly depend on the anchor variable in the low-level intermediate representation to form a function call chain; based on the function call chain, the context related to the anchor variable is determined.
[0072] After selecting and determining the anchor variable, context related to the anchor variable is collected through backward intra-procedural and inter-procedural data dependency analysis, which can be called analysis context. At this stage, definition and usage analysis can be used to track the data flow of each anchor variable, trace the function call chain from the anchor variable, and identify variables that propagate data between multiple functions as transition variables. For example, if function 2 calls function 1, and the anchor variable depends on parameter p in function 1, the corresponding variable v in function 2 can be treated as a transition variable. Iterative function call chain analysis is performed, starting from the transition variable and trying to capture all relevant data dependencies until no more exist. In this process, a set of function call chains can be generated for each anchor variable, each chain representing a different analysis context.
[0073] S108: Inferring the missing source code level information of the low-level intermediate representation by matching the context related to the anchor variable with the known behavior of the library function.
[0074] In one or more embodiments of the present specification, on the one hand, the types of other variables, parameters, or return values that directly or indirectly depend on the anchor variable can be inferred based on the influence of the anchor variable in its related context propagation; on the other hand, the semantics and structural patterns related to the function call point of the anchor variable can be matched.
[0075] The known behaviors of the library function mentioned in step S108 may include: the usage behaviors of the library function known based on experience in the field or within one's own project, such as which high-level uses it is used for based on its own underlying uses; which variables it is used with; which other functions it involves calling or being called; how its behavioral impact will be passed on to other code blocks; and so on.
[0076] These matched structural patterns can be used to infer the appropriate definitions and references of variables and functions. For example, if a function involves memory allocation (such as malloc or realloc), and its name contains the alloc keyword (the keyword here does not refer to the keyword defined in the program, but refers to certain words that this application is concerned with), it may be inferred that the return type of the function should be a pointer.
[0077] The reason is: taking the C language development environment as an example, alloc is most likely the abbreviation of allocation, which means allocation. In this environment, it is most likely to refer to memory allocation, and the return type of memory allocation-related functions is most likely a pointer. Therefore, this inference can be made based on alloc.
[0078] Intuitively, take the following source code as an example:
[0079] "int*arr=(int*)malloc(10*sizeof(int));"
[0080] Here, the malloc function is used to allocate memory space for 10 ints. Similarly, developers can use custom functions to call the malloc function to implement more complex memory allocation functions. When developers name custom functions based on the functional description, they are likely to use keywords such as alloc locally.
[0081] S110: According to the missing source code level information, the low-level intermediate representation is modified to generate a high-level intermediate representation.
[0082] In one or more embodiments of the present specification, based on missing source-level information, the types of parameters or return values of the scope functions of anchor variables in the low-level intermediate representation are corrected. Based on the number of caller functions involved in the anchor variable, the correction is propagated across function boundaries to correct the types of other variables that have dependencies on the anchor variable. For example, for the structural patterns mentioned above, once a matching structural pattern is identified in the low-level intermediate representation, corrections can be performed through intra-procedural and inter-procedural analysis, adjusting the definitions and references of related variables and functions based on the dependencies.
[0083] Furthermore, we ensure information synchronization and consistently propagate all modifications throughout the program. Specifically, we synchronize updates to shared nodes in the call chain and mark changes to function types, variable definitions, and references to guide subsequent iterative analysis.
[0084] The tracking of anchor variables described above relies on data dependency analysis. Heuristic rules can be used to modify related variables during this analysis. Alternatively, these heuristic rules can be replaced with a large language model to analyze code snippets and provide modification suggestions, further helping to reduce costs and improve efficiency.
[0085] In one or more embodiments of this specification, after generating a high-level intermediate representation, the high-level intermediate representation can be further analyzed using downstream source code analysis algorithms or tools. For example, SVF, KLEE, and LLVM infrastructure (such as clang, opt, llvm-dis, etc.) can be used for source code analysis.
[0086] pass Figure 1The proposed method first converts the binary program into an easily accessible low-level intermediate representation, and then, based on clearly defined library functions, more accurately searches for anchor variables in the low-level intermediate representation that can help infer missing source-level information. Then, based on the anchor variables, reasoning is carried out by diverging in the low-level intermediate representation, and matching is performed with the known behavior of the library function as a reference, so as to infer more missing source-level information. The low-level intermediate representation is corrected with this information to generate a high-level intermediate representation. Since the high-level intermediate representation has more complete source-level information, it is helpful to further utilize existing source code analysis algorithms for further analysis without the need to implement redundant and specialized binary analysis algorithms, thus helping to reduce costs and improve efficiency.
[0087] based on Figure 1 This specification also provides some specific implementation plans and extension plans of the method, which will be described below.
[0088] When determining anchor variables, for the variables related to the above-mentioned call points, variables with clear types can be selected as anchor variables as much as possible. However, here, the applicant also considers exploring the possible particularities of the target binary program at the logical level or business level, and based on this, giving priority to anchor variables that can reflect such particularities. This will help to subsequently determine missing source code-level information with more practical value, and will also help to improve subsequent processing efficiency.
[0089] Based on this idea, one or more embodiments of this specification provide a missing source code level information inference solution based on matching known behavior of library functions, see Figure 2 , which shows the process of the scheme.
[0090] Figure 2 The solution in this paper includes the following steps:
[0091] S202: For the variables related to the call point, determine whether the variables have clear operational semantics.
[0092] If it has clear operational semantics, it means that the variable has preliminary value in assisting in inferring the semantics or type of other information. In addition, on this basis, we can also consider the scope of influence of the variable through function calls, references, etc. (for example, how long is the corresponding function call chain, how far the code blocks spanned are, and how low the correlation is, etc.), and give priority to retaining variables with a wider range of influence.
[0093] S204: If yes, calculate the degree of deviation between the operational semantics and the known behavior based on the known behavior of the library function.
[0094] The known behavior of library functions reflects some of their empirical usage. This application focuses on the specific logic or business behavior exhibited by the target binary program. Therefore, we measure the specificity of the operational semantics by calculating the degree of deviation. The higher the degree of deviation, the more specific it is.
[0095] In one or more embodiments of the present specification, when calculating the degree of deviation, the operational semantics of the variable may be directly compared to determine how much the operational semantics of the corresponding part in the known behavior deviate from each other.
[0096] Furthermore, optionally, the overall operational semantics deviation between the operation chain where the variable is located (the operation chain composed of multiple propagation nodes in which the variable plays a major role can be selected) and the operation chain corresponding to the known behavior can be compared. The contribution of the operational semantics of the variable itself to the overall operational semantics deviation can also be calculated according to the node weight.
[0097] S206: If the degree of deviation is greater than a set threshold, the variable is determined as an anchor variable candidate variable, so as to select an anchor variable from the anchor variable candidate variables.
[0098] If the degree of deviation is greater than the set threshold, the variable is considered not only to be reliably defined but also to be sufficiently special. Therefore, it can be prioritized as an alternative anchor variable.
[0099] pass Figure 2 The proposed scheme helps to eliminate some anchor variable noises of little value and focus more on recovering more valuable source-level missing information.
[0100] In the pair Figure 1 In the description of , the idea of inferring missing source code level information is introduced. More intuitively, in order to facilitate specific implementation, one or more embodiments of this specification provide a missing source code level information inference solution based on matching known behavior of library functions, see Figure 3 , which shows a schematic diagram of the process.
[0101] Figure 3 The process in includes the following steps:
[0102] S302: Match the caller function name and hard-coded string involved in the context related to the anchor variable with the known behavior of the library function, wherein the known behavior includes: the behavior of the I / O library function in processing streams or files, and / or the behavior of the memory management library function in performing dynamic memory management.
[0103] S304: Based on the result of successful matching, it is inferred that the variable or return value in the low-level intermediate representation lacks a pointer type (it should originally be a pointer type, but the type information is missing).
[0104] I / O library functions, such as fgets, fputs, or fclose, infer that variables in these contexts are file pointers by matching the caller function name and the specified keywords that may exist in the hard-coded string, such as stream, file, or stdout.
[0105] Memory management library functions, such as malloc, realloc, and free, match the caller function names and specific keywords that may exist in hard-coded strings, such as alloc or memdup, to infer that the relevant variables or return values in these contexts are pointers.
[0106] S306: Match the context related to the anchor variable with the memory operation behavior with length constraint of the memory operation library function.
[0107] It should be noted that step S202 and step S204 may be two parallel sub-solutions, and therefore, their execution order may be flexibly selected as needed.
[0108] S308: Through the matching, if it is determined that the corresponding output buffer pointer and length constraint parameters exist in the context related to the anchor variable, and the output buffer is accessed in a truncated manner, then it is inferred that the definition and reference of the relevant variables in the low-level intermediate representation are missing array patterns.
[0109] Memory operation library functions, such as read, memcpy, strncpy, etc. If there are corresponding output buffer pointers and length constraint parameters, and the output buffer is subsequently accessed in a truncated manner, this often needs to be implemented based on an array (for example, the length constraint parameter is used to avoid buffer overflow, and the truncated access method is used to allow partial writing when the buffer is insufficient rather than crossing the boundary; in this case, the buffer is more likely to be a memory block implemented by an array, rather than a dynamically allocated memory block, because the latter is more flexible and has a relatively low risk of crossing the boundary, and truncated access is often not used). To implement such a set of operations, the characteristics represented are relatively consistent with the array mode, so the missing array mode information can be corrected and supplemented.
[0110] based on Figure 3 Based on the idea of pre-extracting more usage behavior features of library functions, and then matching them with local behaviors of different scales in the context related to the anchor variable, so as to infer missing information such as variable types in the context.
[0111] In one or more embodiments of the present disclosure, statements containing the undefined keyword in a low-level intermediate representation are identified, for example, by scanning the entire low-level intermediate representation file during initial analysis to identify undefined functions in the low-level intermediate representation. Based on the type length indicated in the name of the undefined function, the undefined function is corrected to a statement with a valid type definition. For example, assuming an undefined function "init_int32(ptr)" is identified, the type length indicated by "int32" in its name is likely to be 4 bytes of type int. Based on this, the function is corrected to the following statement with a valid type definition: "memset(ptr,0,4)".
[0112] It should be noted that before proposing the solution of the present application, the applicant had tried the Mcsema simulation intermediate representation and the compiled intermediate representation based on Plankton or RetDec. However, since they also had similar problems and were not convenient for source code level analysis, the solution of the present application was proposed to solve their problems together.
[0113] Regarding this emulation-based intermediate representation: Since variable entities are not restored, McSema uses global or local variables to simulate hardware registers (for example, eax and rbx). Register state is passed through load and store operations, and memory reads and writes are represented by the global memory array (struct.Memory). In addition, McSema emulates the calling convention by explicitly managing the stack and registers. Function parameters and return values are passed through virtual registers or stack addresses. Therefore, this emulation-based intermediate representation can only guarantee a certain level of functional correctness, meaning that symbolic execution can be performed to a certain extent but static analysis cannot be performed.
[0114] For this compiled intermediate representation: the code structure and variable entities are restored, but the variable and function types, especially the pointer types, are not restored. In addition, the high-level structure of the variables (arrays, structures) is not correctly restored, which leads to false positives and missed negatives in static analysis. For example, the function return value is not identified as a pointer. This misclassification will cause false positives in the static analysis engine. This is because the engine relies on the following assumption: if the return type of the function is not a pointer, the memory must be free within the current function to avoid never-free vulnerabilities. In addition, in order to restore the variable entity and code structure, undefined behavior is also introduced. The corresponding function may not be parsed by an external analyzer, thus making symbolic execution unable to analyze.
[0115] In order to improve the execution efficiency of the solution of the present application, the compiled intermediate representation can be used as the low-level intermediate representation in the present application. In order to solve the above problems, through the solution of the present application, the compiled intermediate representation can be further analyzed to restore the variable and function types. After the source code is compiled into a binary program, semantic information such as type and name is lost. The present application uses standard library function calls to recover the lost source-level information. The standard library function has a standard document record for the function. Therefore, at the call point of the library function, the return value type and parameter type of the function are determined. This is used as the anchor point for reasoning to perform reverse data dependency tracking and perform type correction on the associated variables and the functions that process the variables; further, when the anchor variable performs reverse data flow, different branches (function call chains) will be formed at the end, because the call chain can be reasoned and analyzed one by one. When there is information correction, the correction information is synchronized to other call chains for iterative correction.
[0116] In order to facilitate more intuitive understanding and implementation, one or more embodiments of this specification also provide Figure 1 A flow chart of a specific embodiment of the method is shown in FIG. Figure 4 .
[0117] Figure 4 The process mainly includes four steps: preprocessing, anchor variable determination, analysis context collection and intermediate representation optimization.
[0118] For the preprocessing part.
[0119] The format of the binary program file (e.g., ELF or PE, etc.) and its underlying architecture (e.g., x86 or ARM, etc.) are analyzed. Subsequently, the machine code is disassembled and translated into the underlying virtual machine intermediate representation (LLVM IR).
[0120] For the anchor variable determination part.
[0121] The following steps may be involved:
[0122] 1. Locate the function call point: Search for standard library function calls (such as fopen, malloc, read) in the intermediate representation file of the underlying virtual machine (the corresponding .ll file) according to the function name.
[0123] 2. Extract anchor variables: Identify variables related to library function call points as anchor variables (AV). Pay special attention to pointer type variables. Because pointer type variables in the intermediate representation obtained by preprocessing may lose semantics or even types during subsequent processing, affecting the performance of downstream analysis. For example, pointers may be converted to integers during subsequent processing.
[0124] 3. Record anchor variables: Each anchor variable is recorded as a triple: av = (location, name, function), where location represents the code line where the variable is located; name represents the variable name; and function represents the standard library function to which the variable corresponds. All anchor variables are combined into an AV list in binary form, which serves as input to the next stage.
[0125] For the analysis context collection part.
[0126] The analysis context is collected by performing intra-procedural and inter-procedural data dependency analysis. The goal is to trace the flow of data associated with an anchor variable along the function call chain.
[0127] Beginning with Define-Use analysis in the function that defines the anchor variable, the data flow within the function is traced, and the dependency relationship between the anchor variable and the function's parameters or return value is examined. Transition variable identification for interprocedural analysis: If the anchor variable depends on a function parameter or return value, the transition variable corresponding to that parameter or return value is identified at the call site of the caller function. The relationship between the transition variable and the parameter or return value in the caller function is then analyzed. Further iterative dependency analysis is performed, recursively tracing this dependency until no further dependencies are found. A depth-first traversal is performed to form function call chains, identifying all functions that directly or indirectly depend on the anchor variable. Ultimately, a set of call chains corresponding to the anchor variable av is formed and stored in the dictionary CC[av] for the next stage of analysis.
[0128] For the intermediate representation optimization part.
[0129] The analysis context (including information such as anchor variables and transition variables) is used to infer the missing source code level information and correct the low-level virtual machine intermediate representation obtained by preprocessing.
[0130] This part consists of three subtasks: pattern matching, intermediate representation modification, and information synchronization. Functions in a function call chain are treated as nodes, and iterative intermediate representation optimization is performed. First, the intermediate representation is optimized in the anchor variable scope function. Then, the intermediate representation is optimized based on the transition variables in the caller function until the analysis of a call chain is completed. Finally, the modified information is synchronized with other call chains. All anchor variables and corresponding call chains are traversed in this manner.
[0131] For the pattern matching subtask:
[0132] Identify semantic and structural patterns in the context of anchor variables. The goal is to match these patterns with known behaviors of standard library functions. For I / O functions such as fgets, fputs, and fclose that handle streams or files, match keywords (such as stream, file, and stdout) in the caller function name and the included hard-coded strings to infer that the variables in these contexts are file pointers. For memory management functions such as malloc, realloc, and free that involve dynamic memory management, combine the return value of the dynamic memory management function with the caller function name and the keywords (such as alloc and memdup) in the included hard-coded strings to determine that the type of the relevant variables and return values should be pointers. For memory operations such as read, memcpy, and strncpy that have length constraints, analyze their output buffer pointers and length constraint parameters. If the output buffer is subsequently accessed in a truncated manner, correct the definition and reference of the relevant variables to array mode.
[0133] In addition, for undefined functions: This part of the processing is not related to anchor variables, but is used to find undefined functions used in the file, which may break the functionality of the intermediate representation, especially when recompiled. Therefore, during the initial analysis, the entire intermediate representation file can be scanned to determine the statements containing the undefined keyword.
[0134] For the intermediate representation correction subtask:
[0135] The intermediate representation is fixed by resolving definitions and references of undefined functions and matching anchor variables. First, during the initial analysis, undefined functions are identified and fixed throughout the file. Undefined functions often involve variable initialization and are therefore replaced with valid statements based on the type length indicated in the function name.
[0136] Afterwards, intra-procedural and inter-procedural analysis is performed to heuristically correct the definitions and reference statements of anchor variables. This process is repeated for each anchor variable and its corresponding call chain.
[0137] Intra-procedural corrections include: using definition-use analysis to check the scope of anchor variables, definitions of variables within functions, and references to variables within functions. If the function's return type is incorrect (for example, it should actually return a pointer), the return type is modified accordingly, and then the dependency between the anchor variable and the function's return value is corrected.
[0138] Interprocedural corrections include: performing interprocedural analysis after intraprocedural analysis is complete to update function calls in the calling function to reflect the corrected types, and modifying global references if necessary. For function types: when the return value or parameter type changes, the correction is propagated across function boundaries. First, the variable corresponding to the function type change at the call site is analyzed, and the type change information is updated to this variable. Then, based on the updated variable, the type update analysis is continued within the function where the call site is located. This process is repeated until the call chain analysis is completed or the function type is no longer affected. For global variables: For anchor variables of file pointer type, whether they depend on any global variables is tracked. If the dependent global variables are not found in the function where the anchor variable is located, the function parameters that the anchor variable depends on along the function call chain will continue to be tracked.
[0139] For the information synchronization subtask:
[0140] After completing the intermediate representation correction, ensure that the changes are applied consistently throughout the program. Specifically, common nodes in different call chains are updated, and modifications to global variables are propagated to all reference locations. Changes in function types, variable definitions, and references are marked to ensure that the impact of the changes can be iteratively analyzed while analyzing other call chains.
[0141] This synchronization process ensures that all function and variable modifications are consistently applied throughout the entire binary program file, allowing the binary program file to be accurately and completely converted into a higher-level intermediate representation of the underlying virtual machine.
[0142] After obtaining this higher-level intermediate representation of the underlying virtual machine, existing source code-level analysis algorithms can be used more conveniently and efficiently to further perform processing such as static analysis, symbolic execution, and reanalysis.
[0143] Based on the same idea, one or more embodiments of this specification also provide devices and apparatuses corresponding to the above methods, such as Figure 5 、 Figure 6 The apparatus and device can accordingly execute the above method and related optional solutions.
[0144] Figure 5 A schematic structural diagram of a binary program processing device provided for one or more embodiments of this specification, the device comprising:
[0145] A low-level intermediate representation acquisition module 502 acquires a low-level intermediate representation obtained by disassembling the target binary program;
[0146] An anchor variable determination module 504 identifies call sites of well-defined library functions in the low-level intermediate representation and determines anchor variables in the low-level intermediate representation based on variables associated with the call sites;
[0147] A context acquisition module 506 acquires context related to the anchor variable in the low-level intermediate representation;
[0148] A missing information inference module 508 infers the missing source code information of the low-level intermediate representation based on the context related to the anchor variable by matching it with the known behavior of the library function;
[0149] The high-level intermediate representation generation module 510 modifies the low-level intermediate representation according to the missing source code level information to generate a high-level intermediate representation.
[0150] Optionally, the anchor variable determination module 504 determines one or more standard library functions with clear types and semantics as target standard library functions;
[0151] identifying, in the low-level intermediate representation, a call site of the target standard library function;
[0152] One or more pointer type variables related to the call site are extracted as anchor variables in the low-level intermediate representation.
[0153] Optionally, the target standard library function includes at least one of the following standard library functions: fopen, malloc, and read.
[0154] Optionally, the anchor variable determination module 504 determines, for the variable related to the call point, whether the variable has clear operational semantics;
[0155] If yes, calculate the degree of deviation between the operational semantics and the known behavior based on the known behavior of the library function;
[0156] If the degree of deviation is greater than a set threshold, the variable is determined as an anchor variable candidate variable, so that an anchor variable is selected from the anchor variable candidate variables.
[0157] Optionally, the context acquisition module 506 determines a function parameter or return value that the anchor variable depends on in the low-level intermediate representation;
[0158] For a call point of a caller function of a function to which the function parameter or return value belongs, identifying a variable in the caller function corresponding to the function parameter or return value as a transition variable;
[0159] Performing iterative dependency analysis based on the transition variable, finding functions that directly or indirectly depend on the anchor variable in the low-level intermediate representation, and forming a function call chain;
[0160] According to the function call chain, a context related to the anchor variable is determined.
[0161] Optionally, the missing information inference module 508 matches the caller function name and hard-coded string involved in the context associated with the anchor variable with known behaviors of the library function;
[0162] According to the result of successful matching, it is inferred that the variable or return value in the low-level intermediate representation is missing a pointer type.
[0163] Optionally, the missing information inference module 508 matches the behavior of an I / O library function in processing streams or files, and / or the behavior of a memory management library function in performing dynamic memory management.
[0164] Optionally, the missing information inference module 508 matches the context related to the anchor variable with a memory operation behavior with a length constraint of a memory operation library function;
[0165] Through the matching, if it is determined that the corresponding output buffer pointer and length constraint parameters exist in the context related to the anchor variable, and the output buffer is accessed in a truncated manner, then the definition and reference missing array pattern of the related variables in the low-level intermediate representation is inferred.
[0166] Optionally, the high-level intermediate representation generation module 510 modifies the type of the parameter or return value of the scope function of the anchor variable in the low-level intermediate representation according to the missing source code level information;
[0167] According to more caller functions involved in the anchor variable, the correction is propagated across function boundaries according to the correction, so that the types of other variables that have a dependency relationship with the anchor variable are also corrected.
[0168] Optionally, the high-level intermediate representation generation module 510 determines a statement containing an undefined keyword in the low-level intermediate representation to identify an undefined function in the low-level intermediate representation;
[0169] According to the type length indicated in the name of the undefined function, the undefined function is corrected to a statement with a valid type definition.
[0170] Optionally, the low-level intermediate representation includes: a low-level underlying virtual machine intermediate representation.
[0171] Figure 6 A schematic diagram of the structure of a binary program processing device provided for one or more embodiments of this specification, the device comprising:
[0172] at least one processor; and,
[0173] a memory communicatively connected to the at least one processor; wherein,
[0174] The memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor to enable the at least one processor to perform:
[0175] Obtaining a low-level intermediate representation obtained by disassembling the target binary program;
[0176] In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site;
[0177] In the low-level intermediate representation, obtaining a context related to the anchor variable;
[0178] Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions;
[0179] The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.
[0180] Based on the same idea, one or more embodiments of this specification further provide a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:
[0181] Obtaining a low-level intermediate representation obtained by disassembling the target binary program;
[0182] In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site;
[0183] In the low-level intermediate representation, obtaining a context related to the anchor variable;
[0184] Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions;
[0185] The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.
[0186] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0187] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0188] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0189] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0190] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0191] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0192] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0194] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0195] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0196] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0197] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0198] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0199] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.
[0200] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0201] The foregoing description is merely one or more embodiments of this specification and is not intended to limit this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of one or more embodiments of this specification are intended to be within the scope of the claims of this specification.
Claims
1. A binary program processing method, comprising: Obtaining a low-level intermediate representation obtained by disassembling the target binary program; In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site; In the low-level intermediate representation, obtaining a context related to the anchor variable; Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions; The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.
2. The method of claim 1 , wherein identifying call sites of well-defined library functions in the low-level intermediate representation and determining anchor variables in the low-level intermediate representation based on variables associated with the call sites comprises: Determine one or more standard library functions with clear types and semantics as target standard library functions; identifying, in the low-level intermediate representation, a call site of the target standard library function; One or more pointer type variables related to the call site are extracted as anchor variables in the low-level intermediate representation.
3. The method according to claim 2, wherein the target standard library function comprises at least one of the following standard library functions: fopen, malloc, and read.
4. The method of claim 1 , wherein determining the anchor variable in the low-level intermediate representation based on the variables associated with the call site comprises: For variables related to the call point, determining whether the variables have clear operational semantics; If yes, calculate the degree of deviation between the operational semantics and the known behavior based on the known behavior of the library function; If the degree of deviation is greater than a set threshold, the variable is determined as an anchor variable candidate variable, so that an anchor variable is selected from the anchor variable candidate variables.
5. The method according to claim 1, wherein obtaining the context related to the anchor variable in the low-level intermediate representation specifically comprises: determining a function parameter or a return value on which the anchor variable depends in the low-level intermediate representation; For a call point of a caller function of a function to which the function parameter or return value belongs, identifying a variable in the caller function corresponding to the function parameter or return value as a transition variable; Performing iterative dependency analysis based on the transition variable, finding functions that directly or indirectly depend on the anchor variable in the low-level intermediate representation, and forming a function call chain; According to the function call chain, a context related to the anchor variable is determined.
6. The method of claim 1 , wherein inferring missing source code information of the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions, specifically comprises: Matching the caller function name and hard-coded string involved in the context associated with the anchor variable with the known behavior of the library function; According to the result of successful matching, it is inferred that the variable or return value in the low-level intermediate representation is missing a pointer type.
7. The method according to claim 6, wherein the matching with the known behavior of the library function specifically comprises: Match the behavior of I / O library functions for handling streams or files, and / or the behavior of memory management library functions for dynamic memory management.
8. The method of claim 1 , wherein inferring missing source code information of the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions, specifically comprises: Matching the context related to the anchor variable with the memory operation behavior with length constraints of the memory operation library function; Through the matching, if it is determined that the corresponding output buffer pointer and length constraint parameters exist in the context related to the anchor variable, and the output buffer is accessed in a truncated manner, then the definition and reference missing array pattern of the related variables in the low-level intermediate representation is inferred.
9. The method of claim 1 , wherein the correcting the low-level intermediate representation based on the missing source-level information comprises: According to the missing source code level information, modify the type of the parameter or return value of the scope function of the anchor variable in the low-level intermediate representation; According to more caller functions involved in the anchor variable, the correction is propagated across function boundaries according to the correction, so that the types of other variables that have a dependency relationship with the anchor variable are also corrected.
10. The method of claim 1, further comprising: Determining a statement containing an undefined keyword in the low-level intermediate representation to identify undefined functions in the low-level intermediate representation; According to the type length indicated in the name of the undefined function, the undefined function is corrected to a statement with a valid type definition.
11. The method according to any one of claims 1 to 10, wherein the low-level intermediate representation comprises: Low-level intermediate representation of the underlying virtual machine.
12. A binary program processing device, comprising: A low-level intermediate representation acquisition module obtains a low-level intermediate representation obtained by disassembling the target binary program; an anchor variable determination module, which identifies call sites of well-defined library functions in the low-level intermediate representation and determines anchor variables in the low-level intermediate representation based on variables associated with the call sites; A context acquisition module, which acquires context related to the anchor variable in the low-level intermediate representation; a missing information inference module, which infers the missing source code level information of the low-level intermediate representation based on the context related to the anchor variable by matching it with the known behavior of the library function; The high-level intermediate representation generation module amends the low-level intermediate representation according to the missing source code level information to generate a high-level intermediate representation.
13. The apparatus according to claim 12, wherein the anchor variable determination module determines one or more standard library functions with clear types and semantics as target standard library functions; identifying, in the low-level intermediate representation, a call site of the target standard library function; One or more pointer type variables related to the call site are extracted as anchor variables in the low-level intermediate representation.
14. The apparatus of claim 12, wherein the missing information inference module matches a caller function name and a hard-coded string involved in the context associated with the anchor variable with known behaviors of the library function; According to the result of successful matching, it is inferred that the variable or return value in the low-level intermediate representation is missing a pointer type.
15. The apparatus of claim 12, wherein the missing information inference module matches the context related to the anchor variable with a memory operation behavior with a length constraint of a memory operation library function; Through the matching, if it is determined that the corresponding output buffer pointer and length constraint parameters exist in the context related to the anchor variable, and the output buffer is accessed in a truncated manner, then the definition and reference missing array pattern of the related variables in the low-level intermediate representation is inferred.
16. The apparatus of claim 12, wherein the high-level intermediate representation generation module determines a statement containing an undefined keyword in the low-level intermediate representation to identify an undefined function in the low-level intermediate representation; According to the type length indicated in the name of the undefined function, the undefined function is corrected to a statement with a valid type definition.
17. The apparatus according to any one of claims 12 to 16, wherein the low-level intermediate representation comprises: Low-level intermediate representation of the underlying virtual machine.
18. A binary program processing device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor to enable the at least one processor to perform: Obtaining a low-level intermediate representation obtained by disassembling the target binary program; In the low-level intermediate representation, identifying a call site with a well-defined library function, and determining an anchor variable in the low-level intermediate representation based on variables associated with the call site; In the low-level intermediate representation, obtaining a context related to the anchor variable; Inferring missing source code level information from the low-level intermediate representation by matching the context associated with the anchor variable with known behaviors of library functions; The low-level intermediate representation is modified according to the missing source-level information to generate a high-level intermediate representation.