Function name recovery method and electronic device

By generating and analyzing the target pseudocode, extracting structural information and converting it into three address codes, combined with LLVM IR optimization, the problem of difficult function names in binary files is solved, and accurate function name recovery and software security guarantee are achieved.

CN119357963BActive Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907503.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-07-18
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In the prior art, software programs remove or obfuscate function names and variable information during compilation into binary files, making it difficult to recover function names in the source code, and fail to accurately locate potential security risks, affecting software security.

Method used

By obtaining the target binary file, generating the target pseudocode, extracting the structure information of the function, converting it into three address codes, and generating function names based on the feature information related to the control flow, using advanced pseudocode to reduce the interference of the underlying architecture information, and using LLVM IR for cross-platform optimization and analysis.

Benefits of technology

It realizes accurate recovery of function names, improves the accuracy and efficiency of function name recovery, can effectively locate key dangerous functions, and ensures software security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357963B_ABST
    Figure CN119357963B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a function name recovery method and an electronic device, which are applicable to the field of software security technology. The function name recovery method includes: obtaining a target binary file; generating target pseudocode corresponding to the target binary file; extracting structural information of a target function from the target pseudocode, where the structural information is used to describe a function, basic blocks in the function, and variables in the function, and the target function is any function in the target pseudocode; based on the structural information of the target function, converting first pseudocode in the target pseudocode into three-address code, where the first pseudocode is related to the target function; determining first feature information related to the control flow of the target function based on the three-address code; and generating the function name of the target function based on the first feature information. Embodiments of the present application can accurately recover the function names of functions in a software program, thereby ensuring software security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software security technologies, and in particular, to a method for restoring function names and an electronic device. Background Art

[0002] Binary reverse analysis is a technology that analyzes the binary files (machine code) of software to understand the behavior and logic of the software. Binary reverse analysis is widely used in various fields such as software security, malware analysis, vulnerability detection, and legacy system maintenance. Since accurate function names can help quickly locate potential security risk points, such as identifying dangerous critical functions, function name restoration is an important task in binary reverse analysis work.

[0003] In related technologies, during the process of compiling a software program from source code into a binary file, the symbol information in the binary file is usually removed or obfuscated. For example, function names and / or variable information are removed, making it difficult to restore the function names in the source code, and thus unable to discover the dangers in the software through function names, making it difficult to ensure software security. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method for restoring function names, an electronic device, a computer-readable storage medium, a chip system, and a computer program product, which can accurately restore the function names of each function in a software program, thereby ensuring software security.

[0005] In a first aspect, embodiments of this application provide a method for restoring function names. This method can be applied to an electronic device, which can be a device such as a terminal or a server. The terminal can be a tablet computer, a vehicle-mounted device, an Augmented Reality (AR) / Virtual Reality (VR) device, a laptop computer, an Ultra-Mobile Personal Computer (UMPC), etc. The server can be a network server, a cloud server, etc. Embodiments of this application do not make any limitations in this regard.

[0006] The method for restoring function names may include: First, the electronic device obtains a target binary file. Then, the electronic device generates target pseudocode corresponding to the target binary file. After that, the electronic device extracts the structural information of the target function from the target pseudocode. Next, the electronic device converts the first pseudocode in the target pseudocode into three-address code based on the structural information of the target function. Finally, the electronic device determines first feature information related to the control flow of the target function based on the three-address code, and generates the function name of the target function based on the first feature information.

[0007] Among them, the structure information is used to describe functions, basic blocks in functions, and variables in functions. The target function (which can also be called the first function) is any function in the target pseudocode. It should be noted that the target pseudocode may include one function or multiple functions. In the case where the target pseudocode includes multiple functions, the electronic device can extract the structure information of each function and generate the function name of the function based on the structure information of each function.

[0008] Among them, the above first pseudocode is related to the target function. That is to say, the first pseudocode is a part of the target pseudocode.

[0009] Optionally, the target binary file is a binary file from which symbol information has been removed. In this application, a binary file refers to a file that carries the machine code of a program. The file format of the binary file may include: object file format (.o file), shared object file format (.so file), executable file format (.exe file), and so on. It can be understood that the file format of the target binary file is not specifically limited in the embodiments of this application.

[0010] In the embodiments of this application, on the one hand, the structure information of a function can accurately describe the characteristics of the function, accurately extract rich and comprehensive structure information of each function, such as function information, basic block information in the function, and variable information in the function, and generate the function name of the function based on the structure information of the function, which helps to improve the accuracy of function name recovery. On the other hand, in the case of accurately recovering the function name, key dangerous functions can be filtered and directly located according to the function name, which helps to ensure software security. On the third hand, only the first pseudocode related to the function structure needs to be converted into three-address code, which can avoid converting the pseudocode unrelated to the function structure and can improve the conversion efficiency.

[0011] Optionally, the first feature information related to the control flow may include the features characterized by a control flow graph (CFG), or may include one or more of the following: the features characterized by a call graph (CG), the features characterized by a data flow graph (DFG), and the features characterized by a program dependence graph (PDG).

[0012] It can be understood that the electronic device can adopt the solutions disclosed in related technologies to convert the three-address code into a control flow graph, a call graph, a data flow graph, a program dependence graph, etc., which will not be elaborated here.

[0013] It should be noted that extracting the characteristics of a function from multiple dimensions can make the extracted characteristics comprehensive and rich, which helps to further improve the accuracy of function name recovery.

[0014] In the first possible implementation of the first aspect, the structural information of the target function may include function information, basic block information in the target function, and variable information in the target function.

[0015] In this case, for the electronic device to extract the structural information of the target function from the target pseudocode, it may include: First, the electronic device traverses the target pseudocode. When accessing the first operation instruction, it extracts the function information of the target function from the first operation instruction. Then, based on each instruction in the target function, the electronic device determines M basic blocks and the basic block information of each basic block in the target function, where M is an integer greater than 0. After that, the electronic device traverses the first basic block and extracts variable information from the first instruction in the first basic block. Here, the first basic block is any one of the M basic blocks, and the first instruction is any instruction in the first basic block. Finally, the electronic device generates an intermediate file, which is used to record the function information, basic block information, and variable information of the target function.

[0016] Among them, the first operation instruction is the instruction that defines the target function in the target pseudocode, and the function information includes the first entry address of the function. The first entry address can also be referred to as the entry address of the function or the starting address of the function. It can be understood that the target pseudocode is composed of instructions, and the instructions in the pseudocode can also be called pseudocode instructions. Each instruction in the pseudocode corresponds to an operation, such as an AND operation, an OR operation, etc. Therefore, the instructions in the pseudocode can also be called operation instructions.

[0017] Among them, the basic block information includes the second entry address of the basic block, the forward node of the basic block, and the backward node of the basic block. The second entry address can also be referred to as the entry address of the basic block or the starting address of the basic block. It should be noted that for any basic block, the forward node of this basic block is the pseudocode instruction (P-Code instruction) that jumps to this basic block, and the backward node of this basic block is the P-Code instruction that this basic block needs to jump to.

[0018] Among them, the first entry address and the second entry address in the intermediate file correspond to the first pseudocode. The first entry address is the storage address of the first pseudocode instruction in the function, and the second entry address is the storage address of the first pseudocode instruction in the basic block. The electronic device can access each pseudocode instruction in the function based on the first entry address and the second entry address.

[0019] In the embodiments of the present application, since the pseudocode has a specific syntax, the electronic device can extract function information, basic block information in the function, and variable information in the function from the target code in combination with the syntax of the pseudocode. It should be noted that since the pseudocode is essentially unordered code, extracting the structural information of each function from the pseudocode can achieve the orderly classification of the unordered pseudocode. In addition, classifying and dividing the pseudocode layer by layer according to multiple levels such as functions, basic blocks in the functions, instructions in the basic blocks, and variables helps to accurately extract the semantic structure of each function, thereby ensuring the accuracy of function name recovery.

[0020] Optionally, the function information may further include: the name of the function (abbreviated as function name), the type of the return value of the function, the size of the return value of the function, etc.

[0021] Optionally, the basic block information may further include the index of the basic block. The index of the basic block is the unique identifier of the basic block, and each basic block has an index. For example, the index of the first basic block in the function can be 0, and the index of the next basic block can be 1.

[0022] In the second possible implementation manner of the first aspect, the variable information includes one or more of the following: the name of the variable, the storage type of the variable, the size of the variable.

[0023] Among them, the storage type of the variable can indicate the storage location of the variable. For example, the register type means that the variable is stored in the register, the memory type means that the variable is stored in the memory, and the stack type means that the variable is stored in the stack.

[0024] In the third possible implementation manner of the first aspect, the language format of the three-address code is the low-level virtual machine intermediate representation LLVM IR. In this case, the electronic device converts the first pseudocode in the target pseudocode into three-address code based on the structural information of the target function, which may include: First, the electronic device generates an initial variable table corresponding to the target function, and the variable table is used to record the value change situation of each variable in the target function during the program operation. Then, the electronic device traverses the first pseudocode based on the intermediate file, and for the second instruction in the first pseudocode, converts the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table.

[0025] Among them, the second instruction is any instruction in the first pseudocode, and the operation types include: the first operation type, the second operation type, the third operation type, the fourth operation type. The first operation type is used to obtain input variables, the second operation type is used to update output variables, the third operation type is used to merge multiple branches, and the fourth operation type is an operation type other than the first operation type, the second operation type, and the third operation type.

[0026] It should be noted that for any function, a variable table can be corresponding, and this variable table is used to record the value change situations of variables in the function during the program running process. The variables recorded in the variable table can include global variables, local variables, temporary variables, etc. Optionally, the initial variable table corresponding to the function can be an empty table by default.

[0027] In the embodiment of the present application, on the first hand, since LLVM IR is an intermediate representation independent of a specific hardware platform, it converts P-Code into a more standardized lower-level representation, and a unified representation independent of platforms and languages can be obtained, which is convenient for analysis and optimization. Moreover, by first converting the intermediate file into an LLVM IR file, the powerful functions of LLVM can be utilized for cross-platform optimization and analysis, and the generation process of the control flow graph can be simplified, making the analysis and optimization more efficient and general. On the second hand, converting the pseudo-code instructions of various operation types described in the intermediate file into corresponding LLVM IR instructions can ensure that each pseudo-code instruction involved in the function is converted into an LLVM IR instruction, that is, it can achieve a comprehensive and accurate conversion of the instructions involved in the function. In this way, it helps to comprehensively and accurately extract the feature information of the function, thereby improving the accuracy rate of function name recovery. On the third hand, since different types of variables have different storage locations, uses, scopes, and life cycles, and as the program runs, the values of variables in each function will change. In the present application, by using the variable table of the function to describe the real-time change situations of each variable in the function, the accuracy of the generated LLVM IR instructions can be ensured. For example, if a variable is not recorded in the variable table when it is updated, it indicates that there is a problem with the use of this variable. Therefore, based on the real-time updated variable table, the accuracy of the generated LLVM IR instructions can be ensured.

[0028] In the fourth possible implementation manner of the first aspect, the electronic device converts the second instruction into an LLVM IR instruction according to the operation type and / or variable table corresponding to the second instruction, which may include: First, if the operation type of the second instruction is the first operation type, the electronic device can obtain the storage type of the input variable from the second instruction. Then, the electronic device can generate an LLVM IR instruction corresponding to the second instruction according to the storage type of the input variable and the variable table.

[0029] Among them, the storage type may include register type, memory type, and stack type.

[0030] Among them, if the storage type of the input variable is the register type, the LLVM IR instruction is the instruction used to load the value into the input variable. If the storage type of the input variable is the memory type, the LLVM IR instruction is the instruction used to obtain the value of the input variable. If the storage type of the input variable is the stack type, the LLVM IR instruction is the instruction indicating to obtain the value of the input variable from the variable table.

[0031] In the embodiment of the present application, the electronic device can convert the second instruction into a corresponding LLVM IR instruction in combination with the operation type of the instruction, and during the instruction conversion process, it can combine the storage type of the input variable to specifically convert the second instruction into an LLVM IR instruction, which can improve the accuracy of instruction conversion and further improve the function name recovery efficiency.

[0032] In the fifth possible implementation manner of the first aspect, after the electronic device generates the LLVM IR instruction corresponding to the second instruction according to the storage type of the input variable and the variable table, it can also add the first variable information of the input variable to the variable table, and the first variable information includes the variable name of the input variable, the address of the input variable, and the value of the input variable.

[0033] In the embodiment of the present application, when the electronic device obtains the input variable, adding the variable information of the input variable to the variable table in a timely manner can enable the variable table to always record the real-time changes of each variable.

[0034] Optionally, when the input variable is of the register type, during the process of the electronic device generating the LLVM IR instruction, if the electronic device detects that the input variable does not exist in the current variable table, it can generate a first prompt message, and the first prompt message is used to prompt that the input variable is not defined, or there is a syntax error in the use of the input variable. That is to say, in the present application, the variable table can be used to check whether there are syntax errors in the source code.

[0035] In the sixth possible implementation manner of the first aspect, the electronic device converting the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table may include: First, if the operation type of the second instruction is the second operation type, obtain the storage type of the output variable from the second instruction. Then, the electronic device generates the LLVM IR instruction corresponding to the second instruction according to the storage type of the output variable.

[0036] Among them, if the storage type of the output variable is a register type, the LLVM IR instructions include a first LLVM IR instruction and a second LLVM IR instruction. The first LLVM IR instruction is used to convert the data type of the output variable into a first data type, and the second LLVM IR instruction is used to store the output variable of the first data type. The first data type is the data type corresponding to the output variable. If the storage type of the output variable is a memory type or a stack type, the LLVM IR instruction is an instruction for storing the output variable.

[0037] In the embodiments of the present application, the electronic device can convert the second instruction into a corresponding LLVM IR instruction in combination with the operation type of the instruction, and in the process of instruction conversion, it can also combine the storage type of the output variable to specifically convert the second instruction into an LLVM IR instruction, which can improve the accuracy of instruction conversion and thus further improve the function name recovery efficiency.

[0038] In the seventh possible implementation manner of the first aspect, after the electronic device generates the LLVM IR instruction corresponding to the second instruction according to the storage type of the output variable, it can also add the second variable information of the output variable to the variable table. The second variable information includes the variable name of the output variable, the address of the output variable, and the value of the output variable.

[0039] In the embodiments of the present application, the electronic device timely adds the latest variable information of the output variable to the variable table, which can realize that the variable table always records the real-time changes of each variable.

[0040] In the eighth possible implementation manner of the first aspect, the electronic device converts the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table, including: First, if the operation type of the second instruction is a third operation type, obtain the N input variables included in the second instruction, where N is an integer greater than 1. Then, the electronic device determines the forward basic block of the second basic block based on the intermediate file. The second basic block is the basic block to which the second instruction belongs, and the forward basic block is the basic block that jumps to the second basic block. After that, the electronic device obtains the latest value of the first input variable from the first forward basic block, where the first forward basic block is the forward basic block containing the first input variable, and the first input variable is any one of the input variables in the second instruction. Finally, the electronic device generates an LLVM IR instruction according to the obtained values of the input variables.

[0041] In the embodiments of the present application, when the operation type of the second instruction is the third operation type, the electronic device can accurately and effectively convert the second instruction in the pseudocode into an LLVM IR instruction. During the instruction conversion process, by combining the operation type of the instruction to specifically convert the second instruction into an LLVM IR instruction, the accuracy of instruction conversion can be improved, thereby further improving the function name recovery efficiency.

[0042] In a ninth possible implementation manner of the first aspect, the electronic device converts the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table, including: if the operation type corresponding to the second instruction is the fourth operation type, then based on a preset first correspondence, the second instruction is converted into an LLVM IR instruction, where the first correspondence is the correspondence between the instructions of the pseudocode and the LLVM IR instructions.

[0043] In the embodiments of the present application, during the instruction conversion process, the electronic device combines the operation type of the instruction to specifically convert the second instruction into an LLVM IR instruction, which can improve the accuracy of instruction conversion, thereby further improving the function name recovery efficiency.

[0044] In a tenth possible implementation manner of the first aspect, the electronic device generates the function name of the target function based on the first feature information, including: First, the electronic device inputs the first feature information into the first model to obtain multiple candidate function names of the target function. Then, the multiple candidate function names, the first feature information, and the three-address code are input into the second model to obtain the function name of the target function.

[0045] Wherein, the first model is a generative model, and the second model is a large language model.

[0046] In the embodiments of the present application, since the generative model can process various complex data types, such as text, images, audio, etc., and can learn the latent representation of the data and adapt to different tasks. In addition, the generative model usually has strong generalization ability and can perform well even in scenarios outside the training data. Therefore, when the first model is a generative model, the accuracy of the obtained multiple candidate function names can be guaranteed. In addition, since the large prediction model is a natural language processing model based on deep learning and can learn and understand the grammar and semantics of natural language, using the large language model to optimize the multiple candidate function names output by the first model helps to further improve the accuracy of function name recovery.

[0047] In an eleventh possible implementation manner of the first aspect, the electronic device generates target pseudocode corresponding to the target binary file, including: generating high-level pseudocode corresponding to the target binary file, and the target pseudocode is high-level pseudocode.

[0048] In the embodiments of the present application, since high-level pseudocode pays more attention to the source code itself and does not focus on the underlying architecture, while low-level pseudocode contains a large amount of information related to the underlying architecture, and this information related to the underlying architecture has nothing to do with the function itself. Therefore, using high-level pseudocode for function name recovery can avoid excessive irrelevant interference information, help reduce the amount of data processing, and improve the efficiency of function name recovery.

[0049] In a second aspect, the embodiments of the present application provide a function name recovery device, which includes:

[0050] An input unit, configured to obtain a target binary file;

[0051] A processing unit, configured to generate target pseudocode corresponding to the target binary file; extract structural information of a target function from the target pseudocode, where the structural information is used to describe the function, basic blocks in the function, and variables in the function, and the target function is any function in the target pseudocode; based on the structural information of the target function, convert the first pseudocode in the target pseudocode into three-address code, where the first pseudocode is related to the target function; based on the three-address code, determine first feature information related to the control flow of the target function; and generate the function name of the target function based on the first feature information.

[0052] As an embodiment of the present application, the function name recovery device can implement the method according to any one of the above first aspects.

[0053] In a third aspect, the embodiments of the present application provide an electronic device, which includes a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the method according to any one of the above first aspects is implemented.

[0054] In a fourth aspect, the embodiments of the present application provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method according to any one of the above first aspects is implemented.

[0055] In a fifth aspect, the embodiments of the present application provide a chip system, which includes a processor. The processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the method according to any one of the above first aspects. The chip system can be a single chip or a chip module composed of multiple chips.

[0056] In a sixth aspect, the embodiments of the present application provide a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to execute the method according to any one of the above first aspects.

[0057] It can be understood that the beneficial effects of the second to sixth aspects described above can be referred to the relevant descriptions in the first aspect, and will not be elaborated here. Description of the Drawings

[0058] Figure 1 Schematic diagram of the source code provided by an embodiment of the present application;

[0059] Figure 2 Schematic diagram of the assembly code corresponding to the binary file without removing the static symbol table provided by an embodiment of the present application;

[0060] Figure 3 Schematic diagram of the assembly code corresponding to the binary file with the static symbol table removed provided by an embodiment of the present application;

[0061] Figure 4 Schematic diagram of the effect of function name restoration provided by an embodiment of the present application;

[0062] Figure 5 Schematic diagram of the flow of a function name restoration method provided by an embodiment of the present application;

[0063] Figure 6 Schematic diagram of the process of extracting the structural information of a function provided by an embodiment of the present application;

[0064] Figure 7 Schematic diagram of the process from a binary file to an intermediate file provided by an embodiment of the present application;

[0065] Figure 8 Schematic diagram of the variable part in the intermediate file provided by an embodiment of the present application;

[0066] Figure 9 Schematic diagram of a certain function in the function part of the intermediate file provided by an embodiment of the present application;

[0067] Figure 10 Schematic diagram of the visual structure of the intermediate file provided by an embodiment of the present application;

[0068] Figure 11 Schematic diagram of the visual structure of a certain function provided by an embodiment of the present application;

[0069] Figure 12 Schematic diagram of the flow of instruction conversion provided by an embodiment of the present application;

[0070] Figure 13 Schematic diagram of the conversion process of the branch merge instruction provided by an embodiment of the present application;

[0071] Figure 14 Schematic diagram of directly converting P-Code code into an LLVM IR file;

[0072] Figure 15 Schematic diagram of converting an intermediate file into an LLVM IR file;

[0073] Figure 16 Schematic diagram of the control flow graph provided by the embodiments of the present application;

[0074] Figure 17 Schematic diagram of a process for training a function name generation model provided by the embodiments of the present application;

[0075] Figure 18 Schematic diagram of a process for optimizing function names provided by the embodiments of the present application;

[0076] Figure 19 Schematic diagram of the flow of another function name recovery method provided by the embodiments of the present application;

[0077] Figure 20 Schematic diagram of the structure of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0078] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0079] Some concepts that may be involved in the embodiments of the present application are described as follows:

[0080] (1) Multiple: Unless otherwise specified, in the embodiments of the present application, multiple means two or more.

[0081] (2) Assembly language, machine language, high-level programming language: In the embodiments of the present application, a high-level programming language (abbreviated as high-level language) is a programming language for human users, with a relatively high level of abstraction and closer to the natural language of humans. High-level languages can include C language, C++ language, Python language, Php language, Java language, etc.

[0082] Machine language is the lowest-level language understood and executed by the central processing unit (CPU) in a computer. Machine language is binary code directly recognizable by computer hardware. Machine language consists of multiple instruction codes (which can also be called machine instructions), each instruction code is composed of 0 and 1, and each instruction code corresponds to a basic operation of computer hardware.

[0083] Assembly Language is a machine-oriented low-level programming language that can be applied to electronic computers, microprocessors, microcontrollers, or other programmable devices. In assembly language, mnemonics are used to replace the operation codes of machine instructions, and address symbols or labels are used to replace the addresses of instructions or operands. Therefore, assembly language is also called symbolic language. During the assembly process, assembly language can be converted into machine language. In different devices, the same assembly code (which can also be called an assembly instruction set) corresponds to different machine codes (which can also be called machine instruction sets).

[0084] Generally speaking, machine language directly corresponds to binary instructions, which are difficult to understand and remember. Assembly language uses mnemonics, which is efficient but lacks generality. High-level languages are easy to write and have good portability.

[0085] Although assembly instructions are easier to read and write than binary machine instructions, they are still not as close to natural language as high-level languages. Therefore, in practical applications, usually high-level languages are used to program to obtain source code, and then a special program is used to translate the source code into machine language.

[0086] In the embodiments of the present application, the program code programmed in a high-level language can be called source code, the source code in the form of machine language can be called machine code, and the source code in the form of assembly language can be called assembly code.

[0087] (3) Program compilation, assembly, disassembly, and decompilation: In the embodiments of the present application, program compilation is the process of converting source code written in a high-level programming language (such as C, C++, Java, Python, etc.) into machine code. Machine code is the code of binary machine language, that is, machine code is binary code. The compiled machine code can be directly executed by a computer, while the source code needs to be converted into machine code by a compiler to run.

[0088] The file carrying machine code can be called a binary file. The formats of binary files can include: object file format (.o file), shared object file format (.so file), executable file format (.exe file), and so on.

[0089] When the compilation environment is different, multiple machine codes compiled from the same source code are different. Among them, the compilation environment is related to the type of compiler, operating system, architecture of the operating system, etc.

[0090] Assembly refers to the process of translating assembly language into machine language.

[0091] Disassembly is the process of restoring binary machine code to assembly language mnemonics for easier understanding and analysis. In practice, disassembly tools such as IDA-Pro and Ghidra can be used to disassemble binary machine code.

[0092] Decompilation is the process of converting binary machine code into high-level language code with the aim of restoring the high-level logic and structure of the program to make it more understandable to humans. In practice, decompilation tools such as IDA-Pro and Ghidra can be used to decompile binary machine code.

[0093] In practical applications, disassembly and decompilation are often used in combination. Disassembly can help analyze the low-level structure and control flow of a program, while decompilation can help analyze the high-level logic and data structure of a program.

[0094] (4) Program debugging: In the embodiments of this application, program debugging is the process of finding errors in a program before the compiled program is put into actual operation, with the aim of detecting whether the program can run properly and identifying potential problems. By debugging the program, the robustness of the program can be improved and the probability of program crashes can be reduced. That is to say, program debugging is an important means to improve program quality. Program debugging can correct syntax errors and logical errors in the program.

[0095] (5) Binary reverse engineering: In the embodiments of this application, binary reverse engineering is a technique for reverse engineering binary files (such as executable files, library files, firmware, etc.), aiming to understand the internal structure, functions, and behaviors of the files. Binary reverse engineering is widely used in various fields such as software security, malware analysis, vulnerability detection, and legacy system maintenance.

[0096] Binary reverse engineering is achieved through binary reverse engineering tools. Binary reverse engineering tools are used to convert binary programs (machine code) into assembly programs (assembly language mnemonics) for easier analysis of the program's execution logic.

[0097] Relatively commonly used binary reverse engineering tools can include: Interactive DisAssembler Professional (IDA-Pro) and Ghidra Reverse Engineering Framework. Among them, the Ghidra Reverse Engineering Framework can be abbreviated as Ghidra.

[0098] (6) Static Symbol Table, Dynamic Symbol Table, Debugging Information: In the embodiments of the present application, the static symbol table and the dynamic symbol table are two different symbol tables. The symbol table is used to record symbol information such as function names and variable names.

[0099] The static symbol table is created during compilation and is typically constructed during the compilation phase of the program (especially during syntax analysis and semantic analysis). The static symbol table contains all the statically defined symbols in the program (such as global variables, functions, classes, constants, etc.). The static symbol table remains unchanged throughout the life cycle of the program. The static symbol table is mainly used to manage the metadata of all static symbols in the source code (such as the types, scopes, addresses, etc. of functions and global variables). Its role is to support syntax and semantic analysis during compilation, as well as symbol linking when generating the executable file.

[0100] The dynamic symbol table is created and maintained during program execution. It is usually generated during the execution of the program and stores symbols related to the program execution environment (such as dynamic libraries, function calls, local variables, etc.). The dynamic symbol table is usually dynamically managed by the operating system or the program's runtime system (such as the linker, loader). The dynamic symbol table is mainly used to manage symbols during program execution, especially symbols related to dynamic linking, shared libraries, and function calls. Its role is to provide symbol lookup and mapping functions for dynamic linking and loading during program execution.

[0101] Debugging information refers to the additional data stored in the binary file (such as the executable file or object file) after the program is compiled, which is used to support the debugging process. The debugging information contains metadata that helps in debugging and analyzing the program, which can include function names, variable names, data structures, source file names, and line numbers, etc. The debugging information can help developers trace problems in the source code during development and debugging. The debugging information usually does not directly participate in the execution of the program, but provides a mapping relationship from the source code to the machine code for debugging tools. For example, the debugging information can map a certain line of binary machine code to a specific line in the source code.

[0102] (7) Strip Technology: In computer programming, Strip is a commonly used tool or technology for removing debugging information and symbol tables from binary files, thereby reducing the size of the files.

[0103] After using Strip to remove the symbol table, it becomes more difficult for reverse engineering tools to analyze the internal implementation of the program (such as function names, variable names, etc.), which helps to protect the program code. In addition, for the program code processed by Strip, since unnecessary data is removed, the loading speed and execution efficiency of the program may be slightly improved.

[0104] (8) Symbol table stripping: In the embodiments of the present application, symbol table stripping refers to removing the symbol table information in the program generated by compilation during the software development process. For example, Strip can be used to remove the symbol table information in the program.

[0105] (9) Symbol table obfuscation: In the embodiments of the present application, symbol table obfuscation refers to making the code difficult to be analyzed by reverse engineering by modifying or hiding the symbol table information in the program. By obfuscating the symbol table, the code can be made difficult to understand and analyze, thereby achieving the purpose of protecting the program.

[0106] The methods of symbol table obfuscation mainly include removing symbols and renaming symbols. Among them, removing symbols means using tools such as Strip to remove the symbol table information in the program. Renaming symbols means using meaningless characters or random characters to replace the original names, and renaming symbols such as function names and variable names in the program, so that it is difficult for reverse engineering to understand the function of the code by name.

[0107] (10) Key dangerous functions: In the embodiments of the present application, key dangerous functions usually refer to functions that may cause security problems in programming. If these functions are misused, they may lead to serious security vulnerabilities or data leakage.

[0108] (11) Pseudo Code: In the embodiments of the present application, pseudo code can also be called intermediate representation code or P-Code. P-Code is an abstract intermediate representation code used in the reverse engineering process. The role of P-Code is to convert the underlying machine code into more understandable abstract code, which makes reverse engineering more efficient and concise and can help analysts better understand the high-level logic and behavior of the program.

[0109] Among them, P-Code can include High P-Code and Low P-Code.

[0110] Among them, High P-Code is closer to the representation of high-level languages and is a high-level intermediate representation. High P-Code is usually used to represent the high-level structure and logic of the program. It contains information such as function calls and control flows, and is close to the abstraction of the source code. Since High P-Code is closer to the source code, it usually retains more details related to the program behavior and semantics, such as function parameters, local variables, and conditional judgments. High P-Code can help analysts understand the decompiled code logic.

[0111] Low P-Code is closer to the underlying representation and is a low-level intermediate representation. Low P-Code is closer to the execution details of machine instructions and assembly code. It loses some of the abstraction information in high-level languages and focuses more on the control and operation of low-level instructions. Low P-Code pays more attention to the control flow and data flow of the program and is closer to actual machine instructions or assembly instructions in terms of form, rather than the structures in high-level languages. Low P-Code is mainly used to more accurately restore the underlying behavior of the program, especially when the program is based on machine code or assembly code, and Low P-Code provides a more direct mapping of low-level instructions.

[0112] In practical applications, binary reverse analysis tools such as Ghidra can be used to convert binary machine code (machine instruction set) into P-Code (P-Code instruction set). Each P-Code instruction in the P-Code instruction set corresponds to a P-Code operation. A P-code operation is the basic unit in P-Code, and each P-code operation represents an atomic operation. There may be 63 kinds of atomic operations involved in P-Code, covering basic arithmetic operations, logical operations, comparison operations, memory operations, control flow operations, and architecture-specific operations.

[0113] Specifically, in the P-Code, the execution logic of the program is represented by various types of nodes. These nodes represent the basic operations executed in the program. Each P-Code node usually corresponds to an operation instruction (P-Code instruction) or a higher-level logical operation. Several common node types in P-Code are: Assignment Nodes, Arithmetic Operation Nodes, Logical Operation Nodes, Comparison Nodes, Control Flow Nodes, Function Call Nodes, Memory Access Nodes, Constant Load Nodes, stack Operation Nodes, Branching Nodes, Exception Handling Nodes, etc. Among them, an Assignment Node represents a value being assigned to a variable or a memory location. An Arithmetic Operation Node represents basic arithmetic operations such as addition, subtraction, multiplication, division, etc. A Logical Operation Node represents logical operations such as AND, OR, XOR, NOT, etc. A Comparison Node is used to perform comparison operations and usually generates a boolean result. A Control Flow Node represents control flow instructions in the program such as jumps, conditional jumps, etc. A Function Call Node represents function calls in the program, and they usually contain jumps to other functions or subroutines. A Memory Access Node represents read or write operations to memory locations, and they are usually used to access data in arrays, structure fields, stacks, or heaps. A Constant Load Node represents loading a constant value into a register or a memory location. A Branching Node represents a conditional jump in the program, usually based on the result of a previous comparison or calculation. An Exception Handling Node indicates handling exceptions or specific control flows in the program, such as handling specific errors or signals.

[0114] In addition, in the P-Code, operands and registers are abstracted as variable nodes, which are divided into: register, memory, stack, const, and unique according to the storage location and function.

[0115] (12) Basic Block: In the embodiments of the present application, a basic block is a sequence of statements executed sequentially in the program, which has only one entry and one exit. The entry is its first statement, and the exit is its last statement.

[0116] A basic block can be regarded as the smallest and independent processing unit in a program. Inside this unit, the instructions are executed linearly and there are no jumps (except for the jump at the end of the basic block). The start of a basic block is usually the target to which the control flow instruction jumps, and the end is the point where the control flow changes. Among them, control flow instructions refer to the instructions that can change the execution order of the program, such as branch instructions, jump instructions, return instructions, call instructions, loop instructions, etc. Control flow instructions must specify the target address, that is, where to jump to continue the execution.

[0117] Each function in a program can be composed of one or more basic blocks. The start of a function marks the start of the first basic block, and the control flow instructions (such as conditional branches, loops, etc.) inside the function will cause the end of the basic block and start a new basic block. Therefore, a function can be regarded as a set of basic blocks, and these basic blocks are connected to each other through control flow instructions.

[0118] (13) Low Level Virtual Machine (LLVM): In the embodiments of this application, LLVM is a powerful and flexible compiler framework, which is widely used in fields such as compiler development, program optimization, static analysis, and language implementation. It provides cross-platform compilation support, has a modular design and powerful optimization capabilities, and can effectively generate efficient target code. Through LLVM, developers can focus more on the design and implementation of programming languages without having to develop compilation tools for each hardware platform separately.

[0119] LLVM uses a low-level abstract language called "Intermediate Representation (IR)" as the core of the compilation process. LLVM IR is an intermediate representation independent of the specific hardware platform. It converts the P-Code from high-level languages into a more standardized low-level representation, and can obtain a unified representation independent of the platform and language, which is convenient for analysis and optimization.

[0120] (14) Control Flow Graph (CFG): In the embodiments of this application, the control flow graph is an abstract representation of the program, representing all the paths that will be traversed during the program execution. The control flow graph represents the possible flow directions of all basic blocks in the program in the form of a graph, and also reflects the real-time execution process of the program. The control flow graph consists of nodes and edges. The nodes represent basic blocks, and the edges represent the paths of the control flow, that is, the paths from one basic block to another basic block.

[0121] (15) Three-Address Code (TAC): In the embodiments of this application, the three-address code is an intermediate language. Each three-address code instruction can be decomposed into four tuples: operator, operand1, operand2, result, where operand1, operand2, and result are variables. Since each three-address code instruction contains three variables, it is called a three-address code. It should be noted that each instruction in the three-address code has only one operation, and the processing process is simple.

[0122] Binary reverse analysis is a technique for understanding the behavior and logic of software by analyzing its binary files (machine code). Binary reverse analysis is widely used in various fields such as software security, malware analysis, vulnerability detection, and legacy system maintenance.

[0123] In binary reverse analysis work, function name recovery is an important task. Function names usually reflect the functions and uses of functions. When function names are recovered, on the one hand, function names can help understand the behavior and logical structure of software programs; on the other hand, function names can also help quickly locate specific functional modules or code segments, thereby improving the efficiency of reverse analysis. In security analysis and vulnerability research, accurate function names can help quickly locate potential security risk points and help identify critical dangerous functions.

[0124] In the related art, it is usually necessary to compile the source code written in a high-level language into a binary file (machine code) that can be understood by a machine. In practical applications, during the process of compiling the source code of a software program (such as a closed-source software) into a binary file, the symbol information in the binary file is usually removed or obfuscated. In this way, on the one hand, unnecessary data can be removed, improving the loading speed and execution efficiency of the program; on the other hand, the security of the source code can also be protected. As an example, during the process of compiling the source code into a binary file, the Strip tool can be used to remove the symbol information in the source code. Specifically, one or more of the static symbol table, dynamic symbol table, and debug information can be removed.

[0125] The following combines Figures 1 to 3 , taking the removal of the static symbol table as an example, to elaborate on the differences between the binary files obtained by removing symbol information and not removing symbol information. Among them, Figure 1 is a schematic diagram of the source code provided by the embodiments of this application. Figure 2 is a schematic diagram of the assembly code corresponding to the binary file without removing the static symbol table provided by the embodiments of this application. Figure 3Schematic diagram of the assembly code corresponding to the binary file with the static symbol table removed provided by the embodiments of the present application. It can be understood that since the binary machine language presented by the binary file is not very readable, the example description is carried out through the assembly code corresponding to the binary file.

[0126] It should be noted that in the embodiments of the present application, bold font is used to display important content parts. It can be understood that the embodiments of the present application do not make specific limitations on how to display important content parts. For example, in some application scenarios, underlining can be used to distinguish and display important content parts. Specifically, wavy underlining can be used to display important content parts. In other application scenarios, colors can also be used to distinguish and display important content parts. Specifically, bright colors can be used to display important content parts.

[0127] As Figure 1 shown, the source code is the definition code of the CRYPTO_control function. In the CRYPTO_control function, the CRYPTO_encrypt function and the CRYPTO_decrypt function are called.

[0128] It should be noted that the static symbol table is used to record the symbols defined and referenced in the program (such as global variable names, function names). For example, in combination with Figure 1 , the following symbols can be recorded in the static symbol table: CRYPTO_control, CRYPTO_encrypt, and CRYPTO_decrypt.

[0129] As Figure 2 shown, when symbol information is retained in the binary file, after disassembling the binary file using a disassembler, the function names of each function, that is, CRYPTO_control, CRYPTO_encrypt, and CRYPTO_decrypt, are presented completely, and the assembly code clearly reflects the logic of the source code.

[0130] As Figure 3 shown, when symbol information is not retained in the binary file, after disassembling the binary file using a disassembler, the function names of each function are replaced with symbols without semantic information. Specifically, the function name of the CRYPTO_control function is replaced with FUN_00101209, the function name of the CRYPTO_encrypt function is replaced with FUN_00101270, and the function name of the CRYPTO_decrypt function is replaced with FUN_00101359. Combining with Figure 3 the assembly code shown, it is very difficult for users to quickly understand the logic of the source code from it, and it is also impossible to filter and directly locate some key dangerous functions based on the function names.

[0131] From Figures 1 to 3 It can be found that when the symbol information is removed from the compiled binary file, decompiling the binary file cannot restore the function names in the source code in the obtained assembly code, and it is impossible to filter and directly locate the key dangerous functions according to the function names.

[0132] In order to restore the function names in the source code from the binary file, an optional approach is: the electronic device directly calls binary reverse analysis tools, such as IDA-Pro, Ghidra, etc., and the binary reverse analysis tools analyze to obtain the function names in the source code corresponding to the binary file. However, the accuracy of the function names generated directly using binary reverse analysis tools is relatively low. Taking the binary reverse analysis tool Ghidra as an example, if function 1 is defined in a certain.so file, and the.so file records the function name and function address of function 1, and the function address of function 1 in the.so file is address 1. If the.o file calls function 1, and the function address of function 1 in the.o file is address 2. In this case, since address 1 and address 2 are often not the same, when the function name of function 1 in the.so file is removed, the.o file will not be able to call function 1, Ghidra cannot trace the relevant information of function 1, and thus cannot restore the function name of function 1, resulting in relatively low accuracy of function name restoration.

[0133] To solve the problem of restoring the function names of the source code from the binary file with removed symbol information, in this application, first convert the binary file into P-Code code; then extract the structural information of the function from the P-Code code, and then combine the structural information of the function to generate an intermediate file for recording the structure of the function in the P-Code code; after that, based on the intermediate file, convert the P-Code instructions of various operation types related to the function structure into LLVM IR instructions, so as to obtain an LLVM IR file; finally, generate a control flow graph of the function based on the LLVM IR file, and generate the function name of the function based on the control flow graph.

[0134] Optionally, the P-Code instructions of various operation types related to the function structure may involve operation types including but not limited to: obtaining input variables, updating output variables, branch merging, and conventional types. The conventional type is an operation type other than obtaining input variables, updating output variables, and branch merging. Among them, each P-Code instruction corresponds to a P-Code operation, and each P-Code operation corresponds to an operation type. It should be noted that the P-code operation is the basic unit in the P-Code code, and each P-code operation represents an atomic operation. There may be 63 kinds of atomic operations involved in the P-Code code, covering basic arithmetic operations, logical operations, comparison operations, memory operations, control flow operations, and architecture-specific operations.

[0135] As an example, the instruction corresponding to obtaining an input variable can be a load instruction, the instruction corresponding to updating an output variable can be a store instruction, and the instruction corresponding to branch merging can be a MULTIEQUAL instruction.

[0136] It should be noted that converting the P-Code instructions of various operation types described in the intermediate file into corresponding LLVM IR instructions can ensure that each P-Code instruction involved in the function is converted into an LLVM IR instruction, that is, it can achieve a comprehensive and accurate conversion of the instructions involved in the function.

[0137] Optionally, the P-Code code obtained from the binary file may include high-level pseudocode and low-level pseudocode. The electronic device can analyze using the high-level pseudocode. Specifically, the structural information of the function can be extracted from the high-level pseudocode. It should be noted that since the high-level pseudocode pays more attention to the source code itself and does not pay attention to the underlying architecture, while the low-level pseudocode contains a large amount of information related to the underlying architecture, and this information related to the underlying architecture has nothing to do with the function itself. Therefore, using the high-level pseudocode for function name recovery can avoid excessive irrelevant interference information, help reduce the amount of data processing, and improve the efficiency of function name recovery.

[0138] Combined with Figure 4 , Figure 4 is a schematic diagram of the effect of function name recovery provided by the embodiments of the present application. Figure 4 In (a) shown in Figure 4 is the assembly code corresponding to the target binary file, and

[0139] In Figure 4As shown in (a) of [reference], in the assembly code corresponding to the target binary file, the function name is unknown, or rather, the function name is replaced by a symbol with an unclear meaning.

[0140] Such as Figure 4 As shown in (b) of [reference], after restoring the function name of the target binary file, the function name in the obtained assembly code is restored. In this case, the user can accurately determine whether there are critical dangerous functions in the source code and locate the positions of the critical dangerous functions by combining the assembly code with accurate function names.

[0141] Combined with Figure 4 , it can be found that the function name restoration method provided by this application can accurately restore the function names of each function in the software program, which helps to discover the dangers existing in the software through the function names, thereby ensuring software security.

[0142] The embodiments of this application have at least the following beneficial effects:

[0143] 1. The structural information of the function can accurately describe the characteristics of the function. Accurately extracting rich and comprehensive structural information of each function, such as function information, basic block information in the function, and variable information in the function, and generating the function name of the function based on the structural information of the function helps to improve the accuracy rate of function name restoration.

[0144] 2. Since LLVM IR is an intermediate representation independent of the specific hardware platform, it converts P-Code from high-level languages into a more standardized low-level representation, and a unified representation independent of the platform and language can be obtained, which is convenient for analysis and optimization. First, converting the intermediate file into an LLVM IR file can utilize the powerful functions of LLVM for cross-platform optimization and analysis, and simplify the generation process of the control flow graph, making the analysis and optimization more efficient and general.

[0145] 3. In the case of accurately restoring the function name, the critical dangerous functions can be filtered and directly located according to the function name, which helps to ensure software security.

[0146] 4. Since only the P-Code code part related to the function structure is recorded in the intermediate file, generating the LLVM IR file corresponding to the P-Code code based on the intermediate file can avoid converting the P-Code code part unrelated to the function structure into LLVM IR instructions. That is to say, this application can avoid converting redundant instructions in the P-Code code, which helps to improve the efficiency of function name restoration.

[0147] 5. Converting the P-Code instructions of various operation types described in the intermediate file into corresponding LLVM IR instructions can ensure that each P-Code instruction involved in the function is converted into an LLVM IR instruction, that is, it can achieve a comprehensive and accurate conversion of the instructions involved in the function. In this way, it helps to comprehensively and accurately extract the characteristic information of the function, thereby improving the accuracy of function name recovery.

[0148] 6. Since high-level pseudocode pays more attention to the source code itself and does not focus on the underlying architecture, while low-level pseudocode contains a large amount of information related to the underlying architecture, and this information related to the underlying architecture has nothing to do with the function itself. Therefore, using high-level pseudocode for function name recovery can avoid excessive irrelevant interference information, help reduce the amount of data processing, and improve the efficiency of function name recovery.

[0149] The following describes the usage scenarios of the embodiments of the present application:

[0150] The embodiments of the present application can be applicable to the scenario of recovering function names in software programs.

[0151] The function name recovery method provided by the embodiments of the present application can be applied to an electronic device, and the electronic device can be a device such as a terminal or a server. The terminal can be a tablet computer, a vehicle-mounted device, an augmented reality / virtual reality device, a laptop computer, a super mobile personal computer, etc. The server can be a network server, a cloud server, etc. The embodiments of the present application do not make any limitations in this regard.

[0152] Figure 5 FIG. is a schematic flowchart of a function name recovery method provided by an embodiment of the present application. As Figure 5 shown, the function name recovery method may include the following steps 501 to step 507. It can be understood that the above function name recovery method may include all the steps from step 501 to step 507, or may only include some of the steps. It can be understood that the steps in the function name recovery method can be arbitrarily combined without conflict.

[0153] Step 501, the electronic device generates target pseudocode for the target binary file based on the target binary file.

[0154] Among them, the target pseudocode is the pseudocode corresponding to the target binary file. Pseudocode can also be called P-Code code. P-Code code is a set of P-Code instructions, simply referred to as the P-Code instruction set.

[0155] Optionally, the target binary file is a binary file from which symbol information has been removed. In the present application, a binary file refers to a file that carries the machine code of the program.

[0156] In some scenarios, a binary file can also be referred to as a machine code file or simply as machine code. Machine code is a collection of machine instructions, abbreviated as a machine instruction set. The file format of a binary file can include: object file format (.o file), shared object file format (.so file), executable file format (.exe file), and so on. It can be understood that the embodiments of the present application do not specifically limit the file format of the target binary file.

[0157] P-Code can include High P-Code and Low P-Code. High P-Code and Low P-Code essentially correspond to the logic of the same program.

[0158] High P-Code is closer to the source code and usually retains more details related to the behavior and semantics of the program. That is to say, High P-Code focuses more on the source code itself rather than on the underlying architecture. In contrast, Low-P-Code focuses more on the control and operation of the underlying instructions. That is to say, Low-P-Code contains a lot of information about the underlying architecture.

[0159] Considering that a large amount of information related to the underlying architecture contained in Low-P-Code has nothing to do with the function to be restored in the present application, in order to avoid excessive irrelevant interference information, High P-Code can be used in the function name restoration process in the present application. That is to say, the target pseudocode is usually High P-Code.

[0160] In step 501, the electronic device can call a binary reverse analysis tool, such as Ghidra, to first disassemble the target binary file (i.e., the machine instruction set) into assembly code, and then decompile the assembly code into P-Code. It can be understood that in the case of using High P-Code to restore the function name, only High P-Code needs to be generated in step 501, without the need to generate Low P-Code, which can save unnecessary time consumption and computing resources.

[0161] In step 502, the electronic device parses the target pseudocode and extracts the structural information of each function in the target pseudocode.

[0162] Here, similar to high-level languages (such as the C++ language), in the target pseudo-code (P-Code), it can include various functions corresponding to the source code. Each function contains one or more basic blocks, and each basic block contains a series of P-Code instructions. Each P-Code instruction corresponds to a P-Code operation. The P-code operation is the basic unit in the P-Code, and each P-code operation represents an atomic operation. There may be 63 kinds of atomic operations involved in the P-Code, covering basic arithmetic operations, logical operations, comparison operations, memory operations, control flow operations, and architecture-specific operations.

[0163] Among them, the structural information of the function can include function information, basic block information in the function, and variable information in the function. The function information is used to describe the function, the basic block information is used to describe the basic blocks in the function, and the variable information is used to describe the variables in the function.

[0164] Among them, the function information can include but is not limited to: the name of the function (abbreviated as the function name), the starting address of the function, the type and size of the return value of the function. The starting address of the function is the entry address of the function. The entry address of the function is the storage address of the first P-Code instruction in the function.

[0165] The basic block information can include but is not limited to: the forward node information of the basic block, the backward node information of the basic block, the index of the basic block, and the starting address of the basic block. The starting address of the basic block is also called the entry address of the basic block. The starting address of the basic block is the storage address of the first P-Code instruction in the basic block. Among them, the index of the basic block is the unique identifier of the basic block, and each basic block has an index. For example, the index of the first basic block in the function can be 0, and the index of the next basic block can be 1, and so on. Since the basic block is a sequence of statements executed sequentially in the program, with only one entry and one exit, the entry is its first statement, and the exit is its last statement. Therefore, for any basic block, the forward node of this basic block is the P-Code instruction that jumps to this basic block, and the backward node of this basic block is the P-Code instruction that this basic block needs to jump to.

[0166] For example, if there are 3 basic blocks in function 1, namely basic block 0 to basic block 2, and if basic block 0 can jump to basic block 1, and basic block 1 can jump to basic block 2, then for basic block 0, there is no forward node, and the backward node is the first P-Code instruction in basic block 1. For basic block 1, the forward node is the last P-Code instruction in basic block 0, and the backward node is the first P-Code instruction in basic block 2. For basic block 2, there is no backward node, and the forward node is the last P-Code instruction in basic block 1.

[0167] It is understandable that a basic block may have no forward nodes, or may have one or more forward nodes. A basic block may have no backward nodes or may have one or more backward nodes.

[0168] Variable information may include but is not limited to the storage type of the variable, the variable size, and the variable name. Optionally, the storage types of the variables involved in the variable information may include: register type, memory type, stack type.

[0169] In the embodiments of the present application, since the P-Code has a specific syntax, the electronic device can extract function information, basic block information in the function, and variable information in the function from the P-Code in combination with the syntax of the P-Code.

[0170] It should be noted that since the P-Code is essentially unordered code, extracting the structural information of each function from the P-Code can achieve an ordered classification of the unordered P-Code. In addition, classifying and dividing the P-Code layer by layer according to multiple levels such as functions, basic blocks in the functions, instructions in the basic blocks, and variables helps to accurately extract the semantic structure of each function, thereby ensuring the accuracy of function name recovery.

[0171] The following combines Figure 6 to illustrate how to extract the structural information of functions from the P-Code. Among them, Figure 6 is a schematic diagram of the process for extracting the structural information of functions provided by the embodiments of the present application.

[0172] Step 601, the electronic device traverses the P-Code, and when accessing the first operation instruction, extracts the function information of the first function from the first operation instruction.

[0173] Among them, the first operation instruction is an instruction for defining a function in the P-Code (which can be simply referred to as a definition instruction). The first function is the function defined by the first operation instruction. The first function (or the target function) can be any function in the P-Code. It is understandable that the P-Code is composed of instructions, and the instructions in the P-Code can also be called pseudo-code instructions or P-Code instructions. Each instruction in the P-Code corresponds to an operation, such as an AND operation, an OR operation, etc., so the instructions in the P-Code can also be called operation instructions.

[0174] In practice, in P-Code, a defined symbol can be used to define a function. The defined symbol is a pre-set symbol for defining a function. The defined symbol usually includes a start symbol and an end symbol. The start symbol indicates the start of defining a function, and the end symbol indicates the end of defining the function. For example, "BEGIN FUNC" can be used to indicate the start of defining a certain function, and "END FUNC" can be used to indicate the end of defining the function. Another example is that "DEFINE" can be used to indicate the start of defining a certain function, and "END" can be used to indicate the end of defining the function. It can be understood that for different binary reverse analysis tools, the defined symbols may be different, and the embodiments of the present application do not limit the specific implementation of the defined symbols.

[0175] Here, the function information of the first function may include, but is not limited to: the name of the first function (abbreviated as the function name), the starting address of the first function, the type of the return value of the first function, and the size of the return value of the first function. The starting address of the first function can also be referred to as the entry address of the first function.

[0176] Here, the electronic device can determine whether the currently accessed instruction is the first operation instruction based on the defined symbol. If it is the first operation instruction, the first function is identified from the first operation instruction and the function information of the first function is obtained.

[0177] Step 602, the electronic device determines basic blocks and the basic block information of the basic blocks from the first function according to each P-Code instruction in the first function.

[0178] Here, each basic block in the first function may include one or more P-Code instructions.

[0179] The basic block information may include, but is not limited to: the forward node information of the basic block, the backward node information of the basic block, the index of the basic block, and the starting address of the basic block. Among them, the index of the basic block is the unique identifier of the basic block, and each basic block has an index. For example, the index of the first basic block in the function can be 0, and the index of the next basic block can be 1, and so on. For any basic block, the forward node of the basic block is the P-Code instruction that jumps to the basic block, and the backward node of the basic block is the P-Code instruction that the basic block needs to jump to.

[0180] It should be noted that in the P-Code, the starting address of the first basic block within a function is the entry address of the first P-Code instruction in that function. The last P-Code instruction of a basic block can be a jump instruction or not. Among them, a jump instruction usually includes a jump symbol and the entry address of the jump target, and the jump target is the address of the first P-Code instruction of the next basic block. It can be understood that the entry address of a basic block is also the address of the first P-Code instruction in the basic block. A jump instruction is used to jump from the current position to another position for execution. Among them, the jump symbol is a symbol indicating the jump operation. As an example, the jump symbol can be "BRANCH". It should be understood that if the last P-Code instruction of a basic block is not a jump instruction, it means that this function has only one basic block, that is, there is no need to jump to other basic blocks within the function.

[0181] Here, the electronic device can identify each basic block in the first function based on the jump symbols in each P-Code instruction traversed in the first function, and obtain the basic block information of the basic block.

[0182] Step 603, the electronic device traverses each P-Code instruction in the first basic block and extracts the variable information of each variable in the P-Code instruction.

[0183] Among them, the first basic block is any basic block in the first function.

[0184] Among them, the variable information can include but is not limited to storage type, variable size, and variable name.

[0185] Optionally, the storage types of the variables involved in the variable information can include: register type, memory type, stack type. It should be noted that there are the following 5 types of storage types for variables in P-Code: register type, memory type, stack type, const type, unique type. Among them, since the values of const type variables (referred to as const variables for short) and unique type variables (referred to as unique variables for short) do not change, only extracting the variable information of register type variables (referred to as register variables for short), memory type variables (referred to as memory variables for short), and stack type variables (referred to as stack variables for short) can reduce the amount of data to be extracted, thereby improving the function name recovery efficiency.

[0186] Here, for each basic block, the electronic device can traverse each P-Code instruction in the basic block and obtain the variable information from the P-Code instruction.

[0187] FromFigure 6 It can be found that by traversing the P-Code, the electronic device can extract the structural information of each function from the P-Code.

[0188] It should be noted that since the structural information of a function can accurately describe the semantic features of the function, accurately extracting the structural information of each function helps to improve the accuracy of function name recovery.

[0189] Step 503, the electronic device generates an intermediate file for recording the structural information of each function in the target pseudo-code based on the structural information of each function in the target pseudo-code.

[0190] Among them, the intermediate file includes the structural information of each function in the target pseudo-code (P-Code). The file format of the intermediate file can be Extensible Markup Language (XML) format or JavaScript Object Notation (JSON) format. It can be understood that the embodiments of the present application do not specifically limit the file format of the intermediate file.

[0191] Here, the electronic device can write the structural information of each function in the P-Code into the intermediate file. In this case, the intermediate file can intuitively reflect the structure of the P-Code.

[0192] The following further combines Figure 7 to elaborate on the process from the binary file to the intermediate file. Among them, Figure 7 is a schematic diagram of the process from the binary file to the intermediate file provided by the embodiments of the present application.

[0193] As Figure 7 shown, the intermediate file corresponding to the binary file can be obtained through the following steps 701 to 710.

[0194] Step 701, the electronic device obtains the binary file.

[0195] Optionally, the binary file obtained by the electronic device can be a binary file from which symbol information has been removed.

[0196] Step 702, the electronic device converts the binary file into P-Code.

[0197] Here, for the operation of converting the binary file into P-Code, reference can be made to the corresponding embodiment part of step 501, which will not be elaborated here.

[0198] Step 703, the electronic device traverses the functions in the P-Code.

[0199] Here, the electronic device can refer to Step 601 to access the functions in the P-Code.

[0200] Step 704, when the first function is accessed, the electronic device extracts the function name, start address, return value type, and return value size of the first function. When the first function is not accessed, the structural information of each function that has been extracted is saved to an intermediate file.

[0201] Among them, for any function, the structural information of the function can include function information, basic block information in the function, and variable information in the function.

[0202] Among them, the first function is the currently accessed function. The first function can be any function in the P-Code.

[0203] Here, if the first function is not accessed, it means that the functions in the P-Code have been traversed.

[0204] The function name, start address, return value type, and return value size of the function can refer to the function information in Step 601. The electronic device can refer to the corresponding embodiment part of Step 601 to extract the function information of the function, which will not be elaborated here.

[0205] Step 705, the electronic device traverses the basic blocks in the first function.

[0206] Here, the electronic device can refer to the corresponding embodiment part of Step 602 to access each basic block in the first function.

[0207] Step 706, when the first basic block is accessed, the electronic device extracts the index, start address, forward node, and backward node of the first basic block. When the first basic block is not accessed, the electronic device returns to Step 703 to traverse the next function in the P-Code.

[0208] Among them, the first basic block is the currently accessed basic block. The first basic block can be any basic block in the first function.

[0209] Here, if the first basic block is not accessed, it means that the basic blocks in the first function have been traversed.

[0210] The index, start address, forward node, and backward node of the first basic block can refer to the basic block information in Step 602. The electronic device can refer to the corresponding embodiment part of Step 602 to extract the basic block information of each basic block in the function, which will not be elaborated here.

[0211] Step 707, the electronic device traverses each P-Code instruction in the first basic block.

[0212] Here, the electronic device can refer to the embodiment part corresponding to step 603 to access each P-Code instruction in the first basic block.

[0213] Step 708, when the P-Code instruction is accessed, the electronic device extracts the operation name, output variables, and input variables of the P-Code instruction. When the P-Code instruction is not accessed, the electronic device returns to step 705 to continue traversing the next basic block in the first function.

[0214] Here, since the P-Code instruction is not accessed, it means that the P-Code instructions in the first basic block have been traversed. The electronic device can directly extract the operation name, output variables, and input variables of the P-Code instruction from the currently accessed P-Code instruction.

[0215] Step 709, the electronic device traverses each variable in the P-Code instruction.

[0216] Step 710, when a variable is accessed, the storage type, size, and name of the variable are extracted, and the variables with storage types of stack, register, and memory are recorded. When no variable is accessed, the electronic device returns to step 707 to continue traversing the next P-Code instruction in the first basic block.

[0217] Here, since no variable is accessed, it means that all variables in the currently accessed P-Code instruction have been traversed. The storage type, size, and name of the variable can be referred to the variable information in step 603. The electronic device can refer to the embodiment part corresponding to step 603 to extract the variable information of each variable in the function, which will not be elaborated here.

[0218] The following combines Figures 8 to 11 , and elaborates the structure of the intermediate file. Among them, the intermediate file can include a variable part and a function part. Figure 8 This is a schematic diagram of the variable part in the intermediate file provided by the embodiment of the present application. Figure 9 This is a schematic diagram of a certain function in the function part of the intermediate file provided by the embodiment of the present application. Figure 10 This is a schematic diagram of the visual structure of the intermediate file provided by the embodiment of the present application. Figure 11 This is a schematic diagram of the visual structure of a certain function provided by the embodiment of the present application.

[0219] Such as Figure 8As shown, lines 2 to 16 record multiple register variables in the P-Code, and the register variables are global variables. Lines 17 to 23 record the memory variables in the P-Code. Lines 24 to 26 record the stack variables in the P-Code. Taking line 3 as an example, line 3 describes a register variable whose name is "EAX" and size is 4 bytes. It should be understood that Figure 8 the line numbers of each line are only used for distinguishing descriptions and not for specific limitations on each line.

[0220] Combined with Figure 8 it can be found that the intermediate file clearly records the information of each variable in the P-Code. The intermediate file can record register variables, memory variables, and stack variables in the P-Code, such as the name and size of each variable.

[0221] Such as Figure 9 shown, lines 1 to 39 are the description content of a function. Among them, line 1 describes the entry address and function name of the function. Line 2 describes the output information of the function, and line 3 describes the input information of the function. Lines 4 to 38 describe the basic blocks in the function.

[0222] Lines 5 to 37 describe the first basic block in the function, that is, block_0. Among them, line 6 describes the entry address of block_0, line 7 describes the index of block_0, line 8 describes the forward node of block_0, and line 9 describes the backward node of block_0.

[0223] Lines 10 to 36 describe the P-Code instructions in block_0. Among them, lines 11 to 18 describe the first P-Code instruction, that is, pcode_0, lines 19 to 26 describe the second P-Code instruction, that is, pcode_1, and lines 27 to 35 describe the third P-Code instruction, that is, pcode_2.

[0224] Taking pcode_0 as an example, line 12 describes the information of the output variable of pcode_0. Line 13 describes the operation name of pcode_0. Lines 14 to 17 describe the information of the input variables of pcode_0, and among them, lines 15 to 16 describe the information of the first input variable in pcode_0.

[0225] Combined with Figure 9 , Figure 9The starting address (i.e., the entry address of the function) of the shown function (abbreviated as the diagrammatic function) is 00101000, and the function name is "init" (see Figure 9 in line 1). The return value type of the diagrammatic function is "int" (see Figure 9 in line 2). The name of the input parameter of the diagrammatic function is "ctx", and the type is "EVP_PKEY_CTX" (see Figure 9 in line 3). The diagrammatic function has only one basic block, that is, block_0 (see Figure 9 from line 4 to line 38).

[0226] The starting address (i.e., the entry address) of block_0 (the first basic block) of the diagrammatic function is "00101004", the index is 0, there are no forward nodes, and there are no backward nodes (see Figure 9 in lines 6 to 9).

[0227] There are 3 P-Code instructions in block_0, namely pcode_0 (see Figure 9 in lines 11 to 18), pcode_1 (see Figure 9 in lines 19 to 26), and pcode_2 (see Figure 9 in lines 27 to 35).

[0228] The output variable of pcode_0 is EAX, the storage type is register type, and the size is 4 bytes (see Figure 9 in line 12). The operation type of pcode_0 is CALL operation (see Figure 9 in line 13). The input variable of pcode_0 is "A_00105050:8", the storage type is memory type, and the size is 8 bytes (see Figure 9 in line 15).

[0229] The output variable of pcode_1 is EAX, the storage type is register type, and the size is 4 bytes (see Figure 9 in line 20). The operation type of pcode_1 is COPY operation (see Figure 9 in line 21). The input variable of pcode_1 is EAX, the storage type is register type, and the size is 4 bytes (see Figure 9 in line 23).

[0230] The operation type of pcode_2 is RETURN operation (see Figure 9 in line 28). The first input variable of pcode_2 is "0x0", the storage type is constant type, and the size is 8 bytes (seeFigure 9 in line 30). The second input variable of pcode_2 is EAX, with a storage type of register and a size of 4 bytes (see Figure 9 in line 32).

[0231] It can be understood that Figure 9 the function shown is only an example and does not specifically limit the function.

[0232] From Figure 9 it can be found that the intermediate file clearly records the structural information of each function in the P-Code. It should be noted that the electronic device can access the P-Code instructions in the function based on the entry address of the function recorded in the intermediate file (which can also be called the start address or the first entry address of the function), and the electronic device can access the P-Code instructions in the basic block based on the entry address of the basic block recorded in the intermediate file (which can also be called the start address or the second entry address of the basic block). In this way, based on the intermediate file, only the P-Code instructions related to the function can be converted into instructions in the form of three-address code, such as being converted into LLVM IR instructions. Since converting the P-Code instructions into instructions in the form of three-address code can facilitate the extraction of the control flow information of the function, and the control flow information of the function can accurately describe the semantics of the function, that is to say, converting the P-Code instructions related to the function into instructions in the form of three-address code can achieve the accurate and effective extraction of the semantic features of the function.

[0233] As Figure 10 shown, the intermediate file can include 4 information blocks, namely the global information block (see Figure 10 lines 3 and 4), the memory variable information block (see Figure 10 lines 5 and 6), the stack variable information block (see Figure 10 line 7), and the function information block (see Figure 10 lines 8 to 17). Among them, 13 register variables are recorded in the global information block. 32 memory variables are recorded in the memory variable information block. 38 functions are recorded in the function information block.

[0234] Further referring to Figure 11 , Figure 11 is a schematic structural diagram of the function numbered 17 in the intermediate file (abbreviated as function 17). As Figure 11 shown, the structure of function 17 can include 5 parts, namely: the output part (see Figure 11 lines 2 and 3), the input part (see Figure 11 line 4), the basic block part (see Figure 11lines 5 to 11), the entry address part (see Figure 11 line 12), the function name part (see Figure 11 line 13).

[0235] Figure 11 In, the output part describes that function 17 has an output and the output type is void.

[0236] The input part describes that function 17 has 5 inputs.

[0237] The basic block part describes that function 17 has a basic block, that is, block_0, and block_0 has 5 items, namely: entry address, index, forward node, backward node, P-Code instruction (see Figure 11 lines 6 to 11). Specifically, the entry address in block_0 is "00101350", the index is 0, there is no forward node, and there is no backward node. There are 7 P-Code instructions in block_0.

[0238] The entry address part describes that the entry address of function 17 is: 00101359.

[0239] The function name part describes that the function name of function 17 is: CRYPTO_decrypt.

[0240] From Figures 8 to 11 it can be found that the intermediate file clearly describes the semantic structure of each function in the P-Code code.

[0241] Step 504, the electronic device generates an LLVM IR file corresponding to the target pseudocode based on the intermediate file.

[0242] Here, the electronic device can traverse the intermediate file. During the traversal of the intermediate file, when accessing the entry address, such as the entry address of a function or the entry address of a basic block in a function, it can access each P-Code instruction in the function based on the entry address and convert the P-Code instruction into an LLVM IR instruction. In this way, the LLVM IR file corresponding to the entire P-Code code can be obtained.

[0243] It should be noted that since only the P-Code code part related to the function structure is recorded in the intermediate file, generating the LLVM IR file corresponding to the P-Code code based on the intermediate file can avoid converting the P-Code code part that is not related to the function structure into LLVM IR instructions. That is to say, this application can avoid converting redundant instructions in the P-Code code, which helps to improve the function name recovery efficiency.

[0244] The following combines with Figure 12 to elaborate on how to generate an LLVM IR file corresponding to the P-Code code based on the intermediate file. Among them, Figure 12 is a schematic flowchart of the instruction conversion provided by an embodiment of the present application. As Figure 12 shown, the process of instruction conversion may at least include the following steps 5041 to step 5042.

[0245] Step 5041, the electronic device traverses each function in the intermediate file, and based on the entry address of the currently accessed function and the entry addresses of each basic block in the function, accesses each P-Code instruction in the current function.

[0246] Among them, the current function is the currently accessed function.

[0247] Here, the code of the function, that is, the function body, is stored in the memory. The electronic device can access the function from the memory through the entry address of the function. Similarly, the electronic device can access each basic block in the function based on the entry addresses of each basic block in the function, or access each P-Code instruction in the basic block.

[0248] Step 5042, the electronic device determines the operation type corresponding to the first P-Code instruction in the current function, and converts the first P-Code instruction into an LLVM IR instruction based on the operation type corresponding to the first P-Code instruction, so as to obtain an LLVM IR file corresponding to the intermediate file.

[0249] Among them, the first P-Code instruction is the currently accessed P-Code instruction.

[0250] Among them, the operation type may include but is not limited to: obtaining input variables, updating output variables, branch merging, and conventional types. The conventional type is an operation type other than obtaining input variables, updating output variables, and branch merging. Among them, each P-Code instruction corresponds to a P-Code operation, and each P-Code operation corresponds to an operation type. It should be noted that the P-code operation is the basic unit in the P-Code code, and each P-code operation represents an atomic operation. There may be 63 atomic operations involved in the P-Code code, covering basic arithmetic operations, logical operations, comparison operations, memory operations, control flow operations, and architecture-specific operations.

[0251] The following combines with the operation type of the first P-Code instruction to describe how to convert the first P-Code instruction into an LLVM IR instruction. It can be understood that the electronic device can determine the operation type of the P-Code instruction through the operator of the P-Code instruction.

[0252] Case 1, the operation type of the first P-Code instruction is a conventional type.

[0253] In the case of Case 1, the electronic device can convert the first P-Code instruction into an LLVM IR instruction based on a pre-set operation mapping relationship (or referred to as the first mapping relationship). Among them, the operation mapping relationship is used to describe the conversion relationship between the P-Code instruction and the LLVM IR instruction.

[0254] The following describes how to convert the P-Code instruction into an LLVM IR instruction in conjunction with Table 1. Among them, Table 1 is part of the operation mapping relationship provided by the embodiments of the present application.

[0255] Table 1:

[0256]

[0257] In Table 1, the first column is the operation name of the P-Code operation corresponding to the P-Code instruction, the second column is the data processed by the P-Code operation, and the third column is the LLVM IR instruction corresponding to the P-Code instruction.

[0258] The following describes the content of each row in Table 1 in order from top to bottom.

[0259] In the first row, BOOL_NEGATE represents an operation of negating a variable, and the return value type is boolean.!v0 represents an operation of negating variable v0. In the instruction “%negated = xor i1 %v0, true”, %v0 is a variable of i1 type, and i1 represents a 1-bit integer type. The xor instruction is used to perform an exclusive or operation on %v0 and true (i.e., 1), and the result is stored in %negated. In this way, if %v0 is 1 (true), %negated will be 0 (false), and vice versa, if %v0 is 0 (false), %negated will be 1 (true).

[0260] In the second row, INT_NEGATE represents an operation of bitwise negating a variable. ~ v0 represents an operation of bitwise negating variable v0. In the instruction “%negated = xor i32 %v0, -1”, %v0 is a variable of i32 type, and i32 represents a 32-bit integer type. The xor instruction is used to perform an exclusive or operation on %v0 and -1. In binary representation, -1 is equivalent to a number with all bits being 1 (for example, for a 32-bit integer, -1 is represented as 0xFFFFFFFF), so performing an exclusive or operation with -1 will achieve the effect of bitwise negation. The result is stored in %negated.

[0261] In line 3, INT_2COMP represents the two's complement negation operation on a variable, and -v0 represents the two's complement negation operation on v0. In the instruction "%negated = sub i32 0, %v0", %v0 is the variable to be subjected to the two's complement negation operation. The sub instruction is used to subtract %v0 from 0, achieving the effect of two's complement negation. The result is stored in %negated.

[0262] In line 4, FLOAT_NEG represents the negation operation, and the return value is of floating-point (float) type. f-v0 represents the negation operation on v0. In the instruction "%negated = fsub float 0.0, %v0", %v0 is the variable to be negated. The fsub instruction is used to subtract %v0 from 0.0, achieving the negation effect. The result is stored in %negated.

[0263] In line 5, INT_MULT represents the multiplication operation, and v0 * v1 represents multiplying v0 by v1. In the instruction "%result = mul i32 %v0, %v1", %v0 and %v1 are integer variables to be multiplied. The mul instruction is used to multiply %v0 and %v1. The result is stored in %result.

[0264] In line 6, INT_DIV represents the division operation, and v0 / v1 represents dividing v0 by v1. In the instruction "%result = udiv i32 %v0, %v1", %v0 and %v1 are integer variables to be divided. The udiv instruction is used to divide %v0 by %v1. The result is stored in %result.

[0265] In line 7, INT_SDIV represents the signed division operation, and v0 s / v1 represents performing a signed division operation on v0 and v1. In the instruction "%result = sdiv i32 %v0, %v1", %v0 and %v1 are integer variables to be subjected to the signed division operation. The sdiv instruction is used to divide %v0 by %1, performing signed integer division. The result is stored in %result.

[0266] In line 8, INT_REM represents the remainder operation, and v0 % v1 represents performing a remainder operation on v0 and v1. In the instruction "%result = urem i32 %v0, %v1", %v0 and %v1 are integer variables to be subjected to the remainder operation. The urem instruction is used to calculate the remainder of %v0 divided by %v1, performing unsigned integer remainder. The result is stored in %result.

[0267] In line 9, INT_SREM represents a signed integer remainder operation. v0 s% v1 represents performing a signed integer remainder operation on v0 and v1. In the instruction "%result = srem i32 %v0, %v1", %v0 and %v1 are integer variables on which the remainder operation is to be performed, and the srem instruction is used to calculate the remainder of %v0 divided by %v1, performing a signed integer remainder. The result is stored in %result.

[0268] As can be seen from Table 1, the electronic device can convert the P-Code instruction into an LLVM IR instruction based on the operation mapping relationship.

[0269] In Case 2, the operation type of the first P-Code instruction is to obtain an input variable. In this case, the first P-Code instruction is an instruction to obtain an input variable. For example, it can be a load instruction.

[0270] In the case of Case 2, the electronic device can obtain the storage type of the input variable in the first P-Code instruction, and based on the storage type of the input variable and the variable table corresponding to the current function, convert the first P-Code instruction into an LLVM IR instruction.

[0271] Among them, the storage type can include: register type, memory type, stack type, const type, unique type.

[0272] Among them, the variable table is used to record the real-time change situation of the variables in the function. The variable table can include multiple pieces of variable information, and each piece of variable information can include the variable name, the address of the variable, and the value of the variable. It should be noted that for any function, there can be a corresponding variable table, which is used to record the value change situation of the variables in the function during the program running process. The variables recorded in the variable table can include global variables, local variables, temporary variables, etc. Optionally, the initial variable table corresponding to the function can be an empty table by default.

[0273] Since variables of different types have different storage locations, uses, scopes, and lifecycles, and as the program runs, the values of the variables in each function will change. In this application, by using the variable table of the function to describe the real-time change situation of each variable in the function, the accuracy of the generated LLVM IR instructions can be guaranteed. For example, if a variable is not recorded in the variable table when it is updated, it means there is a problem with the use of this variable. Therefore, based on the real-time updated variable table, the accuracy of the generated LLVM IR instructions can be guaranteed.

[0274] (1) If the storage type of the input variable is of the register type, the electronic device generates a first target instruction for loading the value into the register variable (i.e., the input variable), and writes the variable information of the input variable into the variable table. In this case, the first target instruction is an LLVM IR instruction corresponding to the first P-Code instruction.

[0275] It should be noted that when the input variable is of the register type, since a variable of the register type (which can also be called a register variable) does not have a fixed memory address, the initial value of the register variable is usually stored at a certain address in the memory. During program execution, this value needs to be loaded into the register.

[0276] (2) If the storage type of the input variable is of the memory type, the electronic device can generate a second target instruction for obtaining the value of the input variable, and writes the variable information of the input variable into the variable table. In this case, the second target instruction is an LLVM IR instruction corresponding to the first P-Code instruction.

[0277] It should be noted that when the storage type of the input variable is of the memory type, the input variable is assigned a fixed memory address, so the value of the input variable can be directly read from the memory address.

[0278] (3) If the storage type of the input variable is of the stack type, the electronic device can generate a third target instruction for obtaining the value of the input variable from the variable table, and writes the variable information of the input variable into the variable table. In this case, the third target instruction is an LLVM IR instruction corresponding to the first P-Code instruction.

[0279] It should be noted that since stack variables are usually local variables with a scope within a function, the initial values of stack variables are usually passed from other global variables, such as register variables or memory variables. In addition, obtaining the value of the input variable from the variable table can achieve the effective transfer of variable values.

[0280] The form of LLVM IR instructions is in Static Single Assignment (SSA) form, that is, each variable should only be assigned a value once (written only once). As an example, if the eax variable is written twice, then the first write of 1 to eax can be expressed as eax_0 = 1, and the second write of incrementing the value of eax by 1 can be expressed as eax_1 = eax_0 + 1. That is to say, when using the SSA form to express instructions, by renaming variables, each variable can be assigned a value only once. In the embodiments of this application, writing variable information into the variable table in a timely manner can ensure that the variable table can always record the latest value of the variable. During the process of variable assignment, the cumbersome recursive search for values can be avoided, and the P-Code instructions can be quickly and accurately converted into LLVM IR instructions, thereby improving the accuracy of function name recovery.

[0281] Optionally, when the input variable is of the stack type, during the process of the electronic device generating the third target instruction to obtain the value of the input variable from the variable table, if the electronic device detects that the input variable does not exist in the current variable table, it can generate a first prompt message, which is used to prompt that the input variable is not defined, or there is a syntax error in the use of the input variable.

[0282] Situation 3, the operation type of the first P-Code instruction is to update the output variable. In this case, the first P-Code instruction is an instruction to update the output variable. For example, it can be a store instruction.

[0283] In the case of Situation 3, the first P-Code instruction includes the output variable to be updated and the input variable used for the update. The electronic device can obtain the storage type of the output variable in the first P-Code instruction, and based on the storage type of the output variable and the variable table corresponding to the current function, convert the first P-Code instruction into an LLVM IR instruction.

[0284] (1) If the storage type of the output variable is of the register type, the electronic device generates a fourth target instruction for converting the data type of the output variable into a first data type, generates a fifth target instruction for storing the output variable of the first data type, and adds the latest variable information of the output variable to the variable table. The first data type is the data type corresponding to the output variable. In this case, the fourth target instruction and the fifth target instruction are LLVM IR instructions corresponding to the first P-Code instruction. As an example, the fourth target instruction can be bitcast i64* @"RAX" to i64**, which means converting the data type of the register variable @"RAX" to the i64** type. Among them, i64* @"RAX" indicates that the data type of the register variable @"RAX" is the i64* type, that is, a pointer to a 64-bit integer. i64** represents a pointer to i64*.

[0285] It should be noted that the storage type can indicate the storage location of the variable. For example, the register type means storing in the register, the memory type means the variable is stored in the memory, and the stack type means the variable is stored in the stack. The data type of the variable refers to the type of data in the variable, such as pointer type, vector type, etc.

[0286] (2) If the storage type of the output variable is of the memory type, the electronic device can generate a sixth target instruction for storing the output variable and add the latest variable information of the output variable to the variable table. In this case, the sixth target instruction is an LLVM IR instruction corresponding to the first P-Code instruction.

[0287] (3) If the storage type of the output variable is of the stack type, the electronic device can generate a seventh target instruction for storing the output variable and add the latest variable information of the output variable to the variable table. In this case, the seventh target instruction is an LLVM IR instruction corresponding to the first P-Code instruction.

[0288] It should be noted that memory variables are stored in the memory, and memory variables are global variables, and the data type usually remains unchanged. Register variables are stored in the register. During the process of data update or processing of register variables, the data type may change. The stack is a memory area used to store local variables, function parameters, and return addresses during function calls. Variables on the stack are usually created during function calls and destroyed when the function returns. Stack variables are already in the memory, and the electronic device can directly change the type or value of these variables by updating the entries in the variable table without the need for data type conversion operations such as bitcast operations. Therefore, directly update the variables in the corresponding variable table.

[0289] In the embodiments of the present application, during the process of instruction conversion, in combination with the storage type of the output variable, the first P-Code instruction is specifically converted into an LLVM IR instruction, which can improve the accuracy of instruction conversion and further improve the function name recovery efficiency.

[0290] Case 4, the operation type of the first P-Code instruction is of the type of branch merging. In this case, the first P-Code instruction is a branch merging instruction. For example, it can be a MULTIEQUAL instruction.

[0291] Among them, branch merging means that a variable may have different values under different control flow paths. In practical applications, in the P-Code language, the operator for branch merging can be MULTIEQUAL, and in the LLVM IR language, the operator for branch merging can be phi.

[0292] In the embodiments of the present application, since in the SSA form, each variable is only assigned once, this makes the version management and optimization of variables simpler. However, when the control flow merges (for example, after an if-else branch statement), a variable may originate from multiple different paths, and each path may assign different values to this variable. The branch merging instruction is used to handle this situation, which merges the values on different paths into a new variable value.

[0293] The format of the MULTIEQUAL instruction can be as follows: MULTIEQUAL <output> <input1> <input2> ... <inputn>, where output is the merged variable, input1 is the variable input from the first branch path, and inputN is the variable input from the Nth branch path. N is the total number of branch paths.

[0294] The format of the phi instruction is: %result = phi <type> [ <value1> , <label1> ], [ <value2> , <label2> ], ... [ <valuen> , <labeln>. Among them, label1, label2... labelN are the identifiers of basic blocks respectively. value1 represents the value assigned to the variable in the basic block indicated by label1, value2 represents the value assigned to the variable in the basic block indicated by label2, and labelN represents the value assigned to the variable in the basic block indicated by labelN.

[0295] For example, if there are two basic blocks (label1 and label2), and they both pass a certain value to result, and the LLVM IR is as follows: %result = phi i32 [ 42, %label1 ], [ 100, %label2 ]. In this case, if the program control flow comes from the label1 basic block, the value of result will be 42. If the program control flow comes from the label2 basic block, the value of result will be 100.

[0296] In the case of Scenario 4, the electronic device can Figure 13 through the operations shown, convert the first P-Code instruction into an LLVM IR instruction, or in other words, convert the MULTIEQUAL instruction into a phi instruction. Among them, Figure 13 This is a schematic diagram of the conversion process of the branch merge instruction provided by the embodiment of the present application. As Figure 13 shown, converting the branch merge instruction into an LLVM IR instruction can at least include the following steps 50421 to step 50424.

[0297] Step 50421, if the first P-Code instruction is a branch merge instruction, the electronic device obtains each input variable in the first P-Code instruction from the first P-Code instruction.

[0298] Step 50422, the electronic device obtains the forward node information of the first P-Code instruction and determines the forward basic block where the forward node indicated by the forward node information is located.

[0299] Among them, the forward node information is used to indicate the forward node. The forward node information can be the identifier of the forward node. The forward node is a jump statement (or called a jump instruction) that jumps to the current basic block.

[0300] Among them, the forward basic block is the basic block where the forward node is located.

[0301] Here, the electronic device can obtain the forward node information of the first P-Code instruction and various information of the forward basic block based on the structure information of the P-Code code recorded in the intermediate file.

[0302] Step 50423: The electronic device obtains the latest value of each input variable from the corresponding forward basic block based on the variable name of each input variable.

[0303] Here, the electronic device can use the variable name of the input variable to look up the latest value of the variable with the same name in the forward basic block from the variable table corresponding to the forward basic block.

[0304] Step 50424: The electronic device generates LLVM IR instructions corresponding to the first P-Code instruction based on the latest values of the input variables.

[0305] Here, the electronic device can use the latest values of the input variables in the forward basic block to combine and generate the LLVM IR instructions corresponding to the first P-Code instruction. For example, if the first P-Code instruction is a MULTIEQUAL instruction, the LLVM IR instruction corresponding to the first P-Code instruction is a phi instruction. At this time, the electronic device can use the values of the input variables and the identifier of the forward basic block to combine to obtain the phi instruction.

[0306] Optionally, the electronic device can delete the first P-Code instruction before Step 50424, and during the execution of Step 50424, write the generated LLVM IR instruction at the original position of the first P-Code instruction.

[0307] It should be noted that through the above Step 5042, the P-Code instructions of various operation types described in the intermediate file can be converted into corresponding LLVM IR instructions, which can ensure that each P-Code instruction involved in the function is converted into an LLVM IR instruction, that is, it can achieve a comprehensive and accurate conversion of the instructions involved in the function. In this way, it helps to comprehensively and accurately extract the characteristic information of the function, thereby improving the accuracy of function name recovery.

[0308] The following combines Figure 14 and Figure 15 to elaborate on the differences between the LLVM IR file converted by the above Step 5042 and directly converting the P-Code code obtained in Step 501 into an LLVM IR file. Among them, Figure 14 is a schematic diagram of directly converting the P-Code code into an LLVM IR file, Figure 15 is a schematic diagram of converting the LLVM IR file from the intermediate file. It should be noted that Figure 14 the source code corresponding to the LLVM IR file shown in Figure 15 is the same as the source code corresponding to the LLVM IR file shown in

[0309] It should be noted that in the embodiments of the present application, bold fonts are used to display the jump relationships between basic blocks in a function. It can be understood that the embodiments of the present application do not make specific limitations on how to display the jump relationships between basic blocks in a function.

[0310] In the LLVM IR file, the "br" instruction can be used to represent jumps. For example, br label %"00100712" means jumping to basic block 00100712. Here, 00100712 is the identifier of the basic block. In the present application, the identifier of a basic block can be the entry address of the basic block. preds represents the predecessor basic block. For example, preds=%16 means the predecessor basic block is the basic block with the identifier 16. The @+ character represents a global variable or a function. For example, @rand() represents a function, and @EAX represents a global variable. The %+ number represents a local variable or a basic block. For example, %6 can represent a local variable.

[0311] Since the LLVM IR file is in SSA form, each variable is assigned a value only once. In the LLVM IR file, numbers can be used to represent variables or basic blocks, and the numbers used to represent local variables or basic blocks do not repeat. That is to say, after using the number 1 to represent a certain local variable, the number 1 will not be used to represent other local variables or basic blocks.

[0312] Figure 14 There are 12 basic blocks, namely basic block 0 (i.e., the basic block with the identifier 0), basic block 9, basic block 11, basic block 12, basic block 14, basic block 15, basic block 16, basic block 21, basic block 22, basic block 23, basic block 24, and basic block 25. Others represented by "%+ number" are local variables in the function. For example, %8 is a local variable. It should be noted that considering the relatively long length of the LLVM IR file, Figure 14 in the file, a dashed line is used to separate the first half and the second half of the file. Specifically, the left side of the dashed line is the first half, and the right side of the dashed line is the second half.

[0313] Figure 15 There are 7 basic blocks, namely basic block 001006c3, basic block 00100701, basic block 00100707, basic block 00100730, basic block 0010070d, basic block 00100712, and basic block 00100729.

[0314] Combined with Figure 14 and Figure 15 it can be found that Figure 14 Directly converting the P-Code into an LLVM IR file results in a chaotic jump relationship among basic blocks in the obtained LLVM IR file, making it difficult to describe the structural characteristics of the function. Figure 15 The LLVM IR file converted from the intermediate file has a clear and concise structure, which can clearly reflect the structural and semantic characteristics of the function. That is to say, the function name recovery method provided by the embodiments of the present application can extract an LLVM IR file with a clear and concise structure for the target binary file. This LLVM IR file can clearly reflect the structural and semantic characteristics of the function. That is to say, the function name recovery method provided by the embodiments of the present application can accurately extract the characteristic information of the function.

[0315] Step 505, the electronic device generates a control flow graph corresponding to the LLVM IR file based on the LLVM IR file.

[0316] Among them, the control flow graph is an abstract representation of the program, representing all paths that will be traversed during the program execution. The control flow graph represents the possible flow directions of all basic blocks in the program in the form of a graph, and also reflects the real-time execution process of the program. The control flow graph consists of nodes and edges. The nodes represent basic blocks, and the edges represent the paths of the control flow, that is, the paths from one basic block to another basic block.

[0317] Here, the electronic device can adopt the solutions disclosed in the related technologies to convert the LLVM IR file into a control flow graph, which will not be elaborated here.

[0318] It should be noted that directly generating a control flow graph from the intermediate file may depend on a specific compiler or language implementation. Since LLVM IR is an intermediate representation independent of the specific hardware platform, it converts the P-Code from a high-level language into a more standardized low-level representation, and can obtain a unified representation independent of the platform and language, which is convenient for analysis and optimization. And first converting the intermediate file into an LLVM IR file can utilize the powerful functions of LLVM for cross-platform optimization and analysis, and simplify the generation process of the control flow graph, making the analysis and optimization more efficient and general.

[0319] See Figure 16 , Figure 16 is a schematic diagram of the control flow graph provided by the embodiments of the present application. Among them, Figure 16 The shown control flow graph corresponds to Figure 4 the shown LLVM IR file.

[0320] As Figure 16 shown, the control flow graph can clearly describe the jump relationship inside the function, that is, it can clearly describe the structural characteristics of the function.

[0321] Step 506, the electronic device encodes the control flow graph to obtain a graph vector corresponding to the control flow graph.

[0322] Wherein, the graph vector is a vector corresponding to the control flow graph. The graph vector can represent the feature information related to the control flow of the function.

[0323] Here, the electronic device can adopt the methods disclosed in the related technologies to encode the control flow graph to obtain the graph vector. As an example, the electronic device can input the control flow graph into the graph encoding model to obtain the graph vector corresponding to the control flow graph. Among them, the graph encoding model is a model used to convert the graph into a vector. The graph encoding model can be a graph-to-vector model (Graph2vec), or other models. It can be understood that the embodiments of the present application do not make specific limitations on the graph encoding model.

[0324] It should be noted that since the control flow graph clearly describes the structural features of each function in the program, the graph vector also clearly describes the structural features of each function in the program.

[0325] Optionally, in step 505, the electronic device can also adopt the solutions disclosed in the related technologies to generate one or more of the function call graph, data flow graph, and program dependency graph corresponding to the LLVM IR file. In this case, the graph vector generated in step 506 can include the graph vector corresponding to the control flow graph and one or more of the following: the graph vector corresponding to the function call graph, the graph vector corresponding to the data flow graph, and the graph vector corresponding to the program dependency graph.

[0326] Step 507, the electronic device inputs the graph vector into the pre-trained function name generation model to generate the function names corresponding to each function.

[0327] Wherein, the function name generation model is a pre-trained model for generating function names. The function name generation model is used to describe the relationship between the graph vector and the function name, where the graph vector can describe the function features.

[0328] As an example, the function name generation model can be a model obtained by training an initial model (such as a Convolutional Neural Network (CNN), Residual Network (ResNet), etc.) using machine learning methods based on training samples. It can be understood that the embodiments of the present application do not make specific limitations on the initial model.

[0329] The process of training the function name generation model is described below:

[0330] (1) Training method 1

[0331] First, obtain multiple functions for training, convert the control flow graph corresponding to each function into a graph vector, and convert the function name of the function into a function name vector. The function name vector is the vector corresponding to the function name.

[0332] Here, the control flow graph of the function can be converted into a graph vector in the same way as in step 506, which will not be elaborated here.

[0333] The electronic device can convert the function name into a vector in the manner disclosed in the related art. For example, the one-hot encoding method can be used to convert the function name into a vector.

[0334] Then, use the graph vector and the function name vector corresponding to the same function as positive samples, and use the graph vector and the function name vector corresponding to different functions as negative samples, so as to obtain a training sample set for training the initial function name generation model.

[0335] Among them, the training sample set includes positive samples and negative samples.

[0336] Finally, use the training sample set to train the initial function name generation model, so as to obtain the trained function name generation model.

[0337] Training method 1 can train the function name prediction model more accurately. However, training method 1 is more dependent on the quantity and quality of the training samples and cannot accurately predict the function names for a large range of function inputs.

[0338] (2) Training method 2

[0339] First, load a pre-trained model. Among them, the pre-trained model is usually a large deep learning model trained on a huge dataset. In some embodiments, the pre-trained model can be a generative model. It should be noted that since the generative model can process various complex data types, such as text, images, audio, etc., and can learn the latent representations of the data and adapt to different tasks. In addition, the generative model usually has strong generalization ability and can also perform well in scenarios outside the training data. Therefore, when the pre-trained model is a generative model, the accuracy of the generated function names can be further guaranteed.

[0340] Then, refer to training method 1 to prepare a training sample set.

[0341] Finally, use the training sample set to fine-tune the pre-trained model.

[0342] During the training process, the pre-trained model can learn how to predict the corresponding function name according to the input graph vector (represented in text form), and finally fine-tune a function name generation model suitable for the current scenario.

[0343] The function name generation model trained in training method 2 does not need to be trained with a large amount of training samples because the pre-trained model itself is a large deep learning model trained on a huge dataset.

[0344] It can be understood that the above training method 1 and training method 2 are only examples of training the function name generation model, and are not specific limitations on how to train the function name generation model.

[0345] It can be understood that when the input of the function name generation model is the graph vectors of multiple functions, the output can be the function names for each function. That is to say, the function name generation model can generate the function name of one function or the function names of multiple functions.

[0346] Optionally, if the electronic device trains the pre-trained model in training method 2 to obtain the function name generation model. In this case, when the electronic device uses the function name generation model, it can also optimize the finally obtained function name in combination with the candidate function name list and the large language model (LLM) obtained during the function name generation process by the function name generation model. It should be noted that the function name generation model can output a candidate function name list (abbreviated as candidate function name list) in combination with the characteristics of the function described by the graph vector, or directly output a single function name for the function.

[0347] Taking it a step further as an example, for binary file 1 that records function 1, if the function name generation model outputs candidate function name list 1, the electronic device can combine the LLVM IR file corresponding to binary file 1, the graph vector of the control flow graph corresponding to the LLVM IR file, and function name list 1 to obtain a prompt, and send the prompt to the LLM model. The LLM model can adjust function name list 1 and return the adjusted function names of each function to the electronic device. As an example, the prompt sent by the electronic device to the LLM model can be: "What is the most likely function name? The candidate function names are: function name 1, function name 2...; the graph vector of the control flow graph is...; the LLVM IR file is...". The reply sentence returned by the LLM model for the prompt can be: "I think the most likely function name is function name 1". It can be understood that the above prompt and reply sentence are only examples, and the embodiments of the present application do not specifically limit the content of the prompt and reply sentence.

[0348] It should be noted that since the LLM model is a natural language processing model based on deep learning and it can learn and understand the grammar and semantics of natural language, therefore, using the LLM model to optimize the function name output by the function name generation model helps to further improve the accuracy of function name recovery.

[0349] The following further combines Figure 17 to elaborate on the process of training the function name generation model in the above training method 2. Among them, Figure 17 FIG. is a schematic diagram of a process for training a function name generation model provided by an embodiment of the present application.

[0350] As Figure 17 shown, the process of training the function name generation model can be:

[0351] First, the electronic device can input the control flow graph corresponding to the binary file into the graph encoding model (Graph2vec), and the graph encoding model converts the control flow graph into a function representation. Among them, the operations for obtaining the control flow graph corresponding to the binary file can refer to Figure 5 Steps 501 to 505 in, which will not be elaborated here. The above function representation is the graph vector corresponding to the control flow graph, and the function representation can characterize the features of the function.

[0352] Then, the electronic device can generate positive sample pairs (referred to as positive samples for short) and negative sample pairs (referred to as negative samples for short) based on the function representation. Specifically, the function representation (graph vector) and the function name representation (function name vector) corresponding to the same function can be used as positive samples, and the function representations and function name representations corresponding to different functions can be used as negative samples. Combining Figure 17 , the positive sample pair includes two parts. The left part is the function representation corresponding to the function, and the right part is the function name representation of the function, that is, the function representation and the function name representation come from the same function. The negative sample pair also includes two parts. The left part is the function representation corresponding to the function, and the right part is the function name representation of other functions, that is, the function representation and the function name representation come from different functions.

[0353] Figure 17 In, for any sample pair, it can be determined whether the function representation and the function name representation come from the same function, or whether the sample is a positive sample or a negative sample, by the presence or absence of filling. Specifically, if the function representation and the function name representation come from the same function, then the elements (circles) on the left and right sides of the sample are both filled, and in this case, the sample is a positive sample. If the function representation and the function name representation come from different functions, then one of the elements (circles) on the left and right sides of the sample is filled, and in this case, the sample is a negative sample. It can be understood that other methods can also be used to distinguish positive samples and negative samples, and the embodiments of the present application do not limit this.

[0354] Combining Figure 17 It can be found that the function names in the positive samples (specifically, function name representations) come from the function name corpus, while the function names in the negative samples (specifically, function name representations) do not come from the function name corpus. In this case, the trained function name generation model can be made more robust.

[0355] After that, the electronic device can fine-tune the generative model using the obtained positive and negative samples, or in other words, train the generative model. Among them, the generative model is the initial function name generation model, and the generative model is a pre-trained model. It should be noted that the pre-trained model is usually a large deep learning model trained on a huge dataset.

[0356] Finally, after fine-tuning the generative model, the trained model (i.e., the function name generation model) can have the function representation (graph vector) as the input and the candidate function name list as the output. Optionally, the output of the function name generation model can also be the function name.

[0357] Next, further in conjunction with Figure 18 , it is described how to optimize the candidate function name list based on the large language model to obtain the accurate function name when the output of the function name generation model is the candidate function name list. Among them, Figure 18 is the schematic diagram of the process for optimizing the function name provided by the embodiment of this application.

[0358] As Figure 18 shown, the process for optimizing the function name can be as follows:

[0359] First, the electronic device sends the LLVM IR file, the graph vector of the control flow graph corresponding to the LLVM IR file, and the candidate function name list to the server.

[0360] Then, the server generates a question statement and sends the content of the statement to the large language model.

[0361] Among them, the content of the question statement can include: What is the most likely function name? The candidate function names are: function name 1, function name 2...; Other information is: graph vector, LLVM IR file. It should be understood that the embodiment of this application does not limit the specific content of the question statement.

[0362] After that, the large language model determines the most likely function name based on the question statement and sends the reply statement containing the most likely function name to the server.

[0363] As an example, the reply statement can be: "I think the most likely function name is function name 1". It can be understood that the embodiment of this application does not specifically limit the content of the reply statement.

[0364] It can be understood that when there is one function in the pseudocode, the server can obtain the function name of the function. When there are multiple functions in the pseudocode, the server can obtain the function name of each function.

[0365] Finally, the server sends the function names of the functions in the pseudocode to the electronic device.

[0366] It can be understood that Figure 18 in this case, the server and the electronic device can be two independent devices or the same device. For example, when the function of the server is implemented by the electronic device, the server and the electronic device can be the same device.

[0367] Figure 19 It is a schematic flowchart of a function name recovery method provided by an embodiment of the present application. Refer to Figure 19 This method can be applied to an electronic device.

[0368] As Figure 19 shown, the function name recovery method may include the following steps 1601 to 1606. It can be understood that the above function name recovery method may include all of the steps 1601 to 1606, or may only include some of the steps. It can be understood that the steps in the function name recovery method can be arbitrarily combined without conflict.

[0369] Step 1601, the electronic device obtains a target binary file.

[0370] Here, the electronic device can obtain the target binary file from local or from other electronic devices.

[0371] Optionally, the target binary file is a binary file from which symbol information has been removed. In the present application, a binary file refers to a file carrying the machine code of a program. The file format of the binary file may include: object file format (.o file), shared object file format (.so file), executable file format (.exe file), etc. It can be understood that the embodiment of the present application does not make a specific limitation on the file format of the target binary file.

[0372] Step 1602, the electronic device generates a target pseudocode corresponding to the target binary file.

[0373] Among them, the target pseudocode is the pseudocode corresponding to the target binary file (which can also be called P-Code code).

[0374] The P-Code code is a set of P-Code instructions, abbreviated as the P-Code instruction set.

[0375] Here, for the operation of step 1602, reference can be made to Figure 5 The embodiment part corresponding to step 501 in

[0376] It should be noted that the target pseudocode may include one function or multiple functions. In the case where the target pseudocode includes multiple functions, the electronic device can extract the structural information of each function and generate the function name of the function based on the structural information of each function.

[0377] In some optional implementation manners of the embodiments of the present application, in step 1602, when the electronic device generates the target pseudocode corresponding to the target binary file, it may include: the electronic device generates the high-level pseudocode corresponding to the target binary file, and the target pseudocode is the high-level pseudocode.

[0378] In the embodiments of the present application, since the high-level pseudocode (or called High P-Code) pays more attention to the source code itself and does not pay attention to the underlying architecture, while the low-level pseudocode (or called Low-P-Code) contains a large amount of information related to the underlying architecture, and this information related to the underlying architecture has nothing to do with the function itself. Therefore, using the high-level pseudocode for function name recovery can avoid excessive irrelevant interference information, help reduce the amount of data processing, and improve the function name recovery efficiency.

[0379] Step 1603, the electronic device extracts the structural information of the target function from the target pseudocode.

[0380] Among them, the structural information is used to describe the function, the basic blocks in the function, and the variables in the function. The target function is any function in the target pseudocode. The structural information of the target function may include function information, basic block information in the target function, and variable information in the target function.

[0381] Here, for the operation of step 1603, reference can be made to Figure 5 the embodiment part corresponding to step 502 in

[0382] In some optional implementation manners of the embodiments of the present application, in step 1603, when the electronic device extracts the structural information of the target function from the target pseudocode, it may include:

[0383] First, the electronic device traverses the target pseudocode, and when accessing the first operation instruction, extracts the function information of the target function from the first operation instruction.

[0384] Among them, the first operation instruction is the instruction that defines the target function in the target pseudocode, and the function information includes the first entry address of the function. The first entry address can also be called the entry address of the function or the head address of the function.

[0385] Here, for the operation of the electronic device to extract the function information, reference can be made to Figure 6 The embodiment part corresponding to step 601 in

[0386] Then, according to each instruction in the target function, the electronic device determines M basic blocks and the basic block information of each basic block from the target function.

[0387] Among them, the basic block information includes the second entry address of the basic block, the forward nodes of the basic block, and the backward nodes of the basic block, where M is an integer greater than 0. The second entry address can also be referred to as the entry address of the basic block or the first address of the basic block. It should be noted that for any basic block, the forward node of the basic block is the pseudo-code instruction (P-Code instruction) that jumps to the basic block, and the backward node of the basic block is the P-Code instruction that the basic block needs to jump to.

[0388] Here, for the operation of the electronic device to extract the basic block information of the basic blocks in the function, reference can be made to Figure 6 The embodiment part corresponding to step 602 in

[0389] After that, the electronic device traverses the first basic block and extracts variable information from the first instruction in the first basic block.

[0390] Among them, the first basic block is any one of the M basic blocks, and the first instruction is any instruction in the first basic block.

[0391] Here, for the operation of the electronic device to extract the variable information of the variables in the function, reference can be made to Figure 6 The embodiment part corresponding to step 603 in

[0392] Finally, the electronic device generates an intermediate file.

[0393] Among them, the intermediate file is used to record the function information, basic block information, and variable information of the target function. The first entry address and the second entry address in the intermediate file correspond to the first pseudo-code. It can be understood that the target pseudo-code can include one function or multiple functions. In the case where the target pseudo-code includes multiple functions, in the intermediate file generated by the electronic device, the function information of each function, the basic block information in the function, and the variable information in the function can be recorded.

[0394] Here, the content of the intermediate file can be referred to Figure 8 and Figure 9 , the structure of the intermediate file can be referred to Figure 10 and Figure 11 . For the operation of the electronic device to generate the intermediate file, reference can be made to Figure 5 The embodiment part corresponding to step 503 in

[0395] In the embodiments of the present application, since the pseudo-code has a specific syntax, the electronic device can extract function information, basic block information in the function, and variable information in the function from the target code in combination with the syntax of the pseudo-code. It should be noted that since the pseudo-code is essentially unordered code, extracting the structural information of each function from the pseudo-code can achieve the orderly classification of the unordered pseudo-code. In addition, classifying and dividing the pseudo-code layer by layer according to multiple levels such as functions, basic blocks in the functions, instructions in the basic blocks, and variables helps to accurately extract the semantic structure of each function, thereby ensuring the accuracy of function name recovery.

[0396] Optionally, the variable information includes one or more of the following: the name of the variable, the storage type of the variable, and the size of the variable. Among them, the storage type of the variable can indicate the storage location of the variable. For example, the register type means that the variable is stored in the register, the memory type means that the variable is stored in the memory, and the stack type means that the variable is stored in the stack.

[0397] Optionally, the language format of the three-address code is the Low-Level Virtual Machine Intermediate Representation LLVM IR.

[0398] It should be noted that since LLVM IR is an intermediate representation independent of a specific hardware platform, it converts the P-Code from a high-level language into a more standardized low-level representation, and can obtain a unified representation independent of the platform and language, which is convenient for analysis and optimization. And by first converting the intermediate file into an LLVM IR file, the powerful functions of LLVM can be used for cross-platform optimization and analysis, and the generation process of the control flow graph can be simplified, making the analysis and optimization more efficient and general.

[0399] Step 1604, the electronic device converts the first pseudo-code in the target pseudo-code into three-address code based on the structural information of the target function.

[0400] Among them, the first pseudo-code is related to the target function. That is to say, the first pseudo-code is the part of the target pseudo-code related to the target function.

[0401] Here, for the operation of step 1604, reference can be made to Figure 5 the corresponding embodiment part of step 504 in

[0402] In some optional implementation manners of the embodiments of the present application, the above step 1604 may include the following step one and step two.

[0403] Step 1: The electronic device generates an initial variable table corresponding to the target function. The variable table is used to record the value changes of each variable in the target function during the program execution. It should be noted that for any function, a corresponding variable table can be generated, which is used to record the value changes of the variables in the function during the program execution. The variables recorded in the variable table can include global variables, local variables, temporary variables, etc. Optionally, the initial variable table corresponding to the function can be an empty table by default.

[0404] Step 2: Based on the intermediate file, the electronic device traverses the first pseudocode. For the second instruction in the first pseudocode (see Figure 12 the first P-Code instruction in step 5042), according to the operation type and / or variable table corresponding to the second instruction, the second instruction is converted into an LLVM IR instruction.

[0405] Among them, the second instruction is any instruction in the first pseudocode, and the operation types include: the first operation type, the second operation type, the third operation type, and the fourth operation type. The first operation type is used to obtain input variables, the second operation type is used to update output variables, the third operation type is used to merge multiple branches, and the fourth operation type is an operation type other than the first, second, and third operation types.

[0406] Here, the operations in step 1 and step 2 can be referred to Figure 12 the corresponding embodiment part.

[0407] In the embodiments of the present application, converting the pseudocode instructions of various operation types described in the intermediate file into corresponding LLVM IR instructions can ensure that each pseudocode instruction involved in the function is converted into an LLVM IR instruction, that is, it can achieve a comprehensive and accurate conversion of the instructions involved in the function. In this way, it helps to comprehensively and accurately extract the characteristic information of the function, thereby improving the accuracy of function name recovery. In addition, since different types of variables have different storage locations, uses, scopes, and lifecycles, and as the program runs, the values of the variables in each function will change. In the present application, the real-time changes of each variable in the function are described through the variable table of the function, which can ensure the accuracy of the generated LLVM IR instructions. For example, if a variable is not recorded in the variable table when updating a certain variable, it indicates that there is a problem with the use of the variable. Therefore, based on the real-time updated variable table, the accuracy of the generated LLVM IR instructions can be ensured.

[0408] In some alternative implementations, step two above may include: if the operation type of the second instruction is the first operation type (or the type of obtaining an input variable), obtain the storage type of the input variable from the second instruction; generate an LLVM IR instruction corresponding to the second instruction according to the storage type of the input variable and the variable table. The storage type includes register type (or register type), memory type (or memory type), and stack type (or stack type).

[0409] In this case, if the storage type of the input variable is register type, the LLVM IR instruction is an instruction for loading a value into the input variable. If the storage type of the input variable is memory type, the LLVM IR instruction is an instruction for obtaining the value of the input variable. If the storage type of the input variable is stack type, the LLVM IR instruction is an instruction indicating to obtain the value of the input variable from the variable table.

[0410] Here, when the operation type of the second instruction is the first operation type, the process of the electronic device converting the second instruction into an LLVM IR instruction can refer to the embodiment part corresponding to "Situation 2" in the specification.

[0411] In the embodiment of the present application, during the process of instruction conversion, by combining the storage type of the input variable to specifically convert the second instruction into an LLVM IR instruction, the accuracy of instruction conversion can be improved, thereby further improving the function name recovery efficiency.

[0412] Optionally, after the electronic device generates an LLVM IR instruction corresponding to the second instruction according to the storage type of the input variable and the variable table, the electronic device may also add the first variable information of the input variable to the variable table. The first variable information includes the variable name of the input variable, the address of the input variable, and the value of the input variable.

[0413] In the embodiment of the present application, when the electronic device obtains an input variable, adding the variable information of the input variable to the variable table in a timely manner can enable the variable table to always record the real-time changes of each variable.

[0414] Optionally, when the input variable is of register type, during the process of the electronic device generating an LLVM IR instruction, if the electronic device detects that the input variable does not exist in the current variable table, a first prompt message can be generated. The first prompt message is used to prompt that the input variable is not defined, or there is a syntax error in the use of the input variable. That is to say, in the present application, the variable table can be used to check whether there are syntax errors in the source code.

[0415] In some alternative implementations, step two above may include: If the operation type of the second instruction is the second operation type (or the type of updating the output variable), obtain the storage type of the output variable from the second instruction; generate the LLVM IR instruction corresponding to the second instruction according to the storage type of the output variable.

[0416] In this case, if the storage type of the output variable is the register type, the LLVM IR instruction includes the first LLVM IR instruction and the second LLVM IR instruction. The first LLVM IR instruction is used to convert the data type of the output variable into the first data type, and the second LLVM IR instruction is used to store the output variable of the first data type. The first data type is the data type corresponding to the output variable. If the storage type of the output variable is the memory type or the stack type, the LLVM IR instruction is the instruction for storing the output variable.

[0417] Here, when the operation type of the second instruction is the second operation type, the process of the electronic device converting the second instruction into the LLVM IR instruction can refer to the embodiment part corresponding to "Situation 3" in the specification.

[0418] In the embodiments of the present application, during the process of instruction conversion, by combining the storage type of the output variable to specifically convert the second instruction into the LLVM IR instruction, the accuracy of instruction conversion can be improved, thereby further improving the function name recovery efficiency.

[0419] Optionally, after the electronic device generates the LLVM IR instruction corresponding to the second instruction according to the storage type of the output variable, it further includes: adding the second variable information of the output variable to the variable table, where the second variable information includes the variable name of the output variable, the address of the output variable, and the value of the output variable.

[0420] In the embodiments of the present application, the electronic device timely adds the latest variable information of the output variable to the variable table, which can realize that the variable table always records the real-time changes of each variable.

[0421] In some alternative implementations, step two above may include: If the operation type of the second instruction is the third operation type (or the type of branch merging), obtain the N input variables included in the second instruction; based on the intermediate file, determine the forward basic block of the second basic block; obtain the latest value of the first input variable from the first forward basic block; generate the LLVM IR instruction according to the obtained values of the input variables.

[0422] Wherein, the second basic block is the basic block to which the second instruction belongs, and the forward basic block is the basic block that jumps to the second basic block. N is an integer greater than 1.

[0423] Among them, the first forward basic block is the forward basic block containing the first input variable, and the first input variable is any input variable in the second instruction.

[0424] Here, when the operation type of the second instruction is the third operation type, the process of the electronic device converting the second instruction into an LLVM IR instruction can refer to the embodiment part corresponding to "Situation 4" in the specification.

[0425] In the embodiment of the present application, when the operation type of the second instruction is the third operation type, the electronic device can accurately and effectively convert the second instruction in the pseudocode into an LLVM IR instruction. During the process of instruction conversion, by combining the operation type of the instruction to specifically convert the second instruction into an LLVM IR instruction, the accuracy of instruction conversion can be improved, thereby further improving the function name recovery efficiency.

[0426] In some optional implementation manners, step two above may include: if the operation type corresponding to the second instruction is the fourth operation type (or referred to as the regular type), then based on a preset first correspondence, the second instruction is converted into an LLVM IR instruction. Among them, the first correspondence is the correspondence between the instructions of the pseudocode and the LLVM IR instructions.

[0427] Here, when the operation type of the second instruction is the fourth operation type, the process of the electronic device converting the second instruction into an LLVM IR instruction can refer to the embodiment part corresponding to "Situation 1" in the specification.

[0428] In the embodiment of the present application, during the process of instruction conversion by the electronic device, by combining the operation type of the instruction to specifically convert the second instruction into an LLVM IR instruction, the accuracy of instruction conversion can be improved, thereby further improving the function name recovery efficiency.

[0429] Step 1605, the electronic device determines first feature information related to the control flow of the target function based on the three-address code.

[0430] Among them, the first feature information can be implemented as a graph vector.

[0431] Here, the operation of step 1605 can refer to Figure 5 the embodiment parts corresponding to steps 505 and 506 in

[0432] Step 1606, the electronic device generates the function name of the target function based on the first feature information.

[0433] In some alternative implementation manners of the embodiments of the present application, the above step 1606 may include: First, the electronic device inputs the first feature information into the first model (or called the function name generation model) to obtain multiple candidate function names of the target function. Then, the multiple candidate function names, the first feature information, and the three-address code are input into the second model to obtain the function name of the target function.

[0434] Among them, the first model is a generative model, and the second model is a large language model.

[0435] Here, for the operations of step 1606, reference can be made to Figure 5 the corresponding embodiment part of step 507 in

[0436] In the embodiments of the present application, since the generative model can process various complex data types, such as text, images, audio, etc., and can learn the latent representations of the data and adapt to different tasks. In addition, the generative model usually has strong generalization ability and can also perform well in scenarios outside the training data. Therefore, when the first model is a generative model, the accuracy of the obtained multiple candidate function names can be ensured. In addition, since the large prediction model is a natural language processing model based on deep learning and it can learn and understand the grammar and semantics of natural language, therefore, using the large language model to optimize the multiple candidate function names output by the first model helps to further improve the accuracy of function name recovery.

[0437] It should be understood that the magnitudes of the sequence numbers of the above steps in the embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0438] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0439] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0440] As used in the specification and the appended claims of this application, the term "if" can be construed contextually as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be construed contextually to mean "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".

[0441] In addition, in the description of the specification and the appended claims of this application, the terms "first", "second", "third", etc. are only used for differential description and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first", "second", etc. are used in the text in some embodiments of this application to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0442] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in combination with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0443] In addition, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. In each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0444] The function name restoration method provided by the embodiments of the present application can be applied to an electronic device, which can be a device such as a terminal or a server. The terminal can be a tablet computer, an in-vehicle device, an Augmented Reality (AR) / Virtual Reality (VR) device, a laptop computer, an Ultra-Mobile Personal Computer (UMPC), etc. The server can be a network server, a cloud server, etc. The embodiments of the present application do not make any limitations in this regard.

[0445] To better understand the embodiments of the present application, taking the electronic device as a server as an example, the following Figure 20 introduces the structure of the electronic device in the embodiments of the present application.

[0446] Figure 20 is a schematic structural diagram of a server provided by an embodiment of the present application. The server includes at least one processor 111, a communication bus 112, a memory 113, and at least one communication interface 114.

[0447] The processor 111 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application.

[0448] The communication bus 112 may include a path for transmitting information between the above components.

[0449] The communication interface 114 uses any device such as a transceiver for communicating with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Networks (WLAN), etc.

[0450] The memory 113 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited to this. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.

[0451] Among them, the memory 113 is used to store the application program code for executing the solution of this application, and is controlled by the processor 111 to execute. The processor 111 is used to execute the application program code stored in the memory 113, so as to implement the material matching method in the above embodiments.

[0452] As an embodiment, the processor 111 can include one or more CPUs, such as CPU 0 and CPU 1.

[0453] As an embodiment, the server can include multiple processors, such as Figure 20 two of the processors 111 in. Each processor can be a single-core (ingle-CPU) processor or a multi-core (multi-CPU) processor.

[0454] As an embodiment, the server can also include an output device 115 and an input device 116. The output device 115 communicates with the processor 111 and can display information in various ways. For example, the output device 115 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device or a projector, etc.

[0455] The input device 116 communicates with the processor 111 and can accept user input in various ways. For example, the input device 116 can be a mouse, a keyboard, a touch screen or a sensing device, etc.

[0456] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the server. In other embodiments of the present application, the server may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0457] In addition, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example for illustration. In practical applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. In each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0458] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented.

[0459] The embodiments of the present application provide a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above various method embodiments when executed.

[0460] The embodiments of the present application also provide a chip system. The chip system includes a processor, the processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the steps in the above various method embodiments.

[0461] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0462] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0463] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0464] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0465] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.< / labeln> < / valuen> < / label2> < / value2> < / label1> < / value1> < / type> < / inputn> < / input2> < / input1> < / output>

Claims

1. A method for restoring function names, characterized in that, The method includes: Obtain a target binary file; Generate target pseudocode corresponding to the target binary file; According to the syntax of the target pseudocode, extract the structural information of a target function from the target pseudocode. The structural information includes function information, basic block information in the function, and variable information in the function. The function information is used to describe the function, the basic block information in the function is used to describe the basic blocks in the function, and the variable information in the function is used to describe the variables in the function. The target function is any function in the target pseudocode; Based on the structural information of the target function, convert the first pseudocode in the target pseudocode into three-address code. The first pseudocode is related to the target function; Based on the three-address code, determine first feature information related to the control flow of the target function; Based on the first feature information, generate the function name of the target function.

2. The function name recovery method according to claim 1, wherein The extracting the structural information of the target function from the target pseudocode includes: Traverse the target pseudocode. When accessing the first operation instruction, extract the function information of the target function from the first operation instruction. The first operation instruction is an instruction for defining the target function, and the function information includes the first entry address of the function; According to each instruction in the target function, determine M basic blocks and the basic block information of each basic block from the target function. The basic block information includes the second entry address of the basic block, the forward node of the basic block, and the backward node of the basic block, where M is an integer greater than 0; Traverse the first basic block and extract variable information from the first instruction in the first basic block. The first basic block is any one of the M basic blocks, and the first instruction is any instruction in the first basic block; Generate an intermediate file, which is used to record the function information, the basic block information, and the variable information of the target function. The first entry address and the second entry address in the intermediate file correspond to the first pseudocode. The structural information includes the function information, the basic block information, and the variable information.

3. The function name recovery method according to claim 2, wherein The variable information includes one or more of the following: the name of the variable, the storage type of the variable, and the size of the variable.

4. The function name recovery method according to claim 2, characterized in that, The language format of the three-address code is the low-level virtual machine intermediate representation LLVM IR. The converting the first pseudocode in the target pseudocode into three-address code based on the structural information of the target function includes: Generate an initial variable table corresponding to the target function, which is used to record the value change situation of each variable in the target function during the program running process; Based on the intermediate file, traverse the first pseudocode. For the second instruction in the first pseudocode, convert the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table; Among them, the second instruction is any instruction in the first pseudocode, and the operation types include: the first operation type, the second operation type, the third operation type, and the fourth operation type. The first operation type is used to obtain an input variable, the second operation type is used to update an output variable, the third operation type is used to merge multiple branches, and the fourth operation type is an operation type other than the first operation type, the second operation type, and the third operation type.

5. The function name restoration method according to claim 4, characterized in that Converting the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table includes: If the operation type of the second instruction is the first operation type, obtain the storage type of the input variable from the second instruction, where the storage type includes register type, memory type, and stack type; Generate the LLVM IR instruction corresponding to the second instruction according to the storage type of the input variable and the variable table. Among them, if the storage type of the input variable is the register type, the LLVM IR instruction is an instruction for loading a value into the input variable; if the storage type of the input variable is the memory type, the LLVM IR instruction is an instruction for obtaining the value of the input variable; if the storage type of the input variable is the stack type, the LLVM IR instruction is an instruction indicating to obtain the value of the input variable from the variable table.

6. The method for restoring a function name according to claim 5, wherein After generating the LLVM IR instruction corresponding to the second instruction according to the storage type of the input variable and the variable table, it further includes: Add the first variable information of the input variable to the variable table, where the first variable information includes the variable name of the input variable, the address of the input variable, and the value of the input variable.

7. The method for restoring function names according to claim 4, characterized in that Converting the second instruction into an LLVM IR instruction according to the operation type corresponding to the second instruction and / or the variable table includes: If the operation type of the second instruction is the second operation type, obtain the storage type of the output variable from the second instruction; Generate the LLVM IR instruction corresponding to the second instruction according to the storage type of the output variable. Among them, if the storage type of the output variable is the register type, the LLVM IR instruction includes a first LLVM IR instruction and a second LLVM IR instruction. The first LLVM IR instruction is used to convert the data type of the output variable into a first data type, and the second LLVM IR instruction is used to store the output variable of the first data type. The first data type is the data type corresponding to the output variable; if the storage type of the output variable is the memory type or the stack type, the LLVM IR instruction is an instruction for storing the output variable.

8. The function name recovery method according to claim 7, wherein After generating the LLVM IR instruction corresponding to the second instruction according to the storage type of the output variable, it further includes: Add the second variable information of the output variable to the variable table, where the second variable information includes the variable name of the output variable, the address of the output variable, and the value of the output variable.

9. The function name restoration method according to claim 4, wherein, Converting the second instruction into LLVM IR instructions according to the operation type corresponding to the second instruction and / or the variable table includes: If the operation type of the second instruction is the third operation type, obtain N input variables included in the second instruction, where N is an integer greater than 1; Based on the intermediate file, determine the forward basic block of the second basic block, where the second basic block is the basic block to which the second instruction belongs, and the forward basic block is the basic block that jumps to the second basic block; Obtain the latest value of the first input variable from the first forward basic block, where the first forward basic block is the forward basic block containing the first input variable, and the first input variable is any one of the input variables in the second instruction; Generate the LLVM IR instruction according to the values of the obtained input variables.

10. The method for restoring function names according to claim 4, characterized in that Converting the second instruction into LLVM IR instructions according to the operation type corresponding to the second instruction and / or the variable table includes: If the operation type corresponding to the second instruction is the fourth operation type, convert the second instruction into the LLVM IR instruction based on a preset first correspondence, where the first correspondence is the correspondence between the instructions of the pseudocode and the LLVM IR instructions.

11. The function name recovery method according to claim 1, characterized in that Generating the function name of the target function based on the first feature information includes: Input the first feature information into the first model to obtain multiple candidate function names of the target function; Input the multiple candidate function names, the first feature information, and the three-address code into the second model to obtain the function name of the target function, where the first model is a generative model and the second model is a large language model.

12. The function name restoration method according to any one of claims 1-11, characterized in that, Generating the target pseudocode corresponding to the target binary file includes: Generate high-level pseudocode corresponding to the target binary file, and the target pseudocode is the high-level pseudocode.

13. An electronic device, characterized in that, The electronic device includes a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the computer program, it implements the function name recovery method according to any one of claims 1 to 12.

14. A chip system, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, The processor is used to execute the computer program to implement the function name recovery method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. When the computer program runs on an electronic device, the electronic device is caused to execute the function name recovery method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Semantic-based multi-architecture binary function name prediction method

    CN115357890A

  • Cross-platform binary function naming prediction method based on context semantics

    CN117592478A