Hierarchical graph neural network for cross-architecture software reverse engineering
By converting the stripped binary files into graph representations of graphs and using machine learning models, the problem that reverse engineers have difficulty recovering symbols in old systems is solved, and the effect of automatically recovering symbols and improving reverse engineering efficiency is achieved.
Patent Information
- Application Number
- CN202380071260.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-06
- Filing Date
- 2023-08-16
- Publication Date
- 2025-05-27
AI Technical Summary
In legacy systems, the gap between source code version and binary version makes it difficult for reverse engineers to apply patches to legacy binary files, and recovering symbols from stripped binary files is not easy.
By representing the stripped binary as a graph (GoG) representation of the graph and converting it into an expressive representation of each function, training is used to determine the missing symbols in the function.
It realizes automatic recovery of symbols from stripped binary files, improves the efficiency and accuracy of reverse engineering, and reduces dependence on source code.
Smart Images

Figure CN120051762A_ABST
Abstract
Description
Technical Field
[0001] This application relates to software development. More specifically, this application relates to reverse engineering of binary representations of software source code. Background Art
[0002] Software needs to be updated regularly to address persistent issues (including security vulnerabilities that may be discovered over time). Typically, to address vulnerabilities, vendors provide patches in the form of source code changes based on the current software version (e.g., version 0.9). However, in legacy systems, the only available data is the binary file based on the corresponding source code (e.g., version 0.1). Such a version gap poses challenges in applying patches to legacy binaries, making the only solution for applying patches to legacy software to perform direct binary analysis. Reverse engineers (REs) must utilize software reverse engineering (SRE) tools such as Ghidra, HexRays, and radare2 to first disassemble and decompile the binary into a higher-level representation (e.g., C or C++). Typically, these tools use the binary's debug information, strings, and symbol table to reconstruct function names and variable names. This allows the RE to reconstruct the structure and functionality of the software without access to the source code. For the RE, these symbols encode the context of the source code and provide valuable information for their task. Next, the RE must understand the program's logic to achieve their goal of patching the vulnerable binary. However, to optimize the footprint of the binary in memory-constrained mission-critical legacy systems, symbols are typically excluded to reduce the binary size. Since recovering symbols from stripped binaries is not straightforward, most decompilers assign meaningless symbol names to the encoded elements. To understand the software semantics, the RE needs to utilize their experience and expertise to digest the information and then interpret the semantics of each encoded element. An improved method for determining missing symbols in a binary file is desired. Summary of the Invention
[0003] A method for recovering symbols from a stripped binary file according to an embodiment of the present disclosure includes: representing the stripped binary file as a graph of graphs (GoG) representation of multiple graphs, converting the multiple GoGs into multiple expressive representations for each function in the stripped binary file, using the expressive representations to train a machine learning (ML) model, and determining missing symbols for at least one function in the function based on the output of the ML model. The stripped binary file is further represented as multiple GoGs by: using a decompiler to export P-code and cross-call dependencies for each function included in the stripped binary file, and decompiling each function to extract basic blocks, summary attributes, and corresponding P-code of the function. Using this information, each function can be represented as a graph of basic blocks, and a GoG is created for all functions in the stripped binary file, where nodes contain graphs of basic blocks of the function, and edges represent relationships based on cross-call dependencies between functions.
[0004] According to an embodiment, converting the multiple GoGs into multiple expressive representations can be performed by: inputting the multiple GoGs into a hierarchical graph learning pipeline to generate expressive representations, the hierarchical graph learning pipeline including a control flow graph (CFG) embedding layer and a GoG embedding layer. The hierarchical graph learning pipeline includes performing the following steps: in the CFG embedding layer, using one or more graph convolutional network (GCN) layers to propagate features for each basic block of the function to generate an aggregated vectorized representation, and in the GoG embedding layer, receiving the aggregated vectorized representation, merging call function information and called function information into each function to generate an expressive representation for each function. Training the ML model using the expressive representations of the functions can be achieved via a contrast function similarity learning process. If the expressive representations of two functions indicate that the two functions behave similarly, then based on the embedding of one of the functions in the stripped binary file, the missing symbols of the function are determined. Each of the two functions can be built from different CPU architectures. According to one embodiment, the decompiler is a Ghidra decompiler. In some embodiments, the hierarchical graph learning pipeline can be implemented as a data toolkit for the Ghidra reverse engineering tool using the application programming interface (API) of the Ghidra reverse engineering tool. A software update of the stripped binary file can be performed partially based on the discovered missing symbols.
[0005] A system for recovering symbols from a stripped binary file includes: a computer processor in communication with a non-transitory computer memory, the non-transitory computer memory storing machine-readable instructions that, when executed by the computer processor, cause the computer processor to: represent the stripped binary file as a graph-of-graphs (GoG) representation of a plurality of graphs, convert the plurality of GoGs into a plurality of expressive representations for each function in the stripped binary file, use the expressive representations to train a machine learning (ML) model, and determine missing symbols for at least one function among the functions based on the output of the ML model. The stripped binary file can be represented as a plurality of GoGs using a decompiler and exporting P-code and cross-call dependencies for each function included in the stripped binary file. Each function can be further decompiled to extract the basic blocks, summary attributes, and corresponding P-code of the function, represent each function as a graph of basic blocks, and then create a GoG for all functions in the stripped binary file, where the nodes of the GoG contain graphs of the basic blocks of the functions and the edges represent relationships based on cross-call dependencies between the functions.
[0006] Converting the plurality of GoGs into a plurality of expressive representations is performed by: inputting the plurality of GoGs into a hierarchical graph learning pipeline to produce expressive representations, the hierarchical graph learning pipeline including a control flow graph (CFG) embedding layer and a GoG embedding layer. The hierarchical graph learning pipeline also performs the following operations: in the CFG embedding layer, use one or more graph convolutional network (GCN) layers to propagate features for each basic block of the function to produce an aggregated vectorized representation. In the GoG embedding layer, receive the aggregated vectorized representation and merge call function information and called function information into each function to produce an expressive representation for each function. Embodiments can include a system that also performs the following operations: use the expressive representations to train an ML model, the expressive representations of two functions indicating that the two functions behave similarly, and determine missing symbols for a function based on the embedding of one of the functions. A software update of the stripped binary file can be performed partially based on the missing symbols found. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The above aspects and other aspects of the present invention can be best understood from the following detailed description when read in conjunction with the accompanying drawings. For the purpose of illustrating the present invention, the presently preferred embodiments are shown in the drawings. However, it should be understood that the present invention is not limited to the specific means disclosed. The following figures are included in the drawings: DETAILED DESCRIPTION
[0008] In mission-critical systems, embedded software is crucial in manipulating physical processes and performing tasks that may pose risks to human operators. Recently, the concept of the Internet of Things (IoT) has become increasingly popular, creating a market worth $19 trillion in 2025 and significantly increasing the number of connected devices to approximately 35 billion. However, while the IoT has driven technological growth, it has inadvertently exposed mission-critical systems, especially legacy systems, to new types of vulnerabilities. The number of IoT cyberattacks reported in 2019 increased by 300%, and the number of software vulnerabilities discovered rose from 1,600 to 100,000. For example, the Heartbleed flaw could leak up to 64K of memory, threatening the information security of individuals and organizations. Shellshock is a bash command-line interface shell flaw that has been around for 30 years and remains a threat to enterprises today. For these mission-critical systems, even an accidental interruption lasting only a few hours or minutes can result in millions of dollars in lost revenue. While it may seem simple to patch these vulnerabilities promptly and apply updates to the affected software, in reality, due to mission-criticality or the high cost of updates, mission-critical systems often use software that has been in use for decades. Over time, as technology advances, the software in these systems becomes obsolete, and the number of newly discovered vulnerabilities may increase. The original development environment, maintenance support, or source code may no longer exist or be available for legacy software.
[0009] Typically, to address vulnerabilities, vendors provide patches in the form of source code changes based on the current software version (e.g., version 0.9). However, the only available data in legacy systems is the binary files based on their source code (e.g., version 0.1). Such a version gap poses challenges in applying patches to legacy binaries, making the only solution for applying patches to legacy software direct binary analysis. Nowadays, REs must utilize software reverse engineering (SRE) tools (such as Ghidra, HexRays, and radare2) to first disassemble and decompile the binary files into a higher-level representation (e.g., C or C++). Typically, these tools employ debug information, strings, symbol tables, and the binary files to reconstruct function names and variable names, thus allowing REs to reconstruct the structure and functionality of the software without access to the source code. For REs, these symbols encode the context of the source code and provide valuable information for their tasks. Next, REs must understand the logic of the program to achieve their goal of patching vulnerable binaries. However, to optimize the footprint of the binary files in legacy mission-critical systems with limited memory, symbols are typically excluded to reduce the binary file size. Since recovering symbols from stripped binaries is not straightforward, most decompilers assign meaningless symbol names to the encoded elements. To understand the software semantics, REs must utilize their experience and expertise to digest the information and then interpret the semantics of each encoded element.
[0010] Software reverse engineering (SRE) aims to understand the behavior of a program without access to its source code and is commonly used in many applications such as malware detection, vulnerability discovery, and patching bugs in legacy software. One of the main tools used by reverse engineers (REs) to examine a program is a disassembler, which is a tool that converts a binary program into low-level assembly code. Examples of such tools include GNU Binutils objdump, IDA, Binary Ninja, and Hopper. However, even with the assistance of these tools, reasoning at the assembly level requires a significant amount of cognitive work from RE experts. More recently, REs have used decompilers such as Hex-Rays or Ghidra to reverse the compilation process by further converting the output of the disassembler into code similar to a high-level programming language such as C or C++, thereby alleviating the burden of understanding assembly code. These decompilers can use program analysis and heuristics to reconstruct the variables, types, functions, and control flow structures of the binary file. However, despite the fact that these decompilers generate a higher-level output for better code understanding, decompilation is not complete. This is because the compilation process discards source-level information and reduces its level of abstraction in exchange for a smaller footprint, faster execution time, or even security considerations. Accordingly, while source-level information including comments, variable names, function names, and idiomatic structures may be crucial for understanding a program, it is typically not available in the output of these decompilers.
[0011] Since developers usually deploy their software in binary form, analyzing software binaries to eliminate vulnerabilities is simple and effective for enhancing security. Typically, security experts can conduct their patching processes or vulnerability analysis by understanding the compiled source, function signatures, and variable information. However, after compilation, such information is usually deliberately stripped or obfuscated. In such cases, analyzing the stripped binaries becomes more challenging. Early recovery efforts for binaries focused on manual completion but suffered from inefficiency, high cost, and the error-prone nature of reverse engineering. The improvement in computing power has prompted researchers to use machine learning (ML) methodologies to address this challenge. As ML has made significant progress in its reasoning ability, researchers in this field first utilized general information from each function for in-depth binary code understanding and reconstructed higher-level source code information as an alternative to manual-based methods. The first approach used neural network-based models and graph-based models. It predicted function types to assist reverse engineers in understanding the binaries. Another approach used neural networks to predict function names. It aggregated relevant features of segments of binary vectors. Then, the binaries would consist of feature vectors ready for training. Subsequently, it utilized natural language processing (NLP) to analyze the connections between each part of the code to predict function names. On the other hand, this approach did not use neural networks. It combined decision tree-based classification algorithms and structured prediction with probabilistic graphical models. Then, function names were matched by analyzing symbol names, types, and locations.
[0012] The embodiments described herein present novel methods that include two major contributions to the field of software reverse engineering tools. These contributions will be discussed in detail in the following paragraphs.
[0013] The first objective of the disclosed embodiments models the stripped binary executable as a graph of graphs (GoG) representation. This step transforms the stripped binary into a two-layer graph, which is denoted as the graph of graphs representation. Starting from the stripped binary, a novel Ghidra Data Toolkit (GDT) tool is developed using the state-of-the-art decompiler Ghidra and its Ghidra headless analyzer. GDT is a Java-based metadata extraction script for implementing these tools. GDT incorporates Ghidra's API to export the decompiled P-code information for each function and their cross-call dependencies. For each decompiled binary function, GDT also decompiles it into basic blocks and their control flow dependencies, thereby extracting their summary attributes and corresponding P-code. Finally, GDT integrates all the information and forms the GoG representation for the given stripped binary.
[0014] Figure 1It is a flowchart showing the generation of a GoG from a compiled binary file according to an embodiment of the present disclosure. The binary code file 101 is decompiled into blocks, such as functions executed by the binary file 101. Each function can be characterized as FIGS. 110, 111, 112, 113, 114, 115 based on the actions of the function and the dependencies of the function. The control flow between functions can be used to determine the relationships between functions and is arranged in FIG. 120 of the graph.
[0015] In a second objective, embodiments of the present disclosure use hierarchical graph learning techniques to model the GoG representation in order to reconstruct function names for each binary function. Hierarchical graph learning techniques can be used to model complex GoG representations and perform the reconstruction of function names to improve state-of-the-art decompilers (e.g., Ghidra) and provide more source-level information to reverse engineers.
[0016] Figure 2 It is a block diagram showing the training of a hierarchical graph model according to an embodiment of the present disclosure. For model training, once the GDT converts all stripped binary files into GoGs (as depicted in Figure 1 ), the GoG 120 is input into a hierarchical graph learning pipeline 210, which includes a control flow graph (CFG) embedding layer 211 and a GoG embedding layer 213. The CFG embedding layer 211 includes one or more graph convolutional network (GCN) layers that propagate information (e.g., attributes + P code) for each basic block and utilize control flow-level information to generate an aggregated vectorized representation 212 (e.g., function embedding) for each function. The GoG embedding layer 213 also processes these function embeddings, including additional graph convolutional layers that propagate information from each function, incorporating information about calling or called functions into each function, thereby creating an expressive representation 214 for each function. Then, we use these resulting representations 214 of the functions to train all layers using a contrastive function similarity learning process 220. Specifically, this design allows the model to learn whether two binary functions behave similarly, meaning that even if they are built from different CPU architectures, their function embeddings should be as similar as possible.
[0017] Figure 3It is a process flow diagram of a method for reconstructing missing information from a compiled binary file according to aspects of an embodiment of the present disclosure. P-code and dependencies for each function in the binary file are extracted from the binary file 310. Basic blocks (e.g., if statements or iterative loops, etc.) of each function and the basic control flow dependencies of the blocks are extracted from each of the extracted functions 320. Each function can be characterized by a number and represented as a vectorized representation 330. By comparing the relationships between functions, a graph convolutional layer is also used to process each function to create an expressive representation of each function 340. These representations are used to train a graph learning pipeline to learn the similarity between individual functions 350. These similarities between functions can also be utilized to determine the names of functions using the trained model 360.
[0018] Embodiments of the present disclosure provide improvements to the manual analysis of binary files to allow REs to more efficiently and automatically extract missing information related to source code files from their corresponding compiled binary files.
[0019] Figure 4 An exemplary computing environment 400 is shown within which embodiments of the present invention can be implemented. Computers and computing environments such as computer system 410 and computing environment 400 are known to those skilled in the art and are therefore described briefly herein.
[0020] As Figure 4 shown, computer system 410 may include a communication mechanism such as system bus 421 or other communication mechanisms for transferring information within computer system 410. Computer system 410 also includes one or more processors 420 coupled to system bus 421 for processing information.
[0021] The processor 420 may include one or more central processing units (CPUs), a graphics processing unit (GPU), or any other processor known in the art. More generally, as used in the present invention, a processor is a device for executing machine-readable instructions stored on a computer-readable medium to perform tasks, and may include any one of hardware and firmware or a combination of hardware and firmware. The processor may also include a memory that stores machine-readable instructions executable to perform tasks. The processor acts on information by manipulating, analyzing, modifying, transforming, or sending information for use by an executable program or information device, and / or by routing the information to an output device. The processor may use or include the capabilities of, for example, a computer, a controller, or a microprocessor, and may be adjusted using executable instructions to perform specialized functions not performed by a general-purpose computer. The processor may be coupled (electrically and / or as including executable components) to any other processor to enable interaction and / or communication therebetween. A user interface processor or generator is a known element, including an electronic circuit system or software or a combination of both for generating display images or portions thereof. The user interface includes one or more display images to enable interaction between the user and the processor or other devices.
[0022] Continuing to refer to Figure 4 , the computer system 410 also includes a system memory 430 coupled to the system bus 421 for storing information and instructions to be executed by the processor 420. The system memory 430 may include computer-readable storage media in the form of volatile and / or non-volatile memory (such as read-only memory (ROM) 431 and / or random access memory (RAM) 432). The RAM 432 may include other dynamic storage devices (e.g., dynamic RAM, static RAM, and synchronous DRAM). The ROM 431 may include other static storage devices (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). Additionally, the system memory 430 may be used to store temporary variables or other intermediate information during the execution of instructions by the processor 420. A basic input / output system 433 (BIOS) containing basic routines may be stored in the ROM 431, and the basic routines assist in transferring information between elements within the computer system 410 (such as during startup). The RAM 432 may contain data and / or program modules that are immediately accessible and / or currently being operated on by the processor 420. The system memory 430 may additionally include, for example, an operating system 434, application programs 435, other program modules 436, and program data 437.
[0023] The computer system 410 also includes a disk controller 440 coupled to the system bus 421 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 441 and a removable media drive 442 (e.g., a floppy disk drive, a compact disc drive, a tape drive, and / or a solid state drive). Storage devices can be added to the computer system 410 using an appropriate device interface (e.g., Small Computer System Interface (SCSI), Integrated Device Electronics (IDE), Universal Serial Bus (USB), or FireWire).
[0024] The computer system 410 may also include a display controller 465 coupled to the system bus 421 to control a display or monitor 466 for displaying information to a computer user, such as a cathode ray tube (CRT) or a liquid crystal display (LCD). The computer system includes an input interface 460 and one or more input devices (such as a keyboard 462 and a pointing device 461) for interacting with a computer user and providing information to the processor 420. For example, the pointing device 461 can be a mouse, a light pen, a trackball, or a pointing stick for transmitting direction information and command selections to the processor 420 and for controlling the movement of a cursor on the display 466. The display 466 may provide a touch screen interface that allows input to supplement or replace the transmission of direction information and command selections of the pointing device 461. In some embodiments, an augmented reality device 467 that a user can wear may provide input / output functionality that allows the user to interact with both the physical world and the virtual world. The augmented reality device 467 communicates with the display controller 465 and the user input interface 460, thereby allowing the user to interact with virtual items generated by the display controller 465 in the augmented reality device 467. The user may also provide gestures that are detected by the augmented reality device 467 and sent as input signals to the user input interface 460.
[0025] The computer system 410 may perform some or all of the processing steps of the embodiments of the present invention in response to one or more sequences of one or more instructions included in a memory (such as the system memory 430) being executed by the processor 420. Such instructions may be read into the system memory 430 from another computer-readable medium (such as the magnetic hard disk 441 or the removable media drive 442). The magnetic hard disk 441 may contain one or more data repositories and data files used by the embodiments of the present invention. The contents of the data repositories and the data files may be encrypted to enhance security. The processor 420 may also be employed in a multiprocessing device to execute one or more sequences of instructions included in the system memory 430. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Accordingly, the embodiments are not limited to any specific combination of hardware circuitry and software.
[0026] As stated above, the computer system 410 may include at least one computer-readable medium or memory for storing instructions programmed according to the embodiments of the present invention and for containing the data structures, tables, records, or other data described herein. As used in the present invention, the term "computer-readable medium" refers to any medium that participates in providing instructions to the processor 420 for execution. Computer-readable media may take many forms, including but not limited to non-transitory media, non-volatile media, volatile media, and transmission media. Non-limiting examples of non-volatile media include optical discs, solid-state drives, magnetic disks, and magneto-optical discs, such as the magnetic hard disk 441 or the removable media drive 442. Non-limiting examples of volatile media include dynamic memory, such as the system memory 430. Non-limiting examples of transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up the system bus 421. Transmission media may also take the form of acoustic or light waves, such as those generated during radio-wave and infrared data communications.
[0027] The computing environment 400 may also include a computer system 410 that operates in a networked environment using a logical connection to one or more remote computers (such as the remote computing device 480). The remote computing device 480 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device, or other common network nodes, and typically includes many or all of the elements described above with respect to the computer system 410. When used in a networked environment, the computer system 410 may include a modem 472 for establishing communications over a network 471 (such as the Internet). The modem 472 may be connected to the system bus 421 via the user network interface 470 or via another suitable mechanism.
[0028] Network 471 can be any network or system commonly known in the art, including the Internet, intranet, local area network (LAN), wide area network (WAN), metropolitan area network (MAN), direct connection or series of connections, cellular telephone network, or any other network or medium capable of facilitating communication between computer system 410 and other computers (e.g., remote computing device 480). Network 471 can be wired, wireless, or a combination thereof. The wired connection can be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection commonly known in the art. The wireless connection can be implemented using Wi-Fi, WiMAX, and Bluetooth, infrared, cellular network, satellite, or any other wireless connection methodology commonly known in the art. Additionally, several networks can work alone or communicate with each other to facilitate communication in network 471.
[0029] An executable application as used in the present invention includes code or machine-readable instructions for adjusting a processor to implement a predetermined function (such as those functions of an operating system, context data acquisition system, or other information processing system) in response to, for example, a user command or input. An executable program is a fragment, subroutine, or other distinct section of code or a part of an executable application for performing one or more specific processes. These processes can include receiving input data and / or parameters, performing operations on the received input data and / or functions in response to the received input parameters, and providing the resulting output data and / or parameters.
[0030] A graphical user interface (GUI) as used in the present invention includes one or more display images generated by a display processor and implementing user interaction with the processor or other device, as well as associated data acquisition and processing functions. The GUI also includes an executable program or executable application. The executable program or executable application adjusts the display processor to generate signals representing the GUI display images. These signals are provided to a display device, which displays the images for the user to view. Under the control of the executable program or executable application, the processor manipulates the GUI display images in response to signals received from an input device. In this way, the user can use the input device to interact with the display images, thereby achieving user interaction with the processor or other device.
[0031] The functions and process steps in the present invention can be automatically executed or executed in whole or in part in response to a user command. Automatically executed activities (including steps) are performed in response to one or more executable instructions or device operations without direct initiation of the activity by the user.
[0032] The systems and processes in the figures are not exclusive. Other systems, processes, and menus can be derived in accordance with the principles of the present invention to achieve the same objectives. Although the present invention has been described with reference to specific embodiments, it will be understood that the embodiments and variations shown and described herein are for illustrative purposes only. Those skilled in the art can effect modifications to the current designs without departing from the scope of the present invention.
Claims
1. A method for recovering symbols from a stripped binary file, comprising: representing the stripped binary file as a graph of graphs (GoG) of multiple graphs; converting the multiple GoGs into multiple expressive representations for each function in the stripped binary file; using the expressive representations to train a machine learning (ML) model; and determining missing symbols for at least one of the functions based on the output of the ML model.
2. The method according to claim 1, wherein, representing the stripped binary file as multiple GoGs further comprises: using a decompiler to export P-code and cross-call dependencies for each function included in the stripped binary file; further decompiling each function to extract the basic blocks, summary attributes, and corresponding P-code of the function; representing each function as a graph of basic blocks; and creating a GoG for all functions in the stripped binary file, wherein nodes contain graphs of basic blocks of functions, and edges represent relationships based on the cross-call dependencies between functions.
3. The method according to claim 1, wherein, converting the multiple GoGs into multiple expressive representations includes: inputting the multiple GoGs into a hierarchical graph learning pipeline to generate the expressive representations, the hierarchical graph learning pipeline including a control flow graph (CFG) embedding layer and a GoG embedding layer.
4. The method according to claim 3, the hierarchical graph learning pipeline further performs the following steps: In the CFG embedding layer, using one or more graph convolutional network (GCN) layers to propagate features for each basic block of a function to generate an aggregated vectorized representation; and In the GoG embedding layer, receiving the aggregated vectorized representation, merging call function information and called function information into each function to generate the expressive representation of each function.
5. The method according to claim 4, further comprising: using the expressive representation of the function to train the ML model during a contrast function similarity learning process.
6. The method according to claim 5, further comprising: determining missing symbols for a function based on the embedding of one of the functions under the condition that the expressive representations of two functions indicate that the two functions behave similarly.
7. The method according to claim 6, wherein, each of the two functions is built from a different CPU architecture.
8. The method according to claim 2, wherein, the decompiler is a Ghidra decompiler.
9. The method according to claim 3, further comprising: implementing the hierarchical graph learning pipeline as a data toolkit for a Ghidra reverse engineering tool.
10. The method according to claim 9, wherein, the data toolkit incorporates the application programming interface (API) of the Ghidra reverse engineering tool.
11. The method according to claim 1, further comprising: performing a software update of the stripped binary file partially based on the discovered missing symbols.
12. A system for recovering symbols from a stripped binary file, comprising: A computer processor that communicates with a non-transitory computer memory storing machine-readable instructions that, when executed by the computer processor, cause the computer processor to: Represent the stripped binary file as a graph-of-graphs (GoG) representation of multiple graphs; Convert the multiple GoGs into multiple expressive representations for each function in the stripped binary file; Use the expressive representations to train a machine learning (ML) model; and Determine a missing symbol for at least one of the functions based on the output of the ML model.
13. The system according to claim 12, wherein, Representing the stripped binary file as multiple GoGs further includes: Using a decompiler to export P-code and cross-call dependencies for each function included in the stripped binary file; Further decompiling each function to extract the basic blocks, summary attributes, and corresponding P-code of the function; Representing each function as a graph of basic blocks; and Creating a GoG for all functions in the stripped binary file, where nodes contain graphs of basic blocks of functions and edges represent relationships based on the cross-call dependencies between functions.
14. The system according to claim 12, wherein, Converting the multiple GoGs into multiple expressive representations includes: inputting the multiple GoGs into a hierarchical graph learning pipeline to generate the expressive representations, the hierarchical graph learning pipeline including a control flow graph (CFG) embedding layer and a GoG embedding layer.
15. The method according to claim 14, the hierarchical graph learning pipeline further performs the following steps: In the CFG embedding layer, using one or more graph convolutional network (GCN) layers to propagate features for each basic block of a function to generate an aggregated vectorized representation; and In the GoG embedding layer, receiving the aggregated vectorized representation, merging call function information and called function information into each function to generate the expressive representation of each function.
16. The method according to claim 15, the non-transitory memory further includes instructions that, when executed by the computer processor, cause the computer processor to: use the expressive representation of the function to train the ML model during a contrastive function similarity learning process.
17. The system according to claim 16, the non-transitory memory further includes instructions that, when executed by the computer processor, cause the computer processor to: based on the embedding of one of the functions, determine a missing symbol for the function under the condition that the expressive representations of two functions indicate that the two functions behave similarly.
18. The system according to claim 13, wherein, The decompiler is a Ghidra decompiler.
19. The system according to claim 14, the non-transitory memory further includes instructions that, when executed by the computer processor, cause the computer processor to: implement the hierarchical graph learning pipeline as a data toolkit for a Ghidra reverse engineering tool.
20. The non-transitory memory of the system according to claim 12 further includes instructions that, when executed by the computer processor, cause the computer processor to perform a software update of the stripped binary file based in part on the missing symbols found.