An efficient binary in-memory vulnerability patch code identification method and system
Patent Information
- Application Number
- CN202611045712.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-09-09
- Filing Date
- 2026-07-14
- Publication Date
- 2026-10-09
AI Technical Summary
1)补丁代码模式单一,适用性不足:现有方案只能基于单一模式识别补丁代码,适用性不足
关键点1:本发明提出的三种内存漏洞补丁代码模式中“避开崩溃点”模式涵盖现有方案提出的所有模式,该模式在现实世界内存漏洞补丁中的占比超70%。本发明一并提出其他两种普遍采用的补丁代码模式及对应判定方案,极大提升了方案适用面,对比现有方案能适用于更多种类的目标程序及补丁。
Smart Images

Figure CN122884802A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the patch code identification branch of the program analysis field, and relates to an efficient binary-level memory vulnerability patch code identification method and system. Background Technology
[0002] Existing patch identification technologies mainly include the following methods: 1) Static Binary Difference: This method disassembles and normalizes the binary data before and after the patch, extracting structured representations (function / block fingerprints, control flow graphs, constant / string references, etc.). It then uses fingerprint matching and graph-based context propagation algorithms (such as BinDiff) to identify semantically or structurally similar functions / blocks to locate changes. The advantage of this method is that it works without running samples, is environment-independent, and can quickly provide function-level difference mappings, facilitating offline batch analysis and visual review. Its disadvantages are that it heavily relies on the accuracy of static analysis—compile optimizations (inlining, splitting, register rearrangement), symbol stripping, or obfuscation can disrupt the matching; furthermore, binary differences are not unique to patches, leading to a higher risk of false positives or false negatives, and it is difficult to achieve precision down to the instruction level.
[0003] 2) Dynamic Binary Differentiation: This method collects execution traces (instructions, system calls, memory accesses, etc.) during program runtime and compares the actual behavioral differences between two versions based on behavioral profiles or execution side effects (e.g., system call sequence alignment, function execution read / write differences). The advantages of this method are that it reflects actual semantics and runtime side effects, is more robust to static transformations, filters out insignificant static differences, and is suitable for locating behavioral changes caused by input or environment. Its disadvantages are that it depends on trigger inputs and the runtime environment; incomplete coverage can lead to missed detections; it has high acquisition overhead, is subject to noise and non-determinism, and places high demands on time series alignment, noise reduction, and difference attribution.
[0004] 3) AI-based binary difference: This method uses machine learning / deep learning to embed or classify binary data or its graph / trajectory features, and automatically determines the similarity and changes between functions or segments using similarity metrics. The advantages of this method are that it can automatically learn robust features on a large number of samples, improves the matching ability for complex optimizations / rearrangements, outperforms traditional methods on function-level tasks, and has strong scalability. Its disadvantages are that the granularity is usually limited to the function level, making it difficult to locate instruction-level patches; it relies on a large amount of labeled data and training resources, resulting in training bias, poor interpretability, and vulnerability to adversarial transformations or out-of-domain samples.
[0005] 4) Differential Guided Fuzzing Based on Tail Call Sequence (TCS): This method assumes that patches often cause changes in the tail function call sequence on the trigger path. First, function alignment is performed on the binary representations before and after the patch, and candidate points with TCS differences are statically identified. During runtime, the tail call sequence of each execution is collected, and differential comparison is performed based on the TCS differences. A difference metric is used to guide directional mutation to quickly generate triggering inputs (differential guided fuzzing). The advantages of this method are that it is more robust to optimization noise than pure static differencing; combining static call graphs with dynamic trajectories improves localization accuracy and accelerates PoC generation; and it is highly effective for patches that change call behavior. Its disadvantages are that it is only effective for patches that change the tail call sequence, easily missing other types of modifications; it depends on executable previous / next versions and sufficiently covered inputs; it requires additional tracking and parallel execution overhead, and its accuracy decreases when the call graph is incomplete.
[0006] Taint propagation based on execution traces is a dynamic information flow analysis method that uses the program's execution trace to track how "taints" (bits / bytes / values from untrusted inputs) propagate through registers, memory, and control flow over time. Implementation often involves binary rewriting, dynamic binary translation, or lightweight instrumentation to capture instruction semantics, and maintaining shadow memory / shadow registers at runtime to store taint metadata (bitmasks or tag sets). Key technologies include: fine-grained taint granularity (bytes / bits / symbols), explicit data flow and implicit control flow processing (control dependency propagation), taint merging and propagation rules (arithmetic, logical, load / store, function calls), and trace compression and incremental replay to reduce storage overhead. The typical steps of this technique are: 1) labeling taint sources; 2) synchronously acquiring the trace and maintaining the shadow state during execution; 3) updating taint tags according to the semantics of each instruction (including control dependency processing for conditional branches); 4) identifying and recording sensitive sinks where taints arrive; and 5) generating alerts, root cause backtracking, or minimization reports on the results. Common challenges include performance overhead and taint explosion, which are often mitigated by selective instrumentation, sparse sampling, taint merging strategies, and hardware-assisted mechanisms.
[0007] Binary Diffing is a static / semi-dynamic analysis technique used to compare the semantic similarity and change points of functions / basic blocks in two binary programs (or different versions of the same program). Its working principle can be summarized as follows: First, the binary code is disassembled and normalized (register renaming, elimination of irrelevant padding, constant / address normalization), extracting features of each function and basic block (instruction sequence summary, constant / string references, call edges, control flow structure, basic block metrics, etc.). Then, initial matching is performed based on local hashing and structural features, followed by refinement of the matching within the context using graph matching / iterative propagation algorithms (based on call graphs or control flow graphs). Finally, matching pairs, similarity scores, and difference summaries are output. Typical steps of this technique are: 1) Input and preprocessing: Load two binary codes, disassemble them, and normalize the instructions and addresses, handling potential differences caused by splitting / inlining. 2) Feature extraction: Calculate a summary for each function / basic block (instruction fingerprint, instruction category distribution, constant set, in / out edges, block size, cyclic complexity, etc.). 3) Initial Matching: High-confidence matching is performed using strong hashes, symbols / derivative names, and similarity features. 4) Structured Refinement: Based on neighborhood consistency in the call graph or control flow graph, iterative propagation or maximum matching strategies are used to expand and refine the initial matches, improving robustness to optimization or rearrangement. 5) Similarity Scoring and Difference Identification: Similarity scores are calculated on the matching results to identify inserted / deleted / rewritten functions and basic blocks, and to locate modification points (instruction level or control flow level). 6) Output Reports and Visualization: Matching pairs, difference patch summaries, similarity heatmaps, and interactive graphical views are generated to assist in patch location, vulnerability backtracking, or reverse engineering.
[0008] The closest solution to this invention is the one in the paper 1dFuzz: Reproduce 1-Day Vulnerabilities with Directed Differential Fuzzing. Both solutions are based on code pattern matching to locate patch code and use binary difference analysis technology to assist in patch code location.
[0009] The main drawbacks of existing patch code identification methods include: 1) Limited applicability due to single patch code pattern: Existing solutions can only identify patch codes based on a single pattern, resulting in insufficient applicability. For example, 1dFuzz identifies patch codes based on tail call sequence (TCS) difference patterns, but experiments show that only about 70% of patches conform to this pattern. Relying on a single pattern renders the solution ineffective for patches that do not use that pattern, reducing its applicability.
[0010] 2) Insufficient precision and accuracy: Existing solutions only support locating the function where the patch is located, and cannot locate the patch code at the instruction level, resulting in insufficient precision. Furthermore, existing solutions cannot guarantee accuracy because they often use call graphs to assist analysis, such as generating a sequence of called functions based on the call graph. However, the construction of the call graph itself is limited by the effects of disassembly under static analysis, indirect call resolution, and compiler optimizations, making it impossible to guarantee the accuracy of the construction.
[0011] 3) The analysis overhead is huge and it is difficult to meet the efficiency requirements: In order to compare the call graph / control flow graph, the existing solution repeatedly re-executes during the fuzzing process to complete the call graph and performs static control flow graph traversal on a large number of functions, which can easily lead to path explosion and frequent round-by-round back-up retesting overhead. The iterative cost of dynamic completion plus the cost of static analysis are huge in large projects or cross-version differences, making it difficult for this method to meet the efficiency and real-time requirements when performing real-time emergency reproduction or large-scale automated deployment. Summary of the Invention
[0012] The technical problem to be solved by this invention is as follows: 1) How to summarize rich patch code patterns and design matching schemes for different patch code patterns, thereby improving the applicability of patch identification schemes based on code pattern matching.
[0013] 2) How to improve the accuracy of patch code identification to the instruction level and improve the accuracy problem caused by static analysis.
[0014] 3) How to significantly reduce analysis overhead so that analysis can support large projects or scenarios with large cross-version differences, improve analysis efficiency and meet real-time requirements.
[0015] The technical solution adopted in this invention is as follows: An efficient method for identifying binary-level memory vulnerability patch code includes the following steps: Obtain the binary program of the vulnerable version and the binary program of the patched version, and obtain the inter-version difference data through static binary difference; By using the vulnerability proof-of-concept document, dynamic execution tracing was performed on the vulnerable version of the binary program and the patched version of the binary program, resulting in two execution tracks. Based on the inter-version difference data and two execution trajectories, obtain patch candidate instructions and crash instructions, and determine the code pattern of the target patch based on the patch candidate instructions, crash instructions and two execution trajectories; Obtain patch instructions based on the code pattern of the target patch.
[0016] Furthermore, obtaining inter-version differential data through static binary differential includes: using a static binary differential tool to extract matching and non-matching elements between the vulnerable version's binary program and the patched version's binary program, including matching functions, matching basic blocks, and non-matching instructions.
[0017] Furthermore, the step of obtaining patch candidate instructions and crash instructions based on inter-version difference data and two execution trajectories includes: Based on the matching elements, non-matching elements, and two execution tracks between the vulnerable version of the binary program and the patched version of the binary program, the non-matching instructions appearing in the execution tracks are selected as patch candidate instructions. Execute the vulnerable version of the binary program and receive the vulnerability proof-of-concept file to obtain the core dump file of the program crash. Obtain the crash instructions in the vulnerable version of the binary program from the call stack information at the time of the crash.
[0018] Furthermore, the code pattern for determining the target patch based on patch candidate instructions, crash instructions, and two execution trajectories includes: Comparing the two execution paths, if the execution path of the patched version of the binary program contains a crash instruction and the number of times the crash instruction appears is the same as the number of times it appears in the execution path of the vulnerable version of the binary program, then the code pattern of the target patch is "fixing the crash point". If no crash instruction appears in the execution trace of the patch version of the binary program and the patch candidate instruction contains a crash instruction, then the code mode of the target patch is "delete crash point"; If no crash instruction appears in the execution trace of the patched binary and no crash instruction is included in the patch candidate instructions, then the code mode of the target patch is "avoid crash point".
[0019] Furthermore, obtaining patch instructions based on the code pattern of the target patch includes: If the target patch's code pattern is "delete crash point", then the crash command is the patch command; If the code pattern of the target patch is "fix crash point", then identify the key fix variables in the two execution paths, perform taint analysis based on the key fix variables, and filter the patch instructions from the patch candidate instructions; If the target patch's code pattern is "avoiding crash points", then the critical fix branches within the two execution paths are identified, and taint analysis is performed based on the critical fix branches to filter out patch instructions from the patch candidate instructions.
[0020] Further, the critical repair branches are identified using the following steps; Locate the last occurrence of the crash instruction in the execution trajectory of the vulnerable version of the binary program as the crash point, obtain the call stack A at the crash point, and search for a call stack that is aligned with call stack A in the execution trajectory of the patched version of the binary program. If the search fails, search for a call stack that is aligned with the child call stack of call stack A, and try to align it starting from the largest child call stack, until a call stack in the patched version of the binary program that is aligned with the child call stack is found, which is called call stack B. Functions in call stack B are extracted from top to bottom. Basic block matching is performed within the two execution paths until the first mismatched basic block is found. The branch jump instruction of the last matching basic block is taken as the critical repair branch.
[0021] Furthermore, the taint analysis includes: receiving a patch candidate instruction, two execution trajectories, a critical repair variable or a critical repair branch, and using offline taint propagation based on the execution trajectory to mark the memory and registers modified by the patch candidate instruction as taints one by one. During the propagation process, the flag registers of the critical repair variable or critical repair branch instruction are checked to see if they are tainted. If so, the patch candidate instruction is added to the patch instruction set.
[0022] A highly efficient binary-level memory vulnerability patch code identification system, comprising: The static binary differential module is used to obtain the binary program of the vulnerable version and the binary program of the patched version, and obtain the difference data between the versions through static binary differential. The dynamic execution tracing module is used to dynamically trace the execution of the vulnerable version of the binary program and the patched version of the binary program using the vulnerability proof-of-concept file, and obtain two execution tracks. The pattern recognition module is used to obtain patch candidate instructions and crash instructions based on the version difference data and two execution trajectories, and to determine the code pattern of the target patch based on the patch candidate instructions, crash instructions and two execution trajectories. The patch instruction retrieval module is used to retrieve patch instructions based on the code pattern of the target patch.
[0023] Furthermore, the patch instruction acquisition module includes: The critical fix variable identification module is used to identify critical fix variables within two execution paths when the code mode of the target patch is "fix crash point"; The critical fix branch identification module is used to identify critical fix branches within two execution paths when the code mode of the target patch is "avoid crash points"; The taint analysis module is used to perform taint analysis based on critical repair variables or critical repair branches, and to filter and obtain patch instructions from the patch candidate instructions.
[0024] The beneficial effects of this invention are as follows: Key Point 1: Among the three memory vulnerability patching code modes proposed in this invention, the "avoiding crash points" mode covers all modes proposed in existing solutions, and this mode accounts for over 70% of real-world memory vulnerability patches. This invention also proposes two other commonly used patching code modes and corresponding judgment schemes, greatly improving the applicability of the solution and making it applicable to a wider range of target programs and patches compared to existing solutions.
[0025] Key Point 2: This invention innovatively employs taint analysis technology to filter mismatched instructions, improving the accuracy of patch code identification from the function level to the instruction level in existing solutions. This invention abandons static analysis techniques with obvious flaws, such as call graph construction, and eliminates patch code pattern matching that relies on call graphs. It rationally adopts static analysis techniques such as binary difference, effectively improving analysis accuracy.
[0026] Key Point 3: This invention extensively employs dynamic analysis techniques, such as execution tracing, binary cleanup, and taint analysis based on execution trajectory. It abandons the inefficient strategy of performing a full graph traversal of the control flow graph in existing solutions, shifting to a more efficient strategy of analyzing the PoC execution trajectory separately. This invention innovatively proposes a critical repair branch identification algorithm, transforming the alignment of the entire execution trajectory into call stack alignment plus alignment of basic blocks within functions. This improves analysis efficiency and addresses the shortcomings of insufficient patch code identification efficiency in large projects or scenarios with significant cross-version differences. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating the steps of an efficient binary-level memory vulnerability patch code identification method according to an embodiment of the present invention.
[0028] Figure 2 This is a module composition diagram of an efficient binary-level memory vulnerability patch code identification system according to an embodiment of the present invention.
[0029] Figure 3 This is a block diagram illustrating the principle of an efficient binary-level memory vulnerability patch code identification method according to an embodiment of the present invention.
[0030] Figure 4 This is an example diagram of a patch code pattern according to an embodiment of the present invention.
[0031] Figure 5 This is a flowchart illustrating an efficient binary-level memory vulnerability patch code identification method according to an embodiment of the present invention. Detailed Implementation
[0032] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0033] This invention discloses a highly efficient method for identifying memory vulnerability patch code in binary programs. The method proposes three common memory vulnerability patch code patterns and designs corresponding pattern matching schemes, overcoming the problem of low applicability of single patterns. Based on taint propagation technology, this method analyzes the correlation between mismatched instructions and vulnerability fixes, improving patch identification accuracy to the instruction level. Furthermore, this method utilizes dynamic analysis techniques such as execution tracing to assist in locating key fix variables and key fix branches, avoiding the use of static analysis techniques with obvious flaws or high overhead, thus improving the efficiency and real-time analysis capabilities of patch identification. This method features wide applicability, high accuracy, and fast analysis speed.
[0034] The innovation of this invention lies in: 1) Three common memory vulnerability patch code patterns are proposed, namely “avoiding crash points”, “fixing crash points”, and “removing crash points”. A patch code pattern determination scheme based on execution trajectory is proposed. Compared with existing schemes, the three patch patterns are more widely applicable and support the identification of more patch codes.
[0035] 2) A scheme is proposed to use taint propagation technology to filter patch code from mismatched instructions. Based on the basic logic that "different code that affects vulnerability repair is patch code", the identification accuracy of patch code is improved to the instruction level.
[0036] 3) An analysis bridge between patch code patterns and instruction-level patch code is proposed—critical repair variables and critical repair branches. A critical repair variable identification scheme and a critical repair branch identification algorithm are designed. The critical repair variable identification scheme is based on dynamic analysis technology of binary cleanup. The critical repair branch identification algorithm utilizes call stack alignment and basic block alignment technology based on execution trajectory. It makes extensive use of dynamic analysis technology and abandons static analysis technology with obvious defects or high overhead, thereby improving analysis efficiency.
[0037] The technical terms involved in this invention are explained as follows: Patch code pattern: A series of patch codes have similarities in code structure and code content. This similarity in structure or content is called a patch code pattern.
[0038] Code pattern matching: techniques or solutions for identifying the code pattern to which the target code belongs.
[0039] Call Stack: In a broad sense, the call stack is a memory structure used during program execution to manage function calls and returns. It stores function local variables, parameters, and return addresses in a last-in-first-out (LIFO) manner to support nested calls and program control flow restoration. In a narrow sense, the call stack is a data structure that stores function call relationships in a last-in-first-out (LIFO) manner.
[0040] Basic Block: A linear sequence of instructions in the program control flow that is only entered by the control flow at the entry point and only jumped or terminated at the exit point, with no interruption branches during execution.
[0041] Taint Propagation / Taint Analysis: A program analysis method that identifies potential security risks or sensitive information leakage paths or analyzes dependencies between data by marking "taints" on input data and tracing their computations and data flow during program execution. It is commonly used for vulnerability detection, malware analysis, and data flow tracing.
[0042] Execution tracing is a program analysis method that records the execution path information of a program during runtime, including instruction sequences, function calls, memory and register states, for debugging, performance analysis, security vulnerability location, and behavior reproduction. An execution trace is the output of an execution tracing tool.
[0043] Binary Diffing: A comparative analysis method for executable files or object code. It compares the differences in instructions, functions, or structures between different versions of binary code for patch identification, vulnerability location, and malware variant detection. Matching / non-matching functions / blocks / instructions are the output of binary diffing tools.
[0044] Binary Sanitizing: A security enhancement method that improves software security and reliability by performing runtime inspections on compiled binary programs to identify defects such as out-of-bounds access and uninitialized data.
[0045] A proof-of-concept (PoC) document is a minimal sample of code or data used to verify that a specific vulnerability can be triggered or exploited, demonstrating the existence and feasibility of the vulnerability.
[0046] Critical Variable: A variable in the program that is changed during patching to fix vulnerabilities.
[0047] Critical Branch: A branch jump instruction in the program. When patching, the vulnerability is fixed by adding this instruction or by causing the execution control flow to shift at this instruction.
[0048] Crash Site: The location in the execution path where a program crashes.
[0049] In one embodiment, the present invention provides an efficient binary-level memory vulnerability patch code identification method, such as... Figure 1 As shown, it includes the following steps: Obtain the binary program of the vulnerable version and the binary program of the patched version, and obtain the inter-version difference data through static binary difference; By using the vulnerability proof-of-concept document, dynamic execution tracing was performed on the vulnerable version of the binary program and the patched version of the binary program, resulting in two execution tracks. Based on the inter-version difference data and two execution trajectories, obtain patch candidate instructions and crash instructions, and determine the code pattern of the target patch based on the patch candidate instructions, crash instructions and two execution trajectories; Obtain patch instructions based on the code pattern of the target patch.
[0050] In one embodiment, obtaining inter-version difference data through static binary differential includes: using a static binary differential tool to extract matching and non-matching elements between the vulnerable version of the binary program and the patched version of the binary program, including matching functions, matching basic blocks, and non-matching instructions.
[0051] In one embodiment, obtaining patch candidate instructions and crash instructions based on inter-version difference data and two execution trajectories includes: Based on the matching elements, non-matching elements, and two execution tracks between the vulnerable version of the binary program and the patched version of the binary program, the non-matching instructions appearing in the execution tracks are selected as patch candidate instructions. Execute the vulnerable version of the binary program and receive the vulnerability proof-of-concept file to obtain the core dump file of the program crash. Obtain the crash instructions in the vulnerable version of the binary program from the call stack information at the time of the crash.
[0052] In one embodiment, the code pattern for determining the target patch based on patch candidate instructions, crash instructions, and two execution trajectories includes: Comparing the two execution paths, if the execution path of the patched version of the binary program contains a crash instruction and the number of times the crash instruction appears is the same as the number of times it appears in the execution path of the vulnerable version of the binary program, then the code pattern of the target patch is "fixing the crash point". If no crash instruction appears in the execution trace of the patch version of the binary program and the patch candidate instruction contains a crash instruction, then the code mode of the target patch is "delete crash point"; If no crash instruction appears in the execution trace of the patched binary and no crash instruction is included in the patch candidate instructions, then the code mode of the target patch is "avoid crash point".
[0053] In one embodiment, obtaining patch instructions based on the code pattern of the target patch includes: If the target patch's code pattern is "delete crash point", then the crash command is the patch command; If the code pattern of the target patch is "fix crash point", then identify the key fix variables in the two execution paths, perform taint analysis based on the key fix variables, and filter the patch instructions from the patch candidate instructions; If the target patch's code pattern is "avoiding crash points", then the critical fix branches within the two execution paths are identified, and taint analysis is performed based on the critical fix branches to filter out patch instructions from the patch candidate instructions.
[0054] In one embodiment, the critical repair branch is identified using the following steps; Locate the last occurrence of the crash instruction in the execution trajectory of the vulnerable version of the binary program as the crash point, obtain the call stack A at the crash point, and search for a call stack that is aligned with call stack A in the execution trajectory of the patched version of the binary program. If the search fails, search for a call stack that is aligned with the child call stack of call stack A, and try to align it starting from the largest child call stack, until a call stack in the patched version of the binary program that is aligned with the child call stack is found, which is called call stack B. Functions in call stack B are extracted from top to bottom. Basic block matching is performed within the two execution paths until the first mismatched basic block is found. The branch jump instruction of the last matching basic block is taken as the critical repair branch.
[0055] In one embodiment, the taint analysis includes: receiving a patch candidate instruction, two execution traces, a critical repair variable or a critical repair branch, using offline taint propagation based on the execution trace, marking the memory and registers modified by the patch candidate instruction as taints one by one, and checking whether the flag register of the critical repair variable or critical repair branch instruction is tainted during the propagation process, and if so, adding the patch candidate instruction to the patch instruction set.
[0056] In one embodiment, the present invention provides a highly efficient binary-level memory vulnerability patch code identification system, such as... Figure 2 As shown, it includes: The static binary differential module is used to obtain the binary program of the vulnerable version and the binary program of the patched version, and obtain the difference data between the versions through static binary differential. The dynamic execution tracing module is used to dynamically trace the execution of the vulnerable version of the binary program and the patched version of the binary program using the vulnerability proof-of-concept file, and obtain two execution tracks. The pattern recognition module is used to obtain patch candidate instructions and crash instructions based on the version difference data and two execution trajectories, and to determine the code pattern of the target patch based on the patch candidate instructions, crash instructions and two execution trajectories. The patch instruction retrieval module is used to retrieve patch instructions based on the code pattern of the target patch.
[0057] The patch instruction acquisition module includes: The critical fix variable identification module is used to identify critical fix variables within two execution paths when the code mode of the target patch is "fix crash point"; The critical fix branch identification module is used to identify critical fix branches within two execution paths when the code mode of the target patch is "avoid crash points"; The taint analysis module is used to perform taint analysis based on critical repair variables or critical repair branches, and to filter and obtain patch instructions from the patch candidate instructions.
[0058] In one embodiment, the principle block diagram of the software implementation of the present invention is as follows: Figure 3 As shown. This system receives three input data items: the vulnerable version of the binary program to be analyzed ( Figure 3 Vul icon), and a patched version of the binary program to be analyzed. Figure 3 (Pat icon), vulnerability proof-of-concept document () Figure 3 The code (PoC icon) is processed by six modules: static binary difference module, dynamic execution tracing module, pattern recognition module, critical repair variable recognition module, critical repair branch recognition module, and taint analysis module, and finally outputs patch code instructions.
[0059] The specific analysis process within each module is as follows: Static Binary Differentiation Module: This module receives two binary programs, one with vulnerabilities and the other with patches. It uses static binary differentiation technology (implemented using difference tools such as Bindiff) to extract matching / non-matching elements between the two programs, including matching functions, matching basic blocks, and non-matching instructions, which are then used as analysis materials by subsequent modules.
[0060] Execution Tracing Module: This module receives two binary programs, one with vulnerabilities and one with patches, along with the vulnerability PoC file. Based on dynamic execution tracing technology (implemented using system-wide recording / replay tools such as Panda), it records the execution trajectory of the binary program as it receives the PoC input. The output consists of two execution trajectories, which are then used as analysis material by subsequent modules.
[0061] Pattern Recognition Module: This module receives matching / non-matching elements between two programs and two execution trajectories of the Proof-of-Concept (PoC). First, this module filters out non-matching instructions appearing in the execution trajectories as patch candidates. For example, instructions such as x, y, z, etc., removed from the vulnerable version, and instructions such as m, n, etc., added to the patch version, appear in both the vulnerable version's trajectory and the patch version's trajectory, thus these are all patch candidates. Next, this module executes the vulnerable version program and receives the PoC input, obtaining the core dump file of the program crash. It then extracts the crash instructions from the call stack information at the time of the crash. Finally, based on the crash instructions, patch candidate instructions, and the two execution trajectories, this module determines the code pattern of the target patch and outputs the code pattern. When the patch code pattern is "delete crash point," the crash instructions are the patch instructions and are directly output. The method for determining the patch code pattern is detailed in the "Patch Code Pattern Determination" section.
[0062] Critical Remediation Variable Identification Module: This module functions when the patch code mode is "Fix Crash Point". It utilizes binary sanitization techniques (implemented using binary sanitizers such as Valgrind memcheck) to identify the access length, access address, and allocation length variables associated with vulnerability-related erroneous access. These three variables are called critical remediation variables. First, the module monitors the execution of the PoC on the vulnerable version, obtaining a list of erroneous access information provided by the binary sanitizer. Next, it searches the list for erroneous access information matching the crash command, extracting the access length or access address variable and allocation length variable from the erroneous access information. Finally, the module compares whether the critical remediation variables have changed within the two execution paths, outputting any changed critical remediation variables to the taint analysis module.
[0063] Critical Fix Branch Identification Module: This module functions when the patch code mode is "Avoid Crash Points". Based on the critical fix branch identification algorithm (see the "Critical Fix Branch Identification Algorithm" section), this module identifies critical fix branch instructions and outputs the results to the taint analysis module.
[0064] The taint analysis module receives patch candidate instructions, two execution paths, and critical repair variables or critical repair branches. First, using offline taint propagation based on execution paths, this module marks the memory / registers modified by each patch candidate instruction as tainted. During propagation, it checks whether the flag registers of critical repair variables or critical repair branch instructions are tainted; if so, it adds the patch candidate instruction to the patch instruction set. Next, the module outputs the patch instruction set, which contains all patch instructions.
[0065] Patch code mode determination: Figure 4 Three example diagrams of patch code patterns are shown: (a) and (b) represent "avoiding the crash point," (c) represents "fixing the crash point," and (d) represents "removing the crash point." Comparing the two execution paths, if the execution path of the patched binary contains a crash instruction, and the number of times it appears is the same as the number of times the crash instruction appears in the execution path of the vulnerable binary, then the patch code pattern is "fixing the crash point." If the execution path of the patched binary does not contain a crash instruction, then if the patch candidate instructions include a crash instruction, then the patch code pattern is "removing the crash point"; if the patch candidate instructions do not include a crash instruction, then the patch code pattern is "avoiding the crash point."
[0066] Critical Fix Branch Identification Algorithm: The algorithm receives two tracks and a crash instruction, and outputs a critical fix branch after two processes: call stack alignment and basic block alignment. First, the algorithm locates the last occurrence of the crash instruction in the vulnerable version track, as the crash point. Next, the algorithm obtains the call stack A at the crash point and searches for a call stack aligned with call stack A in the patch version track (two call stacks are aligned if and only if each level of function within the two call stacks is a matching function). If the search fails, it attempts to search for a call stack aligned with a child call stack of call stack A, starting from the largest child call stack, until a call stack in the patch version aligned with child call stack A' is found, called call stack B. This process is called call stack alignment. Subsequently, the algorithm extracts functions from call stack B from top to bottom, performs basic block matching within the functions on the two tracks, until the first mismatched basic block is found. The branch jump instruction (e.g., a jmp instruction) of the last matching basic block is taken as the critical fix branch output. This process is called basic block alignment.
[0067] In one embodiment, the efficient binary-level memory vulnerability patch code identification method of the present invention has the following software implementation logic flow: Figure 5 As shown, the main steps include the following: Step 1: Receive the vulnerable version of the binary program and the patched version of the binary program, and extract the inter-version difference data based on the static binary difference tool; Step 2: Receive the vulnerable version of the binary program and the patched version of the binary program, as well as the vulnerability proof-of-concept (PoC) file. Record the execution trajectory of the binary program receiving the PoC input using an execution tracing tool, resulting in two execution trajectories. Step 3: Receive the inter-version difference data and two execution traces, use the dump file to identify crash instructions, directly filter the patch candidate instructions, and determine the patch code pattern used by the target patch based on the patch code pattern determination scheme. Step 4: Determine if the patch code mode is "Delete crash point". If so, proceed to step 9; otherwise, proceed to step 5. Step 5: Determine if the patch code mode is "Fix crash point". If so, proceed to Step 7; otherwise, if it is "Avoid crash point", proceed to Step 6. Step 6: Receive crash instructions, inter-version differential data, and two execution paths. Use the critical fix branch identification algorithm to identify critical fix branch instructions and proceed to Step 8. Step 7: Receive the vulnerable version of the binary program and the patched version of the binary program, as well as the proof-of-concept (PoC) file. Use the binary cleanup tool to obtain a list of fault access information and extract key fix variables. Step 8: Receive critical fix variables / critical fix branches, two execution paths, and patch candidate instructions. Utilize offline taint propagation based on execution paths to filter patch instructions from among the patch candidate instructions. Step 9: When the patch code mode is "Delete crash point", the crash command is the patch command, and the patch command is output.
[0068] This invention was experimentally verified on a dataset consisting of 13 real CVE vulnerabilities, involving 10 commonly used network service programs. The experiment verified whether this invention could accurately identify security patch instructions; Table 1 shows the experimental results. The results show that this invention successfully identified security patches for 12 CVE vulnerabilities, with only one incorrectly identified instruction, achieving an instruction-level accuracy of 99.5%.
[0069] Table 1 Other embodiments of the present invention: 1) The tools used to implement the various technologies employed in this invention are interchangeable. For example, the binary differential tool does not necessarily have to be Bindiff, the binary cleanup tool does not necessarily have to be Valgrind memcheck, and the execution tracing tool does not necessarily have to be Panda, etc.
[0070] 2) For the critical repair variable identification module, binary sanitization is not always necessary to identify critical repair variables. For example, allocation / access pairing analysis based on execution traces can also identify critical repair variables.
[0071] 3) For the critical repair branch identification module, identifying critical repair branches does not necessarily require the use of a critical repair branch identification algorithm. For example, basic block alignment based on two complete execution trajectories can also identify critical repair branches.
[0072] It should be understood that the methods and systems disclosed in the above embodiments of the present invention can be implemented in other ways. For example, the above module division can be implemented in other ways, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Each step and module in the present invention can be implemented in the form of software functional units and can be stored in a computer-readable storage medium, including several instructions to cause a computer device to execute some or all of the steps of the method described in the present invention. For example, one embodiment of the present invention provides a computer device (computer, server, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for executing each step of the method of the present invention. For example, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk, etc.) storing a computer program, which, when executed by a computer, implements each step of the method of the present invention. For example, another embodiment of the present invention provides a computer program product including a computer program, which, when executed by a computer, implements the steps of the method of the present invention.
[0073] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and to implement it accordingly. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification; the scope of protection of the present invention is defined by the claims.
Claims
1. A highly efficient method for identifying binary-level memory vulnerability patch code, characterized in that, Includes the following steps: Obtain the binary program of the vulnerable version and the binary program of the patched version, and obtain the inter-version difference data through static binary difference; By using the vulnerability proof-of-concept document, dynamic execution tracing was performed on the vulnerable version of the binary program and the patched version of the binary program, resulting in two execution tracks. Based on the inter-version difference data and two execution trajectories, obtain patch candidate instructions and crash instructions, and determine the code pattern of the target patch based on the patch candidate instructions, crash instructions and two execution trajectories; Obtain patch instructions based on the code pattern of the target patch.
2. The method according to claim 1, characterized in that, The process of obtaining inter-version difference data through static binary differential includes: using a static binary differential tool to extract matching and non-matching elements between the vulnerable version's binary program and the patched version's binary program, including matching functions, matching basic blocks, and non-matching instructions.
3. The method according to claim 2, characterized in that, The process of obtaining patch candidate instructions and crash instructions based on inter-version difference data and two execution trajectories includes: Based on the matching elements, non-matching elements, and two execution traces between the vulnerable version of the binary program and the patched version of the binary program, the non-matching instructions appearing in the execution traces are selected as patch candidate instructions. Execute the vulnerable version of the binary program and receive the vulnerability proof-of-concept file to obtain the core dump file of the program crash. Obtain the crash instructions in the vulnerable version of the binary program from the call stack information at the time of the crash.
4. The method according to claim 1, characterized in that, The code pattern for determining the target patch based on patch candidate instructions, crash instructions, and two execution trajectories includes: Comparing the two execution paths, if the execution path of the patched version of the binary program contains a crash instruction and the number of times the crash instruction appears is the same as the number of times it appears in the execution path of the vulnerable version of the binary program, then the code pattern of the target patch is "fix the crash point". If no crash instruction appears in the execution trace of the patch version of the binary program and the patch candidate instruction contains a crash instruction, then the code mode of the target patch is "delete crash point"; If no crash instruction appears in the execution trace of the patched binary and no crash instruction is included in the patch candidate instructions, then the code mode of the target patch is "avoid crash point".
5. The method according to claim 4, characterized in that, The process of obtaining patch instructions based on the code pattern of the target patch includes: If the target patch's code pattern is "delete crash point", then the crash command is the patch command; If the code pattern of the target patch is "fix crash point", then identify the key fix variables in the two execution paths, perform taint analysis based on the key fix variables, and filter the patch instructions from the patch candidate instructions; If the target patch's code pattern is "avoiding crash points", then the critical fix branches within the two execution paths are identified, and taint analysis is performed based on the critical fix branches to filter and obtain patch instructions from the patch candidate instructions.
6. The method according to claim 5, characterized in that, The following steps are used to identify the critical repair branches; Locate the last occurrence of the crash instruction in the execution trajectory of the vulnerable version of the binary program as the crash point, obtain the call stack A at the crash point, and search for a call stack that is aligned with call stack A in the execution trajectory of the patched version of the binary program. If the search fails, search for a call stack that is aligned with the child call stack of call stack A, and try to align it starting from the largest child call stack, until a call stack in the patched version of the binary program that is aligned with the child call stack is found, which is called call stack B. Functions in call stack B are extracted from top to bottom. Basic block matching is performed within the two execution paths until the first mismatched basic block is found. The branch jump instruction of the last matching basic block is taken as the critical repair branch.
7. The method according to claim 5, characterized in that, The taint analysis includes: receiving a patch candidate instruction, two execution trajectories, and a critical repair variable or critical repair branch; using offline taint propagation based on the execution trajectory to mark the memory and registers modified by the patch candidate instruction as taints one by one; and checking whether the flag register of the critical repair variable or critical repair branch instruction is tainted during the propagation process. If so, the patch candidate instruction is added to the patch instruction set.
8. A highly efficient binary-level memory vulnerability patch code identification system, characterized in that, include: The static binary differential module is used to obtain the binary program of the vulnerable version and the binary program of the patched version, and obtain the difference data between the versions through static binary differential. The dynamic execution tracing module is used to dynamically trace the execution of the vulnerable version of the binary program and the patched version of the binary program using the vulnerability proof-of-concept file, and obtain two execution tracks. The pattern recognition module is used to obtain patch candidate instructions and crash instructions based on the version difference data and two execution trajectories, and to determine the code pattern of the target patch based on the patch candidate instructions, crash instructions and two execution trajectories. The patch instruction retrieval module is used to retrieve patch instructions based on the code pattern of the target patch.
9. The system according to claim 8, characterized in that, The patch instruction acquisition module includes: The critical fix variable identification module is used to identify critical fix variables within two execution paths when the code mode of the target patch is "fix crash point"; The critical fix branch identification module is used to identify critical fix branches within two execution paths when the code mode of the target patch is "avoid crash points"; The taint analysis module is used to perform taint analysis based on critical repair variables or critical repair branches, and to filter and obtain patch instructions from the patch candidate instructions.
10. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 7.