Static memory vulnerability mining method based on hook function

Through the static memory vulnerability mining method based on hook function, using DSL to identify the hook function and configure it as a stain, the problem that existing tools are difficult to detect high-risk vulnerabilities in hook function is solved, and automatic modeling and full care of hook function is realized, which significantly improves the security of the system.

CN120124067APending Publication Date: 2025-06-10JIANGSU HOPERUN SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178065.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing static vulnerability mining tools are difficult to effectively detect high-risk vulnerabilities in hook functions, which makes it difficult to reduce system security risks.

Method used

Through the static memory vulnerability mining method based on hook functions, the hook function is identified using DSL and configured as a taint into the static code mining tool ClangStatic Analyzer (CSA) for static analysis, identifying and fixing problems such as null pointers, array out-of-bounds, and memory access out-of-bounds.

Benefits of technology

Automatic modeling and full care of hook functions are realized, the number of alarms is reduced, the effectiveness of alarms is improved, the security of the system is significantly improved, and the memory problems are reduced by 20%-30%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124067A_ABST
    Figure CN120124067A_ABST
Patent Text Reader

Abstract

The invention relates to a hook function-based static memory vulnerability mining method, which comprises the following steps of S1, performing operations such as lexical analysis and the like on a source code through a Clang compiler LLVM, S2, using the LLVM as a generator of a back-end code, and generating an IR, S3, generating a control flow graph CFG and a program dependency graph PDG based on the IR, S4, generating a code attribute graph CPG based on AST, CFG and PDG, and S5, selecting OverflowDB or Neo4J for a graph database, the method comprises the steps of S1, analyzing code features of a hook function, S6, searching for a proper DSL attribute relation according to the analyzed code features to operate a graph database, S7, using Function Access to develop a queried Checker according to the code features of the hook function, S8, finding a static scanning tool supporting a stain scanning model, S9, generating a configuration format matched with a current tool, and S9, executing the step S9. And S10, starting mining according to the hook configuration, the stain model of the static tool and the stain Checker supported by the stain. The problems of high false alarm and low income of current memory problems are solved, and the system integrity risk is reduced and the system robustness is enhanced through deep mining.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vulnerability mining method, in particular to a static memory vulnerability mining method based on hook functions, and belongs to the technical field of static vulnerability mining. Background Art

[0002] Google engineers have counted more than 900 security bugs fixed by Google since 2015 to 2020, and stated that 70% of all serious security vulnerabilities in the current Chrome codebase are security vulnerabilities in memory management. Due to the incorrect management of memory pointers, attackers are given the opportunity to attack Chrome internal components. Immediately afterwards, Microsoft engineers also publicly stated that approximately 70% of Microsoft's security updates in the past 12 years have been to address memory security vulnerabilities. The data on security vulnerabilities from Google and Microsoft are surprisingly similar. The reason for these memory security vulnerabilities is precisely due to problems with programming languages. The main programming languages used by Google and Microsoft are both C and C++.

[0003] The underlying core code of embedded software is mainly in C / CPP language, such as the andriod kernel, HAL layer, etc. In the process of static vulnerability mining for C / CPP languages, null pointer dereference, out-of-bounds memory access, array out-of-bounds, memory leakage, etc. are the most common problems affecting system stability. However, these problems are very difficult to detect due to their particularity. The reason is that such problems are usually context-related and the function call hierarchy is extremely deep. If static tools are used to report such problems layer by layer, the number of alarms will be extremely large, resulting in real alarms being drowned out and the cost of alarm cleaning being huge. If such problems are not reported at all, there are significant security risks in the system and it is extremely easy to be breached from the outside. Currently, some excellent static mining tools in the industry, such as Coverity, CppCheck, PclintPlus, etc., will only report alarms when such problems are clearly present in context inference. However, when the source of such problems is user input and multi-threaded operations, the tools cannot recognize such scenarios and cannot determine the input values, so they cannot deduce that such functions are risky, resulting in such risks not being detected.

[0004] The layered architecture divides software modules into multiple layers in a horizontal splitting manner. A system consists of multiple layers, and each layer consists of multiple modules. At the same time, each layer has its own independent responsibilities, and multiple layers cooperate to provide complete functions. After layering, it is very convenient to extract some modules and make them into an independent system, which has the characteristics of high cohesion, low coupling, easy reuse, and high scalability. The essence of layered design is actually to simplify complex problems. Based on the single responsibility principle, each layer of code performs its own duties, thereby improving the maintainability and scalability of the code.

[0005] Hook functions are important products of a layered architecture. Hooks are usually used as interfaces between components and modules and also as isolation zones in the layered architecture model. Due to their high flexibility, existing static mining tools cannot solve the problem of data disconnection during the analysis of hook functions. Even industry leaders like Coverity have no feasible solutions, which leads to high-risk interfaces like hook functions being analyzed as ordinary functions. Typically, interfaces such as the android underlying hardware interfaces probe, init, and remove also lack effective safeguarding measures. If manual modeling is used to identify such interfaces, the workload is huge (for a certain component in the linux kernel open source repository, the code volume is 6000 kloc and there are more than 2000 hook functions), and it is also easy to have omissions and errors due to function name changes. The code volume of the Linux kernel reaches 27852 kloc, not counting the code volume of the application layer. The workload of identifying hook functions is extremely large.

[0006] The static high-risk vulnerability mining method for hook functions takes a different approach to the breakpoints in the data flow analysis of the above-mentioned static tools to solve this technical problem. After using this technical solution, automatic modeling of hook functions is achieved, so that they can be fully identified and used as taints (input and output parameters) for static vulnerability mining of common injection vulnerabilities such as null pointer dereference, out-of-bounds memory access, and out-of-bounds array access. It not only solves the problem that the disconnection of hook functions cannot be deeply detected but also realizes full-scale safeguarding of high-risk interfaces, thus greatly reducing the security risks of the overall system.

[0007] By identifying and modeling hook functions and focusing on safeguarding and inspecting these interfaces according to the function layering principle, most external injection risk problems can be solved. The security level of the system is greatly improved. It is mainly used to solve the technical problems of low mining efficiency and high false positives in the detection of high-risk vulnerabilities by existing static vulnerability mining technologies. After a certain product uses this technology, the proportion of high-risk problems that are manually inspected and then mined after being enhanced by the tool has increased from 25% to 85%. The false positive rate depends on the false positive rate of the mining tool itself, and the overall number of alarms is within a controllable range. High-risk problems are closed in advance, improving the robustness of the system and the competitiveness of the product. Therefore, there is an urgent need for a new solution to solve this technical problem. Summary of the Invention

[0008] The present invention specifically aims at the technical problems existing in the prior art and provides a static memory vulnerability mining method based on hook functions. This solution involves mining methods for common high-risk vulnerabilities such as null pointers, out-of-bounds arrays, and out-of-bounds memory access. By deeply mining hook functions, it solves the problems of high false positives and low efficiency commonly existing in the business, reduces the overall risk of the system, and enhances the robustness of the system.

[0009] The present invention aims to provide a method for detecting common memory security vulnerabilities in C / C++ language, such as null pointer dereference, array out-of-bounds access, and memory access out-of-bounds. Currently, there are many detection checkers in various static checking tools to solve such problems. However, due to the particularity of injected vulnerabilities and tool limitations, the problems of excessive warning volume and insufficient effectiveness cannot be solved, resulting in ineffective protection of such problems. Based on the idea of a hierarchical architecture, hook functions identify all hook functions in the system through DSL, use these function names as taint configuration inputs, and perform static analysis using the static code detection tool Clang Static Analyzer (CSA), thereby identifying null pointer dereference, array out-of-bounds access, and memory detection problems and fixing them in a timely manner, greatly reducing the number of warnings, improving the effectiveness of warnings, and significantly enhancing the security of the system.

[0010] To achieve the above object, the technical solution of the present invention is as follows: a static memory vulnerability detection method based on hook functions, the method comprising the following steps:

[0011] S1: The source code is subjected to lexical analysis, syntax analysis, semantic analysis, etc. by the Clang compiler LLVM (Low Level Virtual Machine), and the analysis results are converted into an abstract syntax tree AST (Abstract Syntax Tree); this step is mainly used for users to customize static vulnerability detection checkers according to code characteristics. If a mature existing taint checker is used, this step can be directly skipped to S2;

[0012] If a mature existing taint checker is used, this step can be directly skipped to S2;

[0013] Assume that the file name of the code to be tested is: test.cc, and ensure that it can be compiled successfully. Then, the AST information can be printed through the following command:

[0014] clang -Xclang -ast-dump -fsyntax-only test.cc

[0015] S2: After generating the AST from the source code, use LLVM as the backend code generator to generate IR (Intermediate Representation). This step is the basis of static analysis; users do not need to understand the format of IR in detail, and only need to convert it according to the command.

[0016] The command to generate the IR middleware for the above test file is as follows (ll file):

[0017] clang test.cc -O0 -emit-llvm -o -S test.ll

[0018] If the target format is a bc file, the command is as follows:

[0019] clang test.cc -O0 -emit-llvm -o test.bc

[0020] S3: Generate a Control Flow Graph (CFG) and a Program Dependency Graph (PDG) based on the IR;

[0021] S4: Generate a Code Property Graph (CPG) jointly based on the AST, CFG, and PDG;

[0022] S5: As a code property graph, the CPG must find a graph database as a carrier, just like the relationship between the data we commonly use and the SQL database. The graph database can choose to use OverflowDB or Neo4J;

[0023] Here, it should be noted that for the two processes of S3 and S4, you don't need to focus on the details. Just by importing test.cc into Joern, Joern can parse the source code into a CPG and store it in OverflowDB.

[0024] S6: Use a Domain Specific Language (DSL) to operate on the graph database according to the analysis features to find appropriate property relationships. Here, the DSL tool uses the CodeNavi language for analysis. By analyzing the node and node attribute information in CodeNavi, it is known that the attribute of FunctionAccess is convenient for obtaining function pointers; create a file dsl_hook.kirin and use the following query statement:

[0025]

[0026] The DSL executes the following command:

[0027] Dsl --scope=compile_command.json --rule_file=dsl_hook.kirin --result_path=. /

[0028] Here, compile_command.json is the database file generated by the test code (if there is only one file, it is test.cc), rule_file is the rule file of the DSL, and result_path is the path where the query results are saved.

[0029] S7: Develop a Checker using the DSL language based on the code characteristics of the hook function to generate a hook function configuration list. These function pointers will serve as the taint entry for static mining. In this step, it is only necessary to confirm whether the generated configuration file conforms to the format required by the static vulnerability mining tool.

[0030] S8: Configure the input and output parameters of the hook as taints, and then use this configuration as the input of the static vulnerability scanning tool to start vulnerability mining. Suppose the CSA static scanning tool is used here, then the possible command is as follows:

[0031] clang--analyze--analyzer-output html-o <output-dir>-Xclang-analyzer-checker=hook_function_checker

[0032] The analyzer-checker indicates which Checker of CSA is used for vulnerability scanning, and the output-dir specifies the path to save the final alerts of the scan.

[0033] Specifically as follows:

[0034] S1 is specifically as follows: AST is a tree-like data structure, where each node represents a syntactic structure in the source code, including declarations, statements, expressions, etc. By traversing the AST, the structure of the code can be analyzed and the required key information can be extracted. Assuming the file name is: test.cc, then the AST information can be printed through the following command:

[0035] clang -Xclang -ast-dump -fsyntax-only test.cc.

[0036] S2 is specifically as follows:

[0037] IR is closer to machine code, is language-independent, compressed and concise, contains control flow information, and is the basis for static analysis. The command to generate the IR middleware is as follows:

[0038] clang test.cc -O0 -emit-llvm -o -S test.ll.

[0039] S3 is specifically as follows:

[0040] The control flow graph is a graphical representation of the control flow relationship between basic blocks (Basic Blocks) in a program. The opt tool of LLVM can generate a DOT file of the CFG, and then use Graphviz for visualization. The command to generate the control flow graph of the file is as follows:

[0041] opt -dot-cfg test.ll

[0042] LLVM provides llvm::DependenceGraph and llvm::DependenceAnalysis to analyze data and control dependencies. Use the above interfaces to customize an LLVM Pass to analyze the dependencies and generate a PDG. This Pass is saved as MyPDGPass.cpp. Compile using LLVM:

[0043] clang++ -c MyPDGPass.cpp $(llvm-config --cxxflags --ldflags --system-libs --libs core analysis) -o MyPDGPass.o

[0044] Load and run the Pass:

[0045] opt -load . / MyPDGPass.so -mypdg test.ll -o / dev / null to generate a DOT file and visualize it.

[0046] S4 is as follows:

[0047] Use LLVM to generate AST, CFG, and PDG, then write a custom tool to integrate them into CPG, and use the Python library (NetworkX) to represent and manipulate CPG.

[0048] The extraction steps are as follows:

[0049] Step 1: Extract AST nodes and edges. Traverse the AST, extract all nodes (including variables, expressions, statements, etc.) and edges (including parent - child relationships, sibling relationships), and assign a unique identifier (ID) to each AST node.

[0050] Step 2: Extract CFG nodes and edges. Traverse the CFG, extract all basic blocks and jump relationships, and associate CFG nodes with AST nodes (including statements in basic blocks corresponding to statement nodes in the AST).

[0051] Step 3: Extract PDG nodes and edges. Traverse the PDG, extract data - dependency and control - dependency relationships, and associate PDG edges with AST nodes (including data - dependency edges connecting two variable nodes).

[0052] Step 4: Integrate into CPG. Integrate the nodes and edges of AST, CFG, and PDG into a graph, ensuring the uniqueness of nodes and edges.

[0053] S5 is as follows:

[0054] Step 1: Import the code. Import your code file (test.cc) in Joern: importCode(". / test.cc")

[0055] Step 2: Generate CPG. Joern will automatically parse the code into CPG and store it in OverflowDB.

[0056] Step 3: Use OverflowDB to operate on the CPG. OverflowDB provides powerful graph query languages (including Gremlin and Cypher) that can be used to query and analyze the CPG.

[0057] S6 is as follows:

[0058] FunctionAccess is a node type in CodeNavi that represents function access and is typically used to describe function calls. To extract the hook function, use the following description statement:

[0059]

[0060] S7 is as follows:

[0061] The entry point of the tainted function is usually defined in the form of a configuration file. The following is the YAML format of a configuration file:

[0062]

[0063] risk: "Buffer Overflow"

[0064] Automatically generate this configuration file through DSL, clearly describe the entry and exit points and characteristics of the taint, and the tool can match specific functions according to the taint characteristics and perform taint analysis.

[0065] S8 is as follows:

[0066] CSA provides some built-in checkers to detect issues related to taint analysis. alpha.security.taint.TaintPropagation is an experimental checker for taint propagation analysis. core.CallAndMessage is used to detect potential issues in function calls, and unix.API is used to detect security issues related to system APIs (such as strcpy, gets, etc.).

[0067] If more complex taint analysis is required, it can be achieved by writing a custom Clang Static Analyzer checker. Due to the particularity of taint detection, it is necessary to comprehensively consider the number of warnings and the effectiveness of the warnings to determine the checker for taint detection.

[0068] After the above steps, the access points of the input and output parameters of the hook function will be carefully monitored, and issues such as null pointer, memory out-of-bounds access, and array out-of-bounds access can also receive sufficient warning prompts when the parameters are not fully verified.

[0069] The present invention solves the problem of data chain breakage in the data flow analysis process of static scanning tools through the method of strengthening hook functions. This method is simple and easy to use. In actual product applications, it solves 20%-30% of the memory problems at only a minimal cost (efficiency consumption), and the effect is remarkable.

[0070] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0071] (1) Through DSL modeling, the present invention takes the hook function as a configuration file, which can be used as the taint of static mining tools such as CSA, or as the input of dynamic mining tools such as fuzz. It avoids the huge workload of manual modeling and the accuracy of pointer functions.

[0072] (2) The idea provided by the present invention can partially solve the problem of broken chain of hook functions analyzed by current static inspection tools. It can discover potential risk vulnerabilities in high-risk interfaces through alternative solutions and give warning prompts. It can find problems through static scanning after the code is put into the library and close the loop in time, avoiding leaving problems in the product testing or user usage process, resulting in huge operation and maintenance costs. By using this solution, through further interception at the hook layer, 20%-30% of the memory problems can be reduced with very little performance degradation, which is of great significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 is the overall process flow diagram of the present invention

[0074] Figures 2-1, 2-2, and 2-3 are the function pointer forms that can be recognized as hooks by DSL;

[0075] Figure 3 is the schematic diagram of the hook function recognition process;

[0076] Figure 4 is the schematic diagram of the vulnerability mining process based on hook functions. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] In order to deepen the understanding of the present invention, the following will make a detailed description of this embodiment with reference to the accompanying drawings.

[0078] Embodiment: Refer to Figures 1-4 , a static memory vulnerability mining method based on hook functions, the method comprising the following steps: S1: The source code performs lexical analysis, syntax analysis, and semantic analysis operations through the Clang compiler LLVM, and converts the analysis results into an abstract syntax tree AST. This step is mainly used to customize the Checker for static vulnerability mining according to function characteristics.

[0080] S2: After generating the AST from the source code, use LLVM as the backend code generator to generate IR. This step is the basis of static analysis. IR is a low-level, structured representation that is language-independent, removing syntax and complex language features, making the analysis simpler and more straightforward.

[0083] S3: Generate the Control Flow Graph (CFG) and Program Dependence Graph (PDG) based on IR; this step mainly lays the foundation for the data flow analysis of the subsequent Checker.

[0085] S4: Generate the Code Property Graph (CPG) jointly based on the AST, CFG, and PDG. The purpose of this step is to integrate the advantages of multiple program representations and provide more comprehensive program analysis capabilities.

[0087] S5: As a code property graph, CPG must find a graph database as the carrier. The graph database options are OverflowDB or Neo4J. Utilizing the efficient query capabilities of OverflowDB or Neo4J, it is possible to quickly analyze the CPG and discover potential security vulnerabilities or code quality issues.

[0089] S6: As shown in Figure 2, the code characteristics of the hook function can be seen. The hook function mainly serves as the right value of an assignment statement of the pointer type or as a function that declares a pointer type. Based on this idea, find a suitable DSL customization tool. After analysis, the leaf property of FunctionAccess in CodeNavi is convenient for obtaining function pointers, which is very convenient for customization in this scenario. For the code characteristics of the hook function, a certain scenario matching is performed on the function pointers obtained from this property, and the hook function extraction process is completed. Save the customized source code as a custom Checker. Using the efficient query capabilities of the graph database, it is possible to easily find all hook function names that meet the characteristics.

[0091] S7: Develop a Checker using the DSL language according to the code characteristics of the hook function to generate a hook function configuration list. These function pointers will serve as the taint entry for static mining. This step is to generate a configuration file that meets the taint requirements format through the DSL. As Figure 3 shown, it is a schematic diagram of generating a hook function from the source code. After steps S1 - S5, at this time, the source code has been converted into a graph database. As long as a Checker that can extract the hook function is developed, the fast and efficient characteristics of the graph database can be utilized to quickly match the characteristics of the hook function, thereby generating a configuration file that meets the taint format.

[0092] S8: Since the input and output parameters of all hooks can be configured as tainted data, this configuration file is used as the tainted input for static scanning tools, including taint models such as CppCheck and CSA (Clang Static Analyzer). At this time, the input interfaces and parameters of the known tainted data, based on the existing taint models and supported taint Checkers of the static tools, can be used to start static vulnerability mining as shown in Figure 4 This system takes the source code as input, and after passing through the clang compiler, it is interpreted into three forms: AST (Abstract Syntax Tree), PDG (Program Dependence Graph), and CFG (Control Flow Graph). CPG (Code Property Graph) is a combined data structure of AST, PDG, and CFG, which contains the syntactic and semantic features in the source code.

[0096] This system uses DSL to identify and model hook functions through a graph database established by analyzing the source code. After the modeling is completed, the open-source software CSA is used for static analysis. The input and output parameters of the hook functions are used as tainted data to analyze the context flow, so as to identify high-risk interfaces and mine for vulnerabilities.

[0097] Hook hooks are similar to the dynamic instrumentation mechanism. The main purpose is to intercept data or execution logic before the target function or instruction is executed, execute a piece of inserted code first, and then execute the original target function. The common form of hook functions is function pointers.

[0098] The DSL tool used to extract hook functions is CodeNavi. Use its attribute node functionAccess to identify all function pointers, and then perform conditional restrictions based on the common form of hook functions to meet the requirements.

[0099]

[0100] Note: functionAcess indicates that the type to be queried here is the function pointer type. fa invariableDeclaration means that the function pointer is in a variable declaration expression, corresponding to examples like table-driven. And fa in binaryOperation means that the function pointer is part of a binary expression, the operator is "=", and the right expression is the function pointer, for the assignment statement scenario.

[0101] Customize the taint detection rules through the taint checker of CSA, and then save the hook functions queried by DSL as a configuration file in taint format, and then it can be mined through the library Checker and custom Checker of CSA. Of course, other static / dynamic white-box taint tools can also be selected here to further mine the hook functions.

[0102] (1) Here, a simple example of out-of-bounds vulnerability mining is given. The code is the function hook hooking and its underlying application function.

[0103]

[0104]

[0105] From this code, it can be seen that if the external input parameters are controllable, both fid_data and fid_num are tainted data. Without sufficient verification, there are security risks at the read and write points of the above several tainted data.

[0106] (2) After compiling the source code, a compilation database and BC files are generated. Modeling is carried out through DSL rules. The file name is dsl_hook.kirin, and its content is as follows:

[0107]

[0108] After the modeling is completed, running this rule, the possible DSL commands are:

[0109] Dsl--scope=compile_command.json--rule_file=dsl_hook.kirin--result_path=. /

[0110] (3) After the DSL modeling is completed, the results are as follows (its specific form depends on the format required by the taint scanning tool):

[0111]

[0112] (4) Using the CSA tool for static vulnerability mining, its possible commands are:

[0113] clang--analyze--analyzer-output html-o <output-dir>-Xclang-analyzer-checker=hook_function test.c

[0114] (5) The warning results are as follows:

[0115]

[0116] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention, and equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A static memory vulnerability mining method based on hook function, characterized in that: The method comprises the following steps: S1: The source code is analyzed lexically, grammatically, and semantically by the Clang compiler LLVM, and the analysis results are converted into an abstract syntax tree AST. S2: After generating AST from the source code, use LLVM as the backend code generator to generate IR. S3: Generate control flow graph CFG and program dependency graph PDG based on IR; S4: Generate code property graph CPG based on AST, CFG and PDG. S5: As a code property graph, CPG must find a graph database as a carrier. OverflowDB or Neo4J is the graph database. By using the efficient query capabilities of OverflowDB or Neo4J, CPG can be quickly analyzed and potential security vulnerabilities or code quality issues can be discovered. S6: Use DSL language to query the graph database. The DSL tool uses CodeNavi language for analysis. By analyzing the nodes and node attribute information in CodeNavi, we know that the attribute of FunctionAccess is convenient for obtaining function pointers. S7: Develop Checker using DSL language according to the code features of the hook function, generate a hook function configuration list, and these function pointers will be used as taint entries for static mining. S8: Find a static checking tool with taint scanning capability, such as Clang Static Analyzer, CppCheck, or Coverity. The present invention uses the open source CSA tool to develop the taint checker. CSA has outstanding performance in detecting potential problems in C / C++ code, including memory leaks, null pointer dereferences, buffer overflows, etc. Here, alpha.security.taint.TaintPropagation of CSA is selected for taint scanning. S9: Configure the hook's input and output parameters to the taint format required by CSA, including taint source, taint propagation, and taint sink. S10: Use the configuration file of the hook function as the taint input of CSA, and select the customized taint Checker to start static vulnerability mining.

2. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S1 is as follows: AST is a tree data structure, in which each node represents a grammatical structure in the source code, including declarations, statements, and expressions. By traversing AST, the structure of the code is analyzed and the required key information is extracted. Suppose the file name is: test.cc, then print the AST information through the following command: clang-Xclang-ast-dump-fsyntax-only test.cc.

3. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S2 is as follows: The command to generate IR middleware is as follows: clang test.cc-O0-emit-llvm-oS test.ll.

4. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S3 is as follows: The control flow graph is a graphical representation of the control flow relationship between basic blocks in a program. The opt tool of LLVM generates a DOT file of CFG, which is then visualized using Graphviz. The command to generate the control flow graph from the file is as follows: opt-dot-cfg test.ll dot-Tpng cfg.test.dot-o cfg.png LLVM provides llvm::DependenceGraph and llvm::DependenceAnalysis to analyze data and control dependencies. Use the above interfaces to customize an LLVMPass to analyze dependencies and generate PDG. The Pass is saved as MyPDGPass.cpp and compiled using LLVM: clang++-c MyPDGPass.cpp$(llvm-config--cxxflags--ldflags--system-libs--libs core analysis)-o MyPDGPass.o Load and run the Pass: opt-load . / MyPDGPass.so -mypdg test.ll -o / dev / null Generate DOT file and visualize it.

5. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S4 is as follows: Use LLVM to generate AST, CFG, and PDG, then write custom tools to combine them into CPG, use Python's graph library (NetworkX) to represent and manipulate CPG, The extraction steps are as follows: Step 1: Extract AST nodes and edges, traverse AST, extract all nodes (including variables, expressions, statements) and edges (including parent-child relationships and sibling relationships), and assign a unique identifier (ID) to each AST node. Step 2: Extract CFG nodes and edges, traverse CFG, extract all basic blocks and jump relationships, and associate CFG nodes with AST nodes (including statements in basic blocks corresponding to statement nodes in AST). Step 3: Extract PDG nodes and edges, traverse the PDG, extract data dependencies and control dependencies, associate PDG edges with AST nodes (including data dependency edges connecting two variable nodes), Step 4: Integrate into CPG, integrate the nodes and edges of AST, CFG and PDG into one graph, and ensure the uniqueness of nodes and edges.

6. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S5 is as follows: Step 1: Import the code. Import your code file (test.cc) in Joern: importCode(". / test.cc") Step 2: Generate CPG. Joern will automatically parse the code into CPG and store it in OverflowDB. Step 3: Use OverflowDB to operate CPG. OverflowDB provides powerful graph query languages ​​(including Gremlin and Cypher) for querying and analyzing CPG.

7. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S6 is as follows: FunctionAccess is a node type that represents function access in CodeNavi. To extract the hook function, use the following description statement:

8. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S7 is as follows: The entry of the taint function is usually defined in the form of a configuration file. The following is the YAML format of a configuration file: The configuration file is automatically generated through DSL, which clearly describes the entry, exit and characteristics of the taint. The tool can then match specific functions and perform taint analysis based on the taint characteristics.

9. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S8 is as follows: Clang Static Analyzer is based on LLVM / Clang, supports C / C++, has built-in taint analysis capabilities, detects memory leaks, null pointer dereferences, buffer overflows and other taints, and is an open source tool that supports customization. Select CSA as the taint scanning tool; S9 is as follows: CSA provides some built-in checkers to detect problems related to taint analysis. alpha.security.taint.TaintPropagation is an experimental checker for taint propagation analysis, core.CallAndMessage is used to detect potential problems in function calls, and unix.API is used to detect security issues related to system APIs. If more complex taint analysis is required, it can be achieved by writing a custom Clang Static Analyzer checker. Possible taint scanning commands may be as follows: clang--analyze-Xanalyzer-analyzer- checker=alpha.security.taint.TaintPropagation test.c.

10. The method for mining static memory vulnerabilities based on hook functions according to claim 1, characterized in that: S10 is as follows: The hook function is used as the input of the taint configuration model, and the static command for taint scanning of the hook function configuration file is added as follows: clang --analyze --analyzer-output html-o <output-dir> -Xclang-analyzer- checker=hook_function_checker test.c hook_function_checker encapsulates alpha.security.taint.TaintPropagation. You can configure which Checkers need to be tainted to achieve the purpose of customization, such as array out-of-bounds and buffer out-of-bounds access issues that users are concerned about.