Active Identification Method, Device and Medium for Software Security Implicit Threats Based on Depth Map Reasoning

Through the deep graph inference method, combined with recurrent neural networks and graph neural networks, the instruction flow diagram is constructed, which solves the shortcomings of static and dynamic analysis in the security analysis of IoT terminal software and achieves efficient and accurate implicit threat recognition.

CN118761059BActive Publication Date: 2025-07-04ELECTRIC POWER RES INST OF STATE GRID ZHEJIANG ELECTRIC POWER COMAPNY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410911963.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2025-07-04
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

In the security analysis of IoT terminal software in the existing technology, static analysis cannot obtain the intermediate value of variables during program execution. Dynamic analysis requires a lot of time and calculation costs. Dynamic and static fusion analysis has the problems of model input exceeding the limit and low recognition accuracy.

Method used

Using a deep graph inference method, the software code is converted into an instruction superset through a disassembly algorithm, combining recurrent neural networks and graph neural networks, the instruction flow graph is constructed, semantic features and graph features are fused, and the depth graph inference model is trained and classified, and implicit threats are identified.

Benefits of technology

It significantly improves the recognition accuracy and robustness of implicit threats for software security, and realizes efficient identification of dynamic and static fusion analysis without increasing computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118761059B_ABST
    Figure CN118761059B_ABST
Patent Text Reader

Abstract

The present invention relates to an active identification method, device and medium for software security implicit threats based on depth graph reasoning. Aiming at the problem of inaccurate existing identification methods, an active identification method for software security implicit threats based on depth graph reasoning is provided, including the following steps: obtaining an instruction superset through a disassembly algorithm; obtaining the semantic representation of an instruction sequence by constructing a node attribute feature extraction model based on instruction embedding; obtaining a graph representation vector through the constructed instruction flow graph; introducing the obtained semantic representation into a graph neural network to obtain a depth graph reasoning model for software security implicit threats, inputting the semantic representation and the graph representation vector into the constructed depth graph reasoning model for classification to obtain the validity of the instruction; and determining whether it is an implicit threat by testing whether the instructions determined to be invalid / redundant can form a complete function without affecting program execution. The present application integrates the semantic features of the instruction sequence and the instruction flow graph, and combines a recurrent neural network and a graph neural network to significantly improve the accuracy of determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of threat hunting and software security, and particularly to an active identification method, device, and medium for software security implicit threats based on deep graph reasoning. Background Art

[0002] Security incidents of Internet of Things (IoT) terminals occur frequently. For example, in the field of vehicle networking, the internationally renowned automobile brand Tesla had an iBeacon privacy leakage problem in 2020. In medical IoT software, due to its complex and dynamically changing nature, both its network and physical components are vulnerable to cyberattacks. If a network intruder reprograms a sensor or interferes with the communication protocol of a device, it may pose an extremely high risk to patients' lives. The security analysis of binary programs is divided into three categories: static analysis, dynamic analysis, and hybrid static-dynamic analysis. Static analysis first converts binary code into assembly instructions, then converts the assembly instructions into an intermediate language for analysis, and finally realizes vulnerability mining through pattern matching and similarity comparison. However, static analysis cannot obtain the intermediate values of variables during program execution, thus affecting the analysis results. Dynamic analysis can be divided into gray-box, black-box, and white-box testing according to its different guiding methods. It mainly adopts the technique of fuzz testing. However, dynamic analysis requires providing sufficient inputs so that the execution path can cover most of the code, which requires a large amount of time and computational cost. Hybrid static-dynamic analysis first performs static analysis and then uses the results of static analysis to assist dynamic testing. For example, VulDeePecker does not use traditional rule-based techniques but uses deep learning for vulnerability detection. Most static-dynamic analyses combine natural language processing models for analysis. However, the length of the assembly instructions used as input often exceeds the upper limit that the model can bear. Existing research often uses random segmentation and random input, thus affecting the recognition accuracy of the model for malicious instructions. Summary of the Invention

[0003] The purpose of the present invention is to provide an active identification method, device, and medium for software security implicit threats based on deep graph reasoning, which can effectively fuse the semantic features of assembly code and the features of instruction flow graphs and restore function functionality, thereby significantly improving the effectiveness and robustness of active identification of software security implicit threats.

[0004] The purpose of the present invention can be achieved through the following technical solutions:

[0005] An active identification method for software security implicit threats based on deep graph reasoning, characterized by comprising the following steps:

[0006] S1, converting the software code to be detected into an instruction superset through an anti-assembly algorithm, where the instruction superset is an initial set containing all real instructions and redundant instructions;

[0007] S2. By constructing a node attribute feature extraction model based on instruction embedding, obtain the semantic representation of the instruction sequence, including:

[0008] Extract instruction metadata from the instruction superset, convert the instruction metadata into feature vectors, and cyclically learn the obtained feature vectors through a recurrent neural network to obtain the semantic representation of the instruction sequence;

[0009] S3. By constructing a binary instruction flow graph, obtain the graph representation vector, including:

[0010] Take each instruction in the obtained instruction superset or instruction metadata as a node of the instruction flow graph, take the relationship between each instruction as the edge of the instruction flow graph, and input the instruction flow graph into the graph attention model to obtain the graph representation vector;

[0011] S4. By introducing the obtained semantic representation into the graph neural network, obtain a deep graph inference model for software security implicit threats, input the semantic representation and the graph representation vector into the constructed deep graph inference model for classification, and obtain the validity of the instruction;

[0012] S5. Determine whether it is an implicit threat by testing whether the instructions determined to be invalid / redundant can form a complete function without affecting program execution.

[0013] Superset disassembly is mostly used for binary rewriting, which disassembles each executable byte offset. Although most instructions are false positives, the results contain all true positives, so each possible transfer target can be monitored during binary rewriting. Advanced superset decompilation techniques such as probabilistic disassembly use probability models to model the uncertainties caused by interleaved code and data and indirect transfer targets. It considers register definition-use relationships, control flow convergence, control flow crossing, and calculates the probability of each address based on these features. Experiments show that it has no false negative rate, and the average false positive rate is only 3.7%.

[0014] Relational graph convolutional neural network is a graph-based neural network model that mainly runs on Local graph neighborhoods and can process large-scale relational data. To enhance the accuracy and robustness of software security analysis, it is a common method to learn indexable feature representations from the software control flow graph and then calculate the distance between feature representation vectors through a similarity function. Among them, the process of extracting feature representation vectors is called graph embedding. Currently, in the field of software security analysis, especially binary program vulnerability analysis, using graph embedding to construct software execution state representation has become one of the research hotspots.

[0015] Furthermore, in the step S3, the semantic representation corresponding to each instruction serves as the node attribute of the instruction flow graph.

[0016] Furthermore, in the step S3, the relationships include forward, backward, and overlapping.

[0017] Furthermore, the expression of the node hidden state is:

[0018]

[0019] where i is the node index, l is the number of layers of the neural network model; r is the relationship type index, and j represents the node index within the neural network; is the d2-dimensional hidden state of the node v at the l-th layer i ; h i (l+1) is the d2-dimensional hidden state of the node v at the (l + 1)-th layer i ; is the d2-dimensional hidden state of the adjacent nodes of the node v at the l-th layer i ; represents the set of adjacent indices of the node v under the relationship r ∈ R i ; represents the number of nodes in is the weight matrix of the relationship r ∈ R at the l-th layer; is the weight matrix of the node itself in the l-th layer.

[0020] Furthermore, during the training process of the deep graph reasoning model, the following formula is used to train the semantic representation and the graph representation vector:

[0021] J(Θ, p, y) = ∑(-(y · log(p) + (1 - y) · log(1 - p))) (2);

[0022] where J(Θ, p, y) represents the training function, Θ represents the model parameters, p represents the probability, and y is the true label.

[0023] An active identification device for software security implicit threats based on deep graph reasoning, characterized in that

[0024] a processing module that converts the software code to be detected into an instruction superset through a disassembly algorithm, and the instruction superset is an initial set containing all real instructions and redundant instructions;

[0025] a semantic representation construction module that obtains the semantic representation of the instruction sequence by constructing a node attribute feature extraction model based on instruction embedding, including:

[0026] extracting instruction metadata from the instruction superset, converting the instruction metadata into feature vectors, and cyclically learning the obtained feature vectors through a recurrent neural network to obtain the semantic representation of the instruction sequence;

[0027] The graph representation construction module obtains a graph representation vector by constructing a binary instruction flow graph, which includes:

[0028] Regarding each instruction in the obtained instruction superset or instruction metadata as a node of the instruction flow graph, regarding the relationships between the instructions as the edges of the instruction flow graph, and inputting the instruction flow graph into a graph attention model to obtain a graph representation vector;

[0029] The validity judgment module obtains a deep graph inference model for software security implicit threats by introducing the obtained semantic representation into a graph neural network, classifies the semantic representation and the graph representation vector by inputting them into the constructed deep graph inference model, and obtains the validity of the instructions;

[0030] The implicit threat judgment module determines whether it is an implicit threat by testing whether the instructions determined to be invalid / redundant can form a complete function that does not affect program execution.

[0031] An active identification device for software security implicit threats based on deep graph inference, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the active identification method for software security implicit threats based on deep graph inference.

[0032] A computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the active identification method for software security implicit threats based on deep graph inference.

[0033] First, using static analysis, convert the binary program into assembly code, and then convert the original binary code into an instruction superset through disassembly; then, respectively construct a node attribute feature extraction model based on instruction embedding and a binary instruction flow graph, so as to obtain an accurate semantic representation vector of the instruction sequence and a graph representation vector of the instruction flow graph; secondly, construct a deep graph inference model for software security implicit threats, input the node attributes and edges into the graph neural network and the fully connected layer, and infer the validity of each instruction. Finally, restore the function according to the instruction validity. Among them, if the instructions determined to be invalid / redundant can form a complete function that does not affect program execution, they are determined to be implicit threats.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] (1) The designed active identification method for software security implicit threats is based on static analysis, fuses and analyzes the semantic features of the instruction sequence and the call relationship graph features of the instruction flow graph, and comprehensively uses technologies such as recurrent neural networks, graph neural networks, and superset disassembly, significantly improving the accuracy of instruction validity determination.

[0036] (2) In the stage of actively identifying implicit software security threats, the present invention includes restoring function capabilities based on the determination of instruction validity and integrating dynamic analysis methods for test verification, achieving dynamic and static fusion analysis under unknown source code conditions, and improving the identification efficiency and accuracy of instruction-level implicit threats without significantly increasing the computational cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is the overall framework diagram of the method.

[0038] Figure 2 It is the efficiency comparison diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manners and specific operation processes, but the applicable scope of the present invention is not limited to the following embodiments.

[0040] Refer to Figure 1 , in this example, the specific implementation process of the project is shown. The user may submit a binary program to be detected, which mainly includes the following steps:

[0041] S1. Convert the software code to be detected into an instruction superset through a disassembly algorithm. This instruction superset is an initial set containing all real instructions and redundant instructions;

[0042] The given code is split into segments of 15 bytes each. Each instruction is a (Opcode, ModRM, SIB, REX) tuple. Since an instruction can consist of at most 15 bytes, the model can input 15 consecutive bytes. When the number of bytes is less than 15, it is filled with 0x90 (nop). After an instruction is decoded, it may have features such as Opcode, ModRM, SIB, Displacement, and Immediate. However, for the sake of improving efficiency, the present invention only uses four key features, namely REX, Opcode, ModRM, and SIB, as its semantic representation.

[0043] In addition, since the process of disassembling any instruction in the present invention is independent, only n threads need to be created on the GPU to achieve the parallelization of the disassembly process. Therefore, the disassembly efficiency can be significantly improved.

[0044] S2. Obtain the semantic representation of the instruction sequence by constructing a node attribute feature extraction model based on instruction embedding, including:

[0045] Extract instruction metadata from the instruction superset, convert the instruction metadata into feature vectors, and cyclically learn the obtained feature vectors through a recurrent neural network to obtain the semantic representation of the instruction sequence;

[0046] A constant value is added to Opcode, ModRM, SIB, and REX to reduce overlap. Therefore, there are 1025 different OpenCodes, 257 ModRMs, 257 SIBs, and 17 REXs in the present invention. Each field has a reserved value, which is used when the corresponding field does not exist, that is, the input size of the embedding layer is 1556. The instruction metadata obtained in step 1) is first input into the embedding layer for representation, and then the representation vector is injected into the vanilla RNN for cyclic learning to obtain the node attributes corresponding to each instruction.

[0047] At the same time, perform relationship mining and graph modeling on the instruction superset obtained in step 1). For each instruction v i ∈V, an instruction semantic representation x i will be assigned, that is, the instruction semantic representation. Each edge (v i ,r,v j )∈E is labeled with a relationship r∈R to represent the edge type. R = {f, b, o} represent the forward, backward, and overlapping types respectively. If the label r in (v i ,r,v j ) is a forward relationship, it means that the next instruction j of i, that is, i calls j or i jumps to j. If r is an overlapping relationship, it means that instructions i and j overlap each other. That is, the start point of instruction j is inside instruction i. The constructed instruction flow graph will be input into the graph neural network for inference analysis.

[0048] S3. By constructing a binary instruction flow graph, obtain the graph representation vector, including:

[0049] Take each instruction in the obtained instruction superset or instruction metadata as a node of the instruction flow graph, take the relationships between the instructions as the edges of the instruction flow graph, input the instruction flow graph into the graph attention model to obtain the graph representation vector; fuse the semantic representation and the graph representation, and use the following propagation model to update the hidden state of each node v i :

[0050]

[0051] where i is the node index, l is the number of layers of the neural network model; r is the relationship type index, and j represents the node index inside the neural network; is the d2-dimensional hidden state of node v i at the l-th layer; h i (l+1) is the node v at the l+1-th layeri The d2-dimensional hidden state; is the node v at layer l i The d2-dimensional hidden states of adjacent nodes; represents the set of adjacent indices of node v under the relationship r ∈ R i ; represents the number of nodes in is the weight matrix of the relationship r ∈ R at layer l; is the weight matrix of the node itself in layer l.

[0052] During training, each instruction embedding is propagated and updated L times through different relationships: forward, backward, and overlapping, to capture information from adjacent nodes. The final output hiL is fed into a classifier: a fully connected layer to reduce the dimension to 1, and then through a sigmoid activation to generate a probability p. We attempt to minimize the binary cross-entropy loss function:

[0053] J(Θ, p, y) = ∑(-(y · log(p) + (1 - y) · log(1 - p))) (2);

[0054] where J(Θ, p, y) represents the training function, Θ represents the model parameters, p represents the probability, and y is the true label. All trainable modules of the model are connected together and trained in an end-to-end manner.

[0055] S4. By introducing the obtained semantic representation into the graph neural network, a deep graph inference model for software security implicit threats is obtained. The semantic representation and the graph representation vector are input into the constructed deep graph inference model for classification to obtain the effectiveness of the instruction;

[0056] In the test stage, the output of the Sigmoid layer is the effectiveness of each instruction.

[0057] S5. By testing whether the instructions determined to be invalid / redundant can form a complete function that does not affect program execution, it is further determined whether the invalid / redundant instructions are implicit threats.

[0058] After determining the effectiveness of the instructions, the invalid / redundant instructions are extracted for function reconstruction. If a complete function that does not affect program execution can be formed, it is determined to be an implicit threat; if a complete function cannot be formed or the reconstructed function will affect program execution, it is determined to be a non-frequent execution path. For non-frequent execution paths, their security status needs to be further determined through dynamic analysis.

[0059] Performance comparison

[0060] Figure 2Shows the comparative analysis results of the present invention with IDAPro Binary Ninja, Ghidra, Datalog, and XDA in terms of code segment size and disassembly time. The y-axis of this figure is in logarithmic scale. For IDAPro, Binary Ninja, and Ghidra, run them in console / headless mode to avoid unnecessary GUI costs. For DatalogDisassembly, directly adopt the numbers reported by the tool. When testing disassemblers on the CPU, to ensure fairness, only use one CPU core. The present invention has a significant advantage in terms of GPU running time: its throughput is approximately 24.5MB / s, which is about 170 times faster than that on the CPU (146KB / s). Moreover, in the same hardware environment, it is faster than other disassemblers: IDAPro 72KB / s, XDA (GPU) 47KB / s, BinaryNinja 11KB / s, Ghidra 10KB / s, Datalog 5KB / s (file size is about 1MB), and XDA (CPU) 140B / s.

Claims

1. An active identification method for software security implicit threats based on depth map reasoning, characterized in that It includes the following steps: S1. Convert the software code to be detected into an instruction superset through a disassembly algorithm. The instruction superset is an initial set containing all real instructions and redundant instructions; S2. Obtain the semantic representation of the instruction sequence by constructing a node attribute feature extraction model based on instruction embedding, including: Extract instruction metadata from the instruction superset, convert the instruction metadata into feature vectors, and cyclically learn the obtained feature vectors through a recurrent neural network to obtain the semantic representation of the instruction sequence; S3. Obtain a graph representation vector by constructing a binary instruction flow graph, including: Take each instruction in the obtained instruction superset or instruction metadata as a node of the instruction flow graph, take the relationship between each instruction as an edge of the instruction flow graph, and input the instruction flow graph into a graph attention model to obtain a graph representation vector; S4. Introduce the obtained semantic representation into a graph neural network to obtain a deep graph inference model for implicit threats to software security. Input the semantic representation and the graph representation vector into the constructed deep graph inference model for classification to obtain the validity of the instruction. The expression of the deep graph inference model is: Among them, is the d2-dimensional hidden state of the node v at the l-th layer i ; h i (l+1) is the d2-dimensional hidden state of the node v at the (l + 1)-th layer i ; is the d2-dimensional hidden state of the adjacent nodes of the node v at the l-th layer i ; represents the set of adjacent indices of the node v under the relationship r ∈ R i ; represents the number of nodes in ; W_r^l is the weight matrix of the relationship r ∈ R at the l-th layer is the weight matrix of the node itself in the l-th layer; S5. Determine whether it is an implicit threat by testing whether the instructions determined to be invalid / redundant can form a complete function that does not affect program execution.

2. The active identification method for software security implicit threats based on depth map reasoning according to claim 1, wherein In step S3, the semantic representation corresponding to each instruction is used as the node attribute of the instruction flow graph.

3. The active identification method for software security implicit threats based on depth map reasoning according to claim 1, characterized in that In step S3, the relationships between the instructions include forward, backward, and overlapping.

4. An active identification method for software security implicit threats based on depth map reasoning according to claim 1, characterized in that During the training process of the deep graph inference model, the following formula is used to train the semantic representation and the graph representation vector: J(Θ,p,y)=∑(-(y·log(p)+(1-y)·log(1-p)))(2); where J(Θ,p,y) represents the training function, Θ represents the model parameters, p represents the probability, and y is the true label.

5. An active recognition device for implicit threats to software security based on deep graph inference, characterized in that a processing module that converts the software code to be detected into an instruction superset through a disassembly algorithm. The instruction superset is an initial set containing all real instructions and redundant instructions; a semantic representation construction module that obtains the semantic representation of the instruction sequence by constructing a node attribute feature extraction model based on instruction embedding, including: Extract instruction metadata from the instruction superset, convert the instruction metadata into feature vectors, and cyclically learn the obtained feature vectors through a recurrent neural network to obtain the semantic representation of the instruction sequence; a graph representation construction module that obtains a graph representation vector by constructing a binary instruction flow graph, including: Take each instruction in the obtained instruction superset or instruction metadata as a node of the instruction flow graph, take the relationship between each instruction as an edge of the instruction flow graph, and input the instruction flow graph into a graph attention model to obtain a graph representation vector; a validity judgment module that introduces the obtained semantic representation into a graph neural network to obtain a deep graph inference model for implicit threats to software security. Input the semantic representation and the graph representation vector into the constructed deep graph inference model for classification to obtain the validity of the instruction. The expression of the deep graph inference model is: Among them, is the d2-dimensional hidden state of the node v at the l-th layer i ; h i (l+1) is the d2-dimensional hidden state of the node v at the (l + 1)-th layer i ; is the d2-dimensional hidden state of the adjacent nodes of the node v at the l-th layer i ; represents the set of adjacent indices of the node v under the relationship r ∈ R i ; represents the number of nodes in ; is the weight matrix of the relationship r ∈ R at the l-th layer is the weight matrix of the node itself in the l-th layer The implicit threat judgment module determines whether it is an implicit threat by testing whether the instructions determined to be invalid / redundant can form a complete function without affecting program execution.

6. An active recognition device for software security implicit threats based on depth map reasoning, characterized in that, It includes a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for actively identifying software security implicit threats based on depth graph reasoning according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, A program is stored thereon. When the program is executed by a processor, it implements the method for actively identifying software security implicit threats based on depth graph reasoning according to any one of claims 1-4.

Citation Information

Patent Citations

  • Malicious code detection method based on memory evidence obtaining and graph neural network

    CN115659330A

  • Malicious code classification method and device, electronic equipment and storage medium

    CN118070276A