Cross-language malicious software detection system and method based on graph neural network
Through a cross-language malware detection system based on graph neural network, a control flow diagram of Java and local code is generated and integrated, and a gated graph neural network is used for learning, solving the problem of low accuracy in cross-language malware detection, and achieving efficient and accurate malicious code detection.
Patent Information
- Application Number
- CN202510350168.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
Existing malware detection methods cannot effectively deal with cross-language malware, especially ignoring the syntax and semantic differences between Java and local code, resulting in low detection accuracy and prone to omissions or false alarms.
A cross-language malware detection system based on graph neural network is adopted. The control flow graph generation module is used to generate control flow graphs of Java and local codes, and the relationship between the two graphs is extracted and fused, combined with the gated graph neural network for learning, so as to realize cross-language malware detection.
It effectively improves the efficiency and accuracy of cross-language malware detection, completely retains the semantics of the application, reduces the false alarm and missed alarm rates, and improves the robustness and generalization capabilities of detection.
Smart Images

Figure CN120277666A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of malware detection, and particularly relates to a cross - language malware detection system and method based on a graph neural network. Background Art
[0002] Traditional malware detection methods usually extract application features for unified vectorization and use traditional neural networks for training to obtain a classification model. Most of these methods only focus on Java bytecode, abandon the review of the Native part, and the process of converting features into a unified intermediate representation is inefficient and may lose certain program semantic information. Some existing cross - language detection tools only focus on the risk of sensitive information leakage. The difficulty of cross - language analysis lies in the context communication analysis between two languages. Separate analysis and review may lead to the loss of the semantics of the entire application, resulting in the omission or false alarm of cross - language malicious behavior.
[0003] Most classic analysis and review tools, including Flowdroid, TaintDroid, etc., cannot support the analysis and inspection of native code; cross - language analysis and review tools such as JN - SAF, Ndroid, PILdroid, and Soprotector only support cross - language information leakage detection, and the security review of malicious advertisements, phishing software, SMS hijacking, and vulnerable applications is still lacking; existing machine - learning - based tools, such as CDGDroid, perform poorly on applications with native libraries. The syntactic and semantic differences between Java and native languages will affect the review accuracy of detection tools. In addition, converting the CFG (Control Flow Graph) of an application to an intermediate representation (IR) for CNN input is also a time - consuming task. When the application is complex, the conversion may also lead to the loss of application semantics. Most of the malware detection methods in existing invention patents cannot consider the importance of the Native part in cross - language malware and ignore the possible malicious behaviors in this part. Summary of the Invention
[0004] In order to overcome the deficiencies of the above - mentioned prior art, the purpose of the present invention is to provide a cross - language malware detection system and method based on a graph neural network. By representing the application as a graph representation for processing and designing a graph fusion algorithm to combine the graphs of native and Java code, and using a gated graph neural network to learn the combined graph to achieve cross - language malware detection, the efficiency of information dissemination is effectively improved.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] A cross - language malware detection system based on graph neural network, including a control flow graph generation module, a graph fusion module, and a graph learning module;
[0007] The control flow graph generation module is responsible for pre - processing the application program and generating control flow graphs for Java code and native code respectively;
[0008] The graph fusion module is used to extract the relationships between control flow graphs and fuse the two graphs to obtain a fused graph;
[0009] The graph learning module learns from the fused graph, is responsible for vectorizing the nodes in the fused graph, and classifying the control flow graphs of benign software and malware.
[0010] A detection method for a cross - language malware detection system based on graph neural network includes the following steps;
[0011] Step 1: The control flow graph generation module is responsible for pre - processing the application program and generating control flow graphs for Java code and native code respectively;
[0012] Step 2: The graph fusion module is used to extract the relationships between control flow graphs and fuse the two graphs to obtain a fused graph;
[0013] Step 3: The graph learning module learns from the fused graph, is responsible for vectorizing the nodes in the fused graph, and classifying the control flow graphs of benign software and malware.
[0014] The specific steps included in Step 1 are as follows:
[0015] Step 1.1, construct a dataset of benign samples and malicious samples, screen according to the proportion of cross - language applications in the total number of applications in the real part, and store the screened sample dataset locally;
[0016] Among them, benign samples refer to application programs that comply with official security specifications. The functions and behaviors of these applications meet user expectations and there are no behaviors that harm user privacy or security;
[0017] Malicious samples: refer to application programs containing malicious code or harmful behaviors. These programs are designed to perform operations that harm the user's device or privacy security without the user's authorization. Malicious samples may include malicious behaviors such as viruses, trojans, spyware, ransomware, or advertising fraud.
[0018] Step 1.2, decompile each benign sample and malicious sample in the dataset to obtain the dex file, so file, and Manifest file therein;
[0019] Step 1.3: Filter the Java code control flow (Activity) in each benign sample and malicious sample according to the Activity names defined in the Manifest file. Then, extract the control flow graphs of Java code and native code from the so file and dex file respectively, and extract the corresponding activity files.
[0020] In step 1.2, use an apktool-based decompiler to decompile the APK (Android application package) into three parts, namely, the.dex file, the Manifest file and the Res file, and the.so file.
[0021] In step 1.3, by parsing the Activity components defined in the Manifest file, use the regular expression matching method to filter the control flow (Activity) in the Java code, traverse the directory storing the decompiled files to extract the native code of the application, and then use the Native control flow graph generation component to generate the control flow graph (CFG) for the native part; use Dex2Jar to collect Java activity information.
[0022] Use the Java control flow graph generation component and the Native control flow graph generation component to generate the CFG of the Java part and the native part of the app respectively. The Java control flow graph generation component based on the Androguard decompiler receives the ActivityName from the Activity analysis component and uses the CFG generation tool to decompile the APK and build the CFG. The Native control flow graph generation component receives the.so files decompiled by the decompiler as input and passes them to the flow graph generation component to generate the CFG of the native code.
[0023] Use D2j-dex2jar to extract the Activity, operate on the Android Dalvik (.dex) file format and Java's (.class), perform the file format conversion from dex to class, and generate a detailed Java activity file.
[0024] The specific steps of step 2 are as follows:
[0025] Step 2.1: Extract the relationships between Java internals and between Java and native code according to the activity files described in step 1.3. For the relationships between different control flow graphs, construct three types of edges through the relationship extractor: namely, JNI call edges, class import edges, and parallel edges.
[0026] Among them, the JNI call edge refers to the relationship of JNI calls from Java code to native code; the class import edge refers to the relationship of class imports between two control flow graphs that both belong to the Java part; the parallel edge means that if there is no such relationship between two graphs.
[0027] Step 2.2: Receive the control flow graphs corresponding to the Java code and native code generated in Step 1.3, and combine the inter-graph relationships extracted in Step 2.1. The multi-relationship directed graph generator uses the control flow graphs and the relationships between the control flow graphs to construct an application-level multi-relationship directed graph (Multi-Relationship Directed Graph, MRDG).
[0028] During the extraction process of Step 2.1:
[0029] The relationship extractor receives the Manifest file output in Step 1.2, traverses all Activities and keywords defined by the JNI call mechanism, extracts the relationships between Java classes and native files, and submits the relationships and all the control flow graphs generated in Step 1.3 to the multi-relationship directed graph generator component. Then, the multi-relationship directed graph generator component merges the control flow graphs and the relationships between the control flow graphs to construct a fusion graph. This fusion graph not only covers all the separated CFG information of an app, but also contains the relationship connections between each CFG.
[0030] The relationship extractor uses the semantic information and structural information of basic blocks to represent the semantic information of nodes in the graph. A basic block is the smallest unit in program execution and contains a set of sequentially executed instructions; the relationship extractor analyzes the instructions in the basic block, extracts the key semantic information, and uses it to construct the feature representation of the nodes.
[0031] The multi-relationship directed edges of the CFG represent the call relationships between Java codes or between Java nodes and native nodes. According to the different types of call relationships, three additional types of edges are defined: JNICall (JNI call edge), ClassImport (class import edge), and Parallel (parallel edge).
[0032] The JNICall edge is used to record the relationship of JNI calls from Java code to native code, and the JNICall edge is used to connect the basic blocks in the Java part to the native part.
[0033] The ClassImport edge uses a directed ClassImport edge to connect two Java files with an import relationship.
[0034] The Parallel edge uses a Parallel edge to represent a graph without the above two relationships.
[0035] Further, in the obtained CFG, the multi-relational directed graph generator uses a series of code blocks including bytecodes or assembly instructions as graph nodes, and the control flow paths as edges. The Java CFG and the native CFG are respectively defined as:
[0036]
[0037] where G java and G native respectively represent the sets of control flow graphs generated by the Java part and the native part. Among them, the K CFGs of the Java part: belong to G java , and the L CFGs of the native part: belong to G native . V and E represent the nodes and edges in the control flow graph, that is, basic blocks and control flow edges. Their subscripts J and N respectively indicate that they belong to the Java part and the native part.
[0038] The multi-relational directed graph generator merges these two types of CFGs into a multi-relational directed graph through the defined relational edges. There are three types of edges in the graph: JNICall, ClassImport, and Parallel (JNI call edge, class import edge, and parallel edge). There are two types of key APIs for the Java code CFGs and the native code CFGs to call Java and native functions, including System.LoadLibrary() and import;
[0039] Abbreviate the edges and key APIs as Edges = {e jni , e imp , e par} and APIs = {interface load , interface imp}. The input of this component is the CFG sets from both G java and G native , as well as the set of Java files F java corresponding to the control flow graph in Java. The output of the component is the final graph G Final after all graphs are merged.
[0040] Further, the specific process of generating the multi-relational directed graph is as follows, which is used to fuse all the graphs generated from the native part and the Java part:
[0041] 1) Graph fusion stage: Responsible for merging the sub-graph G sub into the final graph G Final . In this stage, a graph is randomly selected from the Java CFG set and is passed into the sub - graph fusion stage to generate a sub - graph G with as the root node sub , and the nodes and edges of G Final and G sub are connected together by the defined Parallel edges;
[0042] 2) Sub - graph fusion stage: Responsible for recursively merging the graph according to the call relationship between the given CFG g java and the API - corresponding graph to generate a sub - graph G with java as the root node sub , and G sub is part of the MRDG. When all sub - graphs are merged into the final graph G Final , and G Final is the finally fused MRDG;
[0043] The sub - graph fusion stage needs to find the corresponding Java file according to g java , traverse the file and judge the call type according to the given API. This stage considers the following two cases:
[0044] a. In the Java file corresponding to g java , if there is a Java class import, the sub - graph fusion stage obtains g' java of the imported class, connects g imp and g' jave with an edge e java , then adds g' java to G sub , and deletes it from G java . After that, recursively merge all graphs in G java ;
[0045] b. In the Java file corresponding to g jave , if there is a JNI - called API, the sub - graph fusion stage obtains g' native of the native loading library, connects g jni and g' java with an edge e native , then adds g' native to G sub , and then returns G sub to terminate the recursion.
[0046] The specific steps of step 3 are as follows:
[0047] Step 3.1, The gated graph neural network uses the vectorized MRDG (the output of the graph fusion module) as the input and uses multiple parallel feature extraction channels. After passing through the feature extraction channels, high - dimensional feature vectors of each node in the MRDG are obtained;
[0048] Each feature extraction channel consists of a gated graph neural layer, a pooling layer, and a convolutional layer connected in sequence;
[0049] The gated graph neural layer is used to capture the global dependencies of graph-structured data, propagate information between adjacent nodes through the MessagePassing mechanism, enabling each node to learn the information of its neighborhood. The GateMechanism allows the model to control the accumulation of information at different time steps, preventing the vanishing gradient problem and improving the modeling ability of long-term dependencies;
[0050] The pooling layer reduces the computational complexity, making the model more efficient; integrates the information of multiple nodes into a fixed-size feature vector for subsequent classification tasks; alleviates the impact of noise and improves the generalization ability.
[0051] The convolutional layer combines with the pooling layer to extract features in the high-level semantic space and enhance the classification ability of the model.
[0052] Step 3.2: Use the high-dimensional feature vectors of each node in the MRDG generated in Step 3.1 as input for MRDG graph embedding representation; convert the feature vectors of each node in the MRDG into a fixed-length digital sequence to ensure that different types of code snippets are learned in a unified format; represent the multi-relational directed edges in the MRDG in the form of an adjacency matrix;
[0053] Step 3.3: Perform training and learning on the MRDG graph embedding representation generated in Step 3.2 using a gated graph neural network. The gated graph neural network model contains two gates, an update gate and a relevance gate; the update gate is used to control the update of memory. If two associated nodes are far apart, it will maintain the state at the last moment, having a memory effect and solving the problem of long sequences; the relevance gate represents the relevance between the previous state and its candidate value. The activation functions of the update gate and the relevance gate are both sigmoid. After iterating for T time steps, use the node embedding set in the current network as the final set and input it into the graph-level classification layer to achieve the prediction of benign and malicious software. The prediction process includes a fully connected convolutional layer and a density layer, which map the final representation set of the graph to a vector to achieve the classification (benign or malicious) of application software.
[0054] Specifically, Step 3.2 is as follows:
[0055] First, vectorize the multi-relational directed graph and represent the vectorized application-level multi-relational directed graph (MRDG) as an adjacency matrix and node embeddings, and then input the adjacency matrix and node embeddings into the multi-relational directed graph neural network to train the malware detection model, that is, the gated graph neural network model in Step 3.3;
[0056] The graph learning phase has two steps, namely node vectorization and GNN learning. The workflow of the graph learning module specifically has the following two steps:
[0057] (1) Node vectorization:
[0058] The Node Vectorization Composer (NVC) component is responsible for mapping a basic block v i That is, after feature extraction in step 3.1, the high-dimensional feature vectors of each node in the MRDG are mapped to a graph node vector h i As the input to the GNN, H i Is defined as follows: h i = NVC(v i ), i = 1, …, |V|;
[0059] Use 20 features (14 in Java code and 6 in native code) for vectorization. For each node in the graph, first initialize a 20-dimensional vector That is, each dimension represents a feature, filled with all zeros. For each code block, that is, each node in the MRDG, determine the code type in the Java part or the native part, and then count the number of occurrences of each feature in each basic block. Arrange these 20 feature values in order to form a numerical vector;
[0060] For the basic block J1, initialize the embedding as a 20-dimensional vector And fill it with all zeros. Count the number of instructions in the block according to the instructions in the Java part, and then arrange the numbers in sequence to obtain a 14-dimensional vector And copy C s To the first 14 dimensions h0 to h v In, at this time h 13 To h 14 To h 20 Are still all filled with 0. Finally, arrange these numerical vectors to the corresponding vertices v s Of the MRDG;
[0061] Each node in the MRDG is finally represented by a numerical vector as g = <H, E>, where H = {h1, …, h |V|} and E are the vectorized basic blocks and multi-relationship edges. In addition, each vertex h i ∈ H represents the initial numerical feature vector, and use G = {g1,.., g n} as the input to the multi-relationship directed graph neural network;
[0062] (2) Multi-relationship directed graph neural network:
[0063] Based on the extracted feature G, malware detection can be defined as a binary classification problem, that is, learning to judge whether a given application is malicious, constructing a dataset {(x i ,y i )|x i ∈X,y i ∈Y}, i ∈ {1, 2, ..., n}, where X represents a series of applications to be learned, Y = {0, 1} n represents an n-dimensional vector, where 1 represents malware, 0 represents benign, and n represents the number of all APKs. The multi-sided graph g i (V, E) ∈ G, i ∈ 1, 2, …, n, merged in the previous stage represents an APK x i , where G represents a set of graphs generated from all instances, and v i ∈V is a vertex of the multi-relational directed graph (MRDG), which contains the content of assembly-level instructions from the Java part or the native part. e = (v s , v t ) ∈ E represents a directed edge v s →v t . For each edge, the graph also contains an edge type l e ∈ {1, …, k}, where k is the number of edge types.
[0064] In step 3.3:
[0065] The update gate is used to control the update of memory: if two associated nodes are connected far away, the state of the previous moment will be maintained, which makes the network have a memory effect and solves the problem of long sequences. The relevant gate represents the correlation between the state of the previous moment and its candidate value;
[0066] Given an embedded graph g i (H, E) ∈ G, using the feature vectors H of all nodes as the node state vectors
[0067]
[0068] For each node v i in the graph g, during the parameter update process, it needs to communicate with neighbor nodes. The process of node aggregation according to the type and direction of the edges between nodes is as follows
[0069]
[0070]
[0071] is the adjacency matrix composed of the set of edges E and the type l e , where e represents from node vs to v t directed edge, l e represents the type of this edge, equal to 1 indicates that node v s , v t is connected by an edge of type l e ; is the weight to be learned, b is the bias, and the process of repeated unfolding is executed for T steps;
[0072]
[0073] where controls the update information, controls the reset information, σ(·) represents the logistic sigmoid function, and the above three expressions show the GRU-like update process, the aggregated information from neighboring nodes, and the preparatory steps for generating the new state of the node;
[0074]
[0075] where is the finally updated node state, including the selected forget and the newly generated information to be remembered ⊙ is the element-wise multiplication;
[0076] After iterating for T time steps, the set of node embeddings in the network is is the final node representation of the node set V of the graph. Then, using the final set as the input to the graph-level classification layer to predict whether the application is malicious, the graph prediction model uses convolutional and dense neural networks to learn graph-level features for more effective graph prediction. The prediction process is expressed as
[0077]
[0078] where the CM(·) function is a convolutional layer containing a pooling layer and a dense layer, which respectively map and h i to a vector;
[0079] Then the framework performs element-wise multiplication on this vector, takes its average value, and uses the sigmoid function for prediction.
[0080] Advantages of the present invention:
[0081] 1) The present invention first proposes to utilize a graph fusion mechanism to retain complete call information between two languages. In the graph fusion module, a relationship extractor is used to match the keywords defined in the JNI call mechanism, and the relationship between Java classes and the Native part is extracted. The effect of this is to construct a call relationship edge between Java and Native. These relationships and their corresponding control flow graphs are submitted to a multi-relational directed graph generator to construct a merged graph that covers all control flow graphs in the APP.
[0082] 2) Traditional convolutional neural networks need to convert the control flow graph into an intermediate representation (IR) as input, which may lead to the loss of application semantics and time consumption when the application is complex. The present invention vectorizes the nodes of the fused graph and uses a graph neural network for training, improving efficiency and completely retaining the semantics of the application.
[0083] Specifically, in step 3.1, the Gated Graph Neural Network (GGNN) uses the fused multi-relational directed graph (MRDG) as input and performs feature learning through multiple parallel feature extraction channels. Each channel contains a gated graph neural layer, a pooling layer, and a convolutional layer, which can efficiently extract structured information at different levels.
[0084] In step 3.2, the node instructions of the MRDG are converted into a fixed-length digital sequence to ensure that different types of nodes can be trained in a unified format. At the same time, the multi-relational information of the edges is stored in the form of an adjacency matrix, enabling the model to capture the dependencies of cross-language code. In addition, in step 3.3, the gated graph neural network model adopts an Update Gate and a Relevance Gate mechanism, allowing nodes to maintain long-term memory when iteratively propagating information during training and solving the problem of loss of long-sequence information.
[0085] 3) The application of the present invention effectively improves the efficiency of information dissemination, enabling each node to fully utilize the information of its neighbor nodes during parameter update, and accumulating global semantics through multiple time-step iterations to ensure the complete retention of the code semantics of the application. Finally, through the graph-level classification layer, the model can accurately distinguish between benign software and malicious software, achieving efficient and accurate malicious code detection and improving the robustness and generalization ability of the detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 It is a flowchart of cross-language malicious software detection based on a graph neural network proposed by the present invention.
[0087] Figure 2 It is a schematic diagram of the graph feature extraction process proposed by the present invention. Detailed implementation mode
[0088] The present invention will be further described in detail below with reference to the accompanying drawings.
[0089] As Figure 1 , Figure 2 shown, a cross - language malware detection system and method based on a graph neural network specifically include:
[0090] 1) Structure description
[0091] Aiming at solving the technical problems to overcome the deficiencies of the prior art, the present invention provides an Android cross - language malware detection system based on GNN.
[0092] This system first decompiles Android applications to generate control - flow graphs corresponding to two languages, then fuses the two graphs to generate a multi - relational directed graph as the input of the gated graph neural network (GGNN) for learning, extracts the features of the graph containing malicious behaviors, and finally obtains the final prediction result through the test set. Since the context transfer between the two languages is considered, it can effectively detect malicious behaviors hidden in native code and reduce the false positive and false negative rates of cross - language malware detection.
[0093] The technical solution to implement the present invention is a cross - language malware detection technology based on a graph neural network, which consists of three modules: a control - flow graph generation module, a graph fusion module, and a graph learning module, including the following steps. The system structure diagram is as Figure 1 shown.
[0094] The control - flow graph generation module is responsible for pre - processing the application program and generating control - flow graphs for Java code and native code respectively. (Android application programs are mainly developed in Java language and may also contain a small amount of native code. This method can be extended to other languages).
[0095] The graph fusion module is used to extract the relationships between graphs and fuse the two graphs.
[0096] The graph learning module is responsible for the vectorization of nodes in the graph and the classification of control - flow graphs of benign software and malware.
[0097] Furthermore, the specific steps included in the control - flow graph generation module are as follows:
[0098] Step 1.1, construct a dataset of benign samples and malicious samples, screen according to the proportion of cross - language applications in the total number of applications in the actual part, and store the application dataset locally.
[0099] Step 1.2: Decompile each sample in the dataset to obtain the dex file, so file, Manifest file, etc. in it.
[0100] Step 1.3: Filter the Java activities according to the ActivityName defined in the Manifest file, extract the control flow graphs from the so file and dex file respectively, and extract the corresponding files.
[0101] Furthermore, the specific steps included in the graph fusion module are as follows:
[0102] Step 2.1: Extract the relationships between Java internal, Java and native according to the activity files extracted in Step 1.3, and construct three types of edges between the corresponding control flow graphs: JNI call edge, class import edge, parallel edge. Among them, the JNI call edge refers to the relationship of JNI calls from Java code to native code; the class import edge refers to the relationship of class imports between two control flow graphs that both belong to the Java part; the parallel edge refers to the situation where there is no such relationship between the two graphs.
[0103] Step 2.2: Receive the intermediate representations corresponding to the native code and Java code generated in Step 1.3, and combine them with the relationships between the graphs extracted in Step 2.1 to construct an application-level multi-relationship directed graph (Multi-Relationship Directed Graph, MRDG).
[0104] Furthermore, the specific steps included in the graph learning module are as follows:
[0105] Step 3.1: The gated graph neural network uses the vectorized MRDG as the input, and uses multiple parallel feature extraction channels. Each feature extraction channel consists of a gated graph neural layer, a pooling layer, and a convolutional layer connected in sequence.
[0106] Step 3.2: Further, in Step 3.1, the instructions in the nodes of the training sample MRDG are counted in order according to the predefined classification, so that each node is converted from an instruction block to a fixed-length digital sequence; the multi-relationship directed edges in the MRDG are represented in the form of an adjacency matrix.
[0107] Step 3.3. Further, when updating the parameters for each node in the graph during the training process, not only does it receive information from adjacent nodes, but also sends information to adjacent nodes. The gated graph neural network model contains two gates, an update gate and a relevance gate. The update gate is used to control the update of the memory. If two associated nodes are far apart, the state at the last moment will be maintained, having a memory effect, which solves the problem of long sequences. The relevance gate represents the relevance between the state at the previous moment and its candidate value. The activation functions of both the update gate and the relevance gate are sigmoid. After iterating for T time steps, the node embedding set in the current network is used as the final set and input into the graph-level classification layer to predict whether the application is malicious or not. The prediction process includes a fully connected convolutional layer and a density layer, which map the final representation set of the graph into a vector.
[0108] 2) Specific implementation steps
[0109] As Figure 1 shown, the technical solution of the present invention is a cross-language malware detection method based on a graph neural network, which consists of three modules: a control flow graph generation module, a graph fusion module, and a graph learning module. The control flow graph generation module is responsible for pre-setting application programs for the Java part and the native part respectively and generating control flow graphs. It is built on top of Androguard and Radare2. Androguard is used to decompile the apk and perform custom analysis on Android applications. Radare2 is used to analyze and disassemble binary code. Two CFGs can be obtained through these two tools respectively. The graph fusion module is the middle layer, assisting in the combination of two types of graphs, the Java part and the native part. The graph learning module is responsible for classifying the CFGs of benign applications and malware. The graph learning module uses a gated graph neural network (GGNN) to model multiple graph relationships through a combined graph, and the combined graph is abstracted from the multi-relational directed graph (MRDG) constructed in the previous stage.
[0110] The present invention implements a cross-language malware detection system based on a graph neural network. The specific implementation includes the following parts:
[0111] A. Control flow graph generation module
[0112] As Figure 1As shown, this module takes the application as input. In the preprocessing step, an Apktool-based decompiler is used to decompile the APK into three parts, namely the.dex file, the Manifest file, the Res file, and the.so file. The control flow graph generation module uses the Activity analysis component to restrict certain methods in the regular expressions during the Java control flow graph generation process, saving a significant amount of time. In addition, the directory storing the decompiled files is traversed to extract the native code of the application, and then the Native control flow graph generation component (Radare2) is used to generate the CFG for the native part. To prepare for extracting the relationship between the two parts, Dex2Jar is used to collect Java activity information. The detailed process is as follows.
[0113] (1) Activity analysis: Android applications are an Activity-based system. Generating the CFG of the Activities defined in the Manifest file improves efficiency. The Activity analysis component receives the Manifest file to extract all the attributes of the ActivityName for control flow graph generation.
[0114] (2) Control flow graph generation: The Java control flow graph generation component and the Native control flow graph generation component are used to generate the CFG of the Java part and the native part of the app respectively. The Java control flow graph generation component based on the Androguard decompiler receives the ActivityName from the Activity analysis component and uses the CFG generation tool to decompile the APK and build the CFG. In addition, the Native control flow graph generation component receives the.so files decompiled by the decompiler as input and passes them to Radare2 to generate the CFG of the native code.
[0115] (3) Activity extraction: This step uses D2j-dex2jar to extract Activities. D2j-dex2jar is a collection of tools that can operate on the Android Dalvik (.dex) file format and Java's (.class), perform the dex-to-class file format conversion, and generate detailed Java activity information to prepare for extracting the call relationship between the Java part and the native part.
[0116] B. Graph fusion module
[0117] The graph fusion module has two main components, namely the relationship extractor and the multi-relationship directed graph generator. The relationship extractor receives the Java Activities output in the first phase, traverses all the Activities and keywords defined by the JNI call mechanism, and extracts the relationships between Java classes and native files. These relationships and all the generated CFGs will be submitted to the multi-relationship directed graph generator component. Then, the multi-relationship directed graph generator component constructs a merged graph that covers all the separate CFGs of an app, connected by the extracted relationships. The process is detailed as follows.
[0118] (1) Relationship extractor: This step uses the semantic information and structural information of basic blocks to represent the semantic information of nodes in the graph. The multi-relationship directed edges of the CFG represent the call relationships between Java codes or between Java nodes and native nodes. According to the different types of call relationships, three additional types of edges are defined: JNICall, ClassImport, and Parallel.
[0119] JNICall edge. To record the relationship of JNI calls from Java code to native code, the JNICall edge is used to connect the basic blocks in the Java part to the native part.
[0120] ClassImport edge. The directed ClassImport edge is used to connect two Java files with an import relationship.
[0121] Parallel edge. To improve the quality during GNN learning, the Parallel edge is used to represent graphs without the above two relationships.
[0122] (2) Multi-relationship directed graph generator: In the obtained CFGs, a series of code blocks including bytecodes or assembly instructions are used as graph nodes, and the control flow paths are used as edges. Then the JavaCFG and native CFG can be defined respectively as:
[0123]
[0124] where G java and G native represent the sets of control flow graphs generated for the Java part and the native part respectively, where the K CFGs in the Java part: belong to G java , and the L CFGs in the native part: belong to G native . V and E represent the nodes and edges in the control flow graph, that is, basic blocks and control flow edges. Their subscripts J and N represent that they belong to the Java part and the native part respectively.
[0125] The multi - relationship directed graph generator merges these two types of CFGs into a multi - relationship directed graph through the relationship edges defined above. There are three types of edges in the graph, namely JNICall, ClassImport, and Parallel. There are two types of key APIs for calling Java and native functions, including System.LoadLibrary() and import.
[0126] For convenience, abbreviate the edges and key APIs as Edges = {e jni , e imp , e par} and APIs = {interface load , interface imp}. The input of this component is the CFG sets from two parts of G java and G native , as well as the set of Java files F java corresponding to the control flow graph in Java. The output of the component is the final graph G Final after all graphs are merged. The specific process of generating the multi - relationship directed graph is as follows, aiming to fuse all graphs generated from the native part and the Java part:
[0127] 1) Graph fusion stage: Responsible for merging the sub - graph G sub into the final graph G Final . This stage randomly selects a graph from the JavaCFG set and passes it into the sub - graph fusion stage to generate a sub - graph G with sub as the root node. The nodes and edges of G Final and G sub are connected together by the defined Parallel edges.
[0128] 2) Sub - graph fusion stage: Responsible for recursively merging the graph according to the call relationship between the given CFG g java and the API - corresponding graph to generate a sub - graph G java with g sub as the root node. First, the sub - graph fusion stage needs to find the corresponding Java file according to g java , traverse the file and judge the call type according to the given API. This stage considers the following two cases:
[0129] a. In the Java file corresponding to g java , if there is a Java class import, the sub - graph fusion stage obtains g′ java of the imported class, and connects g imp and g′ java with the edge e java . Then g′java Added to G sub and deleted from G java After that, recursively merge all the graphs in G java .
[0130] b. In the corresponding Java file of g java , if there is a JNI call API, in the sub-graph fusion stage, obtain the g' of the native loading library native , and connect g jni and g' java with an edge e native . Then add g' native to G sub . After that, return G sub to terminate the recursion.
[0131] C. Graph learning module
[0132] In the graph learning module, first vectorize the multi-relational directed graphs and represent them as adjacency matrices and node embeddings. Then input them into the multi-relational directed graph neural network to train the malware detection model. This stage has two steps, namely node vectorization and GNN learning. The workflow of the graph learning module specifically has the following two steps:
[0133] (1) Node vectorization: The constructed MRDG cannot be directly input into the learning process, so it needs to go through node vectorization to extract graph features. As Figure 1 shown, the node vectorization (NodeVectorizationComposer, NVC) component is responsible for mapping a basic block v i to a graph node vector h i as the input of the GNN. H i is defined as follows: h i = NVC(v i ), i = 1,..., |V|.
[0134] After comparing the bytecodes and binary files of different platforms, referring to the feature representations provided in the existing work, 20 features (14 for Java code and 6 for native code) are used for vectorization. These features vary little across different underlying platforms, different microprocessor architectures, and different compiler optimization configurations, as shown in Table 1. For each node in the graph, first initialize a 20-dimensional vector i.e., each dimension represents a feature, filled with all zeros. For each code block, i.e., each node in the MRDG, determine the code type in the Java part or the native part, and then count the number of each feature in each basic block and arrange them in order to form a numerical vector.
[0135] Table 1 Basic Block Level Feature Classification
[0136]
[0137] For basic block J1, initialize the embedding as a 20 - dimensional vector and fill it with all zeros. Count the number of instructions in the block according to the instruction classification in the Java part of Table 1, and then arrange the numbers in sequence to obtain a 14 - dimensional vector and copy C s to the first 14 dimensions h0 to h v of h 13 At this time, h 14 to h 20 are still all filled with 0. Finally, arrange these numerical vectors to the corresponding vertex v s in MRDG.
[0138] Each node in MRDG is finally represented by a numerical vector g = <H, E>, where H = {h1, …, h |V|} and E are the vectorized basic blocks and multi - relationship edges. In addition, each vertex h i ∈H represents the initial numerical feature vector, and use G = {g1,.., g n} as the input of the multi - relationship directed graph neural network.
[0139] (2) Multi - relationship Directed Graph Neural Network
[0140] Based on the extracted features, malware detection can be defined as a binary classification problem, that is, learning to judge whether a given application is malicious. Construct a dataset {(x i , y i ) | x i ∈X, y i ∈Y}, i ∈ {1, 2,..., n}, where X represents a series of applications to be learned, Y = {0, 1} n represents an n - dimensional vector, where 1 represents malware and 0 represents benign, and n represents the number of all APKs. The multi - graph g i (V, E) ∈ G, i ∈ 1, 2, …, n, merged in the previous stage represents an APK x i , where G represents a set of graphs generated from all instances, and v i ∈V is a vertex of the multi - relationship directed graph (MRDG), which contains the assembly - level instruction content from the Java part or the native part, and e = (v s , v t ) ∈ E represents a directed edge v s →v t . For each edge, the graph also contains the edge type l e∈ {1, …, k}, where k is the number of edge types. The goal of this module is to vectorize the content of the nodes in the graph generated in the previous stage, input it into a gated neural network, learn a mapping from G to Y, and predict whether a given APK is malicious. Therefore, a malware detection model trained using a Gated Graph Neural Network (GGNN) is used.
[0141] GGNN is a classic spatial domain message passing model based on the Gated Recurrent Unit (GRU). Each time the parameters are updated, each node not only receives information from adjacent nodes but also sends information to adjacent nodes. Compared with the standard RNN, the GRU model contains two gates, an update gate and a relevance gate. The update gate is used to control the update of the memory: if two associated nodes are far apart, the state of the previous moment will be maintained, which gives the network a memory effect and solves the problem of long sequences. The relevance gate represents the relevance between the state of the previous moment and its candidate value.
[0142] Given an embedded graph g i (H, E) ∈ G, the feature vectors H of all nodes are used as the node state vectors
[0143]
[0144] For each node v in the graph g i , it needs to communicate with its neighbor nodes during the parameter update process. The process of node aggregation according to the type and direction of the edges between nodes is as follows.
[0145]
[0146] In particular, is the adjacency matrix composed of the set of edges E and the type l e where e represents a directed edge from node v s to v t and l e represents the type of this edge. being equal to 1 means that node v s , v t is connected by an edge of type l e . are the weights to be learned and b is the bias. The process of repeated unfolding is performed for T steps.
[0147]
[0148] where controls the update information, Control the reset information, and σ(·) represents the logical sigmoid function. The above three expressions show the GRU-like update process, the aggregated information from neighboring nodes, and the preparatory steps for generating the new state of the node.
[0149]
[0150] Among them is the finally updated node state, including the selected forget and the newly generated information to be remembered ⊙ is element-wise multiplication.
[0151] After iterating for T time steps, the set of node embeddings in the network is is the final node representation of the node set V of the graph. Then, the final set is used as the input to the graph-level classification layer to predict whether the application is malicious. The graph prediction model uses convolutional and dense neural networks to learn graph-level features for more effective graph prediction. The prediction process can be expressed as
[0152]
[0153] where the CM(·) function is a convolutional layer containing a pooling layer and a dense layer, which respectively map and h i to a vector.
[0154] Then the framework performs element-wise multiplication on this vector, takes its average value, and uses the sigmoid function for prediction.
Claims
1. A cross - language malware detection system based on graph neural network, characterized in that, It includes a control flow graph generation module, a graph fusion module, and a graph learning module; The control flow graph generation module is responsible for preprocessing the application program and generating control flow graphs for Java code and native code respectively; The graph fusion module is used to extract the relationships between the control flow graphs and fuse the two graphs to obtain a fused graph; The graph learning module learns from the fused graph, is responsible for vectorizing the nodes in the fused graph, and classifying the control flow graphs of benign software and malicious software.
2. Detection method of a cross - language malware detection system based on graph neural network, characterized in that, It includes the following steps; Step 1: Preprocess the application program through the control flow graph generation module and generate control flow graphs for Java code and native code respectively; Step 2: Use the graph fusion module to extract the relationships between the control flow graphs and fuse the two graphs to obtain a fused graph; Step 3: Learn from the fused graph through the graph learning module, which is responsible for vectorizing the nodes in the fused graph and classifying the control flow graphs of benign software and malicious software.
3. The detection method of a cross - language malware detection system based on graph neural network according to claim 2, wherein The specific steps included in Step 1 are as follows: Step 1.1, Construct a dataset of benign samples and malicious samples, screen them according to the proportion of cross-language applications in the total number of applications in the real part, and store the screened sample dataset locally; Among them: Benign samples refer to application programs that comply with the official security specifications; Malicious samples refer to application programs that contain malicious code or harmful behaviors; Step 1.2, Decompile each sample in the dataset to obtain the dex file, so file, and Manifest file therein; Step 1.3, According to the Activity names defined in the Manifest file, filter the Java code control flow Activities in each benign sample and malicious sample, and then extract the control flow graphs of Java code and native code from the so file and dex file respectively, and extract the corresponding activity files.
4. The detection method of a cross - language malware detection system based on graph neural network according to claim 3, characterized in that, In Step 1.2, use an Apktool-based decompiler to decompile the APK, that is, the Android application installation package, into three parts, namely the.dex file, the Manifest file and the Res file, and the.so file; In Step 1.3, by parsing the Activity components defined in the Manifest file, use the regular expression matching method to filter the control flow in the Java code, traverse the directory storing the decompiled files to extract the native code of the application program, and then use the Native control flow graph generation component to generate a control flow graph for the native part; use Dex2Jar to collect Java activity information; Use the Java control flow graph generation component and the Native control flow graph generation component to generate the CFGs of the Java part and the native part of the app respectively. The Java control flow graph generation component based on the Androguard decompiler receives the ActivityName from the Activity analysis component and uses the CFG generation tool to decompile the APK and construct the CFG. The Native control flow graph generation component receives the.so files decompiled by the decompiler as input and passes them to the flow graph generation component to generate the CFG of the native code; Use D2j-dex2jar to extract the Activity, operate on the Dalvik file format of Android and the Java one, perform the file format conversion from dex to class, and generate detailed Java activity files.
5. The detection method of a cross - language malware detection system based on graph neural network according to claim 3, wherein, The specific steps of Step 2 are as follows: Step 2.1, according to the activity files described in Step 1.3, extract the relationships between Java internals and between Java and native code. Corresponding to different relationships between control flow graphs, use a relationship extractor to construct three types of edges, namely JNI call edges, class import edges, and parallel edges; Among them, the JNI call edge refers to the relationship of JNI calls from Java code to native code; the class import edge refers to the relationship of class imports between two control flow graphs belonging to the Java part; the parallel edge means that if there is no such relationship between the two graphs; Step 2.2, receive the control flow graphs corresponding to the Java code and native code generated in Step 1.3, and combine the relationships between the control flow graphs extracted in Step 2.
1. The multi-relationship directed graph generator uses the control flow graphs and the relationships between the control flow graphs to construct an application-level multi-relationship directed graph.
6. The detection method of a cross - language malware detection system based on a graph neural network according to claim 5, characterized in that, During the extraction process of Step 2.1 The relationship extractor receives the Manifest file output by Step 1.2, traverses all Activities and keywords defined by the JNI call mechanism, extracts the relationships between Java classes and native files, and submits the relationships and all the control flow graphs generated in Step 1.3 to the multi-relationship directed graph generator component. Then, the multi-relationship directed graph generator component merges the control flow graphs and the relationships between the control flow graphs to construct a fusion graph, which not only covers all the separated CFG information of an app but also contains the relationship connections between each CFG; The relationship extractor uses the semantic information and structural information of basic blocks to represent the semantic information of nodes in the graph; A basic block is the smallest unit in program execution and contains a set of sequentially executed instructions; the relationship extractor analyzes the instructions in the basic block, extracts the key semantic information, and uses it to construct the feature representation of the node; The multi-relationship directed edges of the CFG represent the call relationships between Java codes or between Java nodes and native nodes. According to the different types of call relationships, three additional types of edges are defined: JNICall, ClassImport, and Parallel; The JNICall edge is used to record the relationship of JNI calls from Java code to native code, and the JNICall edge is used to connect the basic blocks in the Java part to the native part; The ClassImport edge uses a directed ClassImport edge to connect two Java files with an import relationship; The Parallel edge uses a Parallel edge to represent a graph without the above two relationships; In the obtained CFG, the multi-relationship directed graph generator takes a series of code blocks including bytecodes or assembly instructions as graph nodes and the control flow path as edges. The JavaCFG and native CFG are respectively defined as: Among which G java and G native respectively represent the sets of control flow graphs generated by the Java part and the native part, where the K CFGs of the Java part: belong to G java , and the L CFGs of the native part: belong to G native , V represents the nodes in the control flow graph, and E represents the edges in the control flow graph, that is, basic blocks and control flow edges, and their subscripts J and N respectively indicate that they belong to the Java part and the native part.
7. The detection method of a cross - language malware detection system based on graph neural network according to claim 6, characterized in that, The multi-relationship directed graph generator merges two types of CFGs, namely the Java code CFGs and the native code CFGs, into a multi-relationship directed graph through three defined types of edges. In the graph, there are three types of edges: JNICall, ClassImport, and Parallel. There are two types of key APIs for calling Java and native functions, including System.LoadLibrary() and import; Abbreviate the edges and critical APIs as Edges = {e jni , e imp , e par} and APIs = {interface load , interface imp}, the input of this component is the CFG sets from two parts of G java and G native , as well as the set of Java files F java corresponding to the control flow graph in Java. The output of the component is the final graph G Final after all graphs are merged; The specific process of generating the multi-relationship directed graph is as follows, which is used to fuse all the graphs generated from the native part and the Java part: 1) Graph fusion stage: responsible for merging the sub-graph G sub into the final graph G Final . In this stage, a graph is randomly selected from the JavaCFG set and passed into the sub-graph fusion stage to generate a sub-graph G with sub as the root node. The nodes and edges of G Final and G sub are connected together by the defined Parallel edges; 2) Sub - graph fusion stage: responsible for recursively merging the graph according to the call relationship between the given CFGg java and the API - corresponding graph to generate g java as the sub - graph G of the root node sub ; G sub is part of the MRDG. When all sub - graphs are merged into the final graph G Final , G Final is the finally fused MRDG; In the sub - graph fusion stage, it is necessary to find the corresponding Java file according to g java and traverse the file to determine the call type according to the given API. The following two cases are considered in this stage: a. In g java In the corresponding Java file, if there is a Java class import, during the sub-graph fusion stage, obtain the g of the imported class j ′ ava , with edge e imp Connect g java and g j ′ ava , then add g j ′ ava to G sub in it, and delete it from G java . After that, recursively merge all the graphs in G java ; b. In g java In the corresponding Java file, if there is a JNI call to the API, during the sub-graph fusion phase, obtain the g of the native loading library ′ native with edge e jni to connect g java and g ′ native and then add g ′ native to G sub After that, return G sub to terminate the recursion.
8. The detection method of a cross - language malware detection system based on a graph neural network according to claim 2, characterized in that, The specific steps of step 3 are as follows: Step 3.1, The gated graph neural network takes the application-level multi-relationship directed graph MRDG after vectorization as input and uses multiple parallel feature extraction channels. After passing through the feature extraction channels, high-dimensional feature vectors of each node in the MRDG are obtained; Step 3.2, Take the high-dimensional feature vectors of each node in the MRDG generated in step 3.1 as input for MRDG graph embedding representation; the feature vectors of each node in the MRDG are converted into a fixed-length digital sequence to ensure that different types of code snippets are learned in a unified format; the multi-relationship directed edges in the MRDG are represented in the form of an adjacency matrix; Step 3.3, Train and learn the gated graph neural network model for the MRDG graph embedding representation generated in step 3.2; The gated graph neural network model contains two gates, an update gate and a relevance gate; the update gate is used to control the update of memory; the relevance gate represents the correlation between the previous state and its candidate value. The activation functions of the update gate and the relevance gate are both sigmoid. After iterating for T time steps, the set of node embeddings in the current network is used as the final set and input into the graph-level classification layer to achieve benign software and malware prediction. The prediction process includes fully connected convolutional layers and density layers, which map the final representation set of the graph to a vector to achieve the classification of the application software as benign or malicious.
9. The detection method of a cross - language malware detection system based on a graph neural network according to claim 8, characterized in that, The specific content of step 3.2 is: First, the multi-relationship directed graph is vectorized, and the application-level multi-relationship directed graph after vectorization is represented as an adjacency matrix and node embeddings. Then, the adjacency matrix and node embeddings are input into the multi-relationship directed graph neural network to train the malware detection model, that is, the gated graph neural network model in step 3.3; There are two steps in the graph learning stage, namely node vectorization and GNN learning (1) Node vectorization: The node vectorization component is responsible for converting a basic block v i That is, after feature extraction in step 3.1, the high-dimensional feature vectors of each node in the MRDG are mapped to a graph node vector h i As the input of the GNN, H i is defined as follows: h i = NVC(v i ), i = 1, …, |V|; Each node in MRDG is finally represented by a numerical vector as g = <H, E>, where H = {h1, …, h |V|} and E are the vectorized basic blocks and multi-relation edges. In addition, each vertex h i ∈ H represents an initial numerical feature vector, and G = {g1, .., g n} is used as the input of the multi-relation directed graph neural network; (2) Multi-relational directed graph neural network: Based on the extracted feature G, malware detection is defined as a binary classification problem, that is, learning to judge whether a given application is malicious, and constructing a data set {(x i ,y i )|x i ∈X,y i ∈Y},i∈{1,2,...,n}, where X represents a series of applications to be learned, and Y = {0,1} n represents an n-dimensional vector, where 1 represents malware, 0 represents benign, n represents the number of all APKs, and the multi-graph g i (V,E)∈G,i∈1,2,…,n, merged in the previous stage represents an APK x i , where G represents a set of graphs generated from all instances, and v i ∈V is a vertex of the multi-relational directed graph, which contains the assembly-level instruction content from the Java part or the native part, and e=(v s ,v t )∈E represents a directed edge v s →v t , and for each edge, the graph also contains an edge type l e ∈{1,…,k}, where k is the number of edge types.
10. The detection method of a cross - language malware detection system based on a graph neural network according to claim 6, characterized in that, In step 3.3: Given an embedded graph g i (H, E) ∈ G, and use the eigenvectors H of all nodes as the node state vectors For each node v in graph g i , communication with neighbor nodes is required during the parameter update process. The process of node aggregation based on the type and direction of the edges between nodes is as follows: is an adjacency matrix composed of a set of edges E and a type l e where e represents a directed edge from node v s to v t and l e represents the type of this edge, equal to 1 means that nodes v s , v t are connected by an edge of type l e ; are the weights to be learned, b is the bias, and the process of repeated unfolding is performed for T steps; Among them Control update information Control reset information, and σ(·) represents the logical sigmoid function; wherein is the finally updated node state, including the option to forget and the newly generated information to be remembered ⊙ is element-wise multiplication; After iterating for T time steps, the set of node embeddings in the network is is the final node representation of the set of nodes V of the graph. Take the set of node embeddings in the current network as the final set. Then, use the final set as the input to the graph-level classification layer to predict whether the application is malicious. The graph prediction model uses convolutional and dense neural networks to learn graph-level features. The prediction process is represented as Among them, the CM(·) function is a convolutional layer that includes a pooling layer and a dense layer, which respectively map and h i to a vector; Then the framework performs element-wise multiplication on the vector, takes its average value, and uses the sigmoid function for prediction.