Mobile Platform Malware Detection Method and Device Based on Multimodal Information Fusion
Through the multimodal information fusion method, the grayscale graph, interface call sequence and function call graph features of binary applications are extracted and fused, and heterogeneous selection and robust fusion network are built, which solves the problem of malware detection being vulnerable to single modal attacks and improves the accuracy and robustness of detection.
Patent Information
- Application Number
- CN202310136086.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-20
AI Technical Summary
Existing malware detection methods are vulnerable to single modal attacks, resulting in the failure of the model and the inability to effectively identify malware.
By using the multimodal information fusion method, the features of the grayscale graph, interface call sequence and function call graph of binary applications are extracted, and the heterogeneous selection network and robust fusion network are constructed by using Grad-CAM, GraphSAGE and Markov chain technologies to improve detection robustness.
Effectively resisting single-modal attacks, improving the accuracy and robustness of malware detection, capturing multiple aspects of the application, and enhancing the accuracy of decision results.
Smart Images

Figure CN116226852B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the analysis technology of binary software, multi-modal feature fusion technology, and malware detection technology, and particularly relates to a malware detection and analysis method for a mobile platform, and provides a malware detection method and device for a mobile platform based on multi-modal information fusion. Background Art
[0002] With the rapid development of mobile system platforms and the mobile application ecosystem, the security risks of mobile systems are increasing. Mobile applications have become the main targets of malware attacks. With the rapid development of mobile platforms, the attack methods of malware have also seen new expansions on mobile platforms. On the Android platform, according to the analysis of a total of 6 million applications in Google Play and domestic app markets by Wang et al. (Wang H Y, Liu Z, Liang J Y, et al. Beyond google play: a large-scale comparative study of Chinese Android app markets. In: Proceedings og the Internet Measurement Conference (IMC), Boston, 2018. 293-307.), approximately 12.3% of the applications in the Chinese Android app market were reported as malicious programs by at least 10 malware detection tools. In addition, a research team from the Institute of Information Engineering, Chinese Academy of Sciences (Chen K, Wang X Q, Chen Y, Wang P, Lee Y, Wang X F, Ma B, Wang A H, Zhang Y J, Zou W. Following devils footprints: Crossplatform analysis of potentially harmful libraries on Android and iOS. In: Proc.of the 37th IEEE Symp.On Security and Privacy, Ser. (S&P 2016). 2016.) studied the malicious code in third-party libraries in the iOS app market and found that 23 potential iOS malware libraries with a total of 706 variants were included in 14,000 iOS applications. Zhou et al. (Zhou Y, Jiang X. Dissecting Android malware: Characterization and evolution. In: Proc.of the 2012 IEEE Symp.on Security and Privacy. IEEE, 2012. 95-109.) summarized the forms, classifications, and evolutions of mobile terminal malware and classified malware into nine types, such as repackaging, updating, induced downloading, privilege escalation, and remote control.
[0003] With the widespread use of machine learning methods, attackers have started to modify malicious programs. While adding malicious code, they do not change the feature description of the malware, enabling the malware to bypass detection based on machine learning algorithms. Zhao et al. (Zhao K F, Zhou H, Zhu Y L, Zhan X, Zhou K, Li J F, Yu L, Yuan W, Luo X P. Structural Attack against Graph Based Android Malware Detection. In Proc. of the 2021 ACM SIGSAC Conference on Computer and Communications Security. ACM, New York, NY, USA, 3218–3235.) used structural attacks to perform operations such as adding, deleting, and modifying nodes and edges based on the graph structure features extracted from the program. Nearly 100% of the Android platform malware samples can escape detection within 500 operations.
[0004] In response to the above problems, researchers have designed different malware detection methods based on the analysis of mobile platform malware. The existing malware detection methods are mainly divided into two types: static detection methods and dynamic detection methods. Static detection methods perform detection before the program runs. Their advantages are low energy consumption and low risk, but they are affected by obfuscation and encryption attacks. Dynamic detection methods perform detection during the program's operation. Dynamic detection has high requirements for real-time performance and the running environment, requires support from mobile devices, takes longer time, but has a higher accuracy rate compared to static detection methods. With the continuous development of various detection methods, malware methods based on machine learning have gradually become a research hotspot. Based on static detection, by extracting features describing malware, using vectors of fixed dimensions to represent malware, then training on known labeled samples with existing machine learning algorithms and constructing a classifier, and finally predicting and judging the software to be tested.
[0005] The malware detection method based on multi-modal is an extension of the machine learning detection method. It fuses different modal features of the software to be detected, extracts a unified feature representation, and uses machine learning algorithms for classification. Traditional multi-modal fusion models are vulnerable to attacks on a single modality, especially open-source mobile platforms such as Android and OpenHarmony are more vulnerable to attacks. Attacks on a single modality can interfere with the correct modality that has not been attacked and cause the model to fail, unable to meet the requirements of malware detection. Summary of the Invention
[0006] The object of the present invention is to provide a more robust malware detection method and device based on multimodal information fusion for single-modal malware attacks, which can counter single-modal attacks against malware detection. Multimodal fusion uses the feature information extracted from multiple modalities for fusion, so as to obtain a better recognition effect than single-modal.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A mobile platform malware detection method based on multimodal information fusion, comprising the following steps:
[0009] 1) For the binary application to be detected, extract its binary sequence and generate a grayscale image, extract its interface call sequence, and extract its function call relationship and generate a function call graph and a function control flow graph. These three extraction results correspond to the grayscale image modality, the interface call sequence modality, and the graph structure modality in sequence;
[0010] 2) Extract image features from the grayscale image, extract call sequence features from the interface call sequence, and extract global graph features from the function call graph and the function control flow graph;
[0011] 3) A feature set is composed of the image features, the call sequence features, and the global graph features. Input this feature set into the heterogeneous selection network, and output the results of the attacked probability of each modality;
[0012] 4) Construct a robust fusion network based on the fusion network. Use the fusion network to fuse the image features, the call sequence features, the global graph features, and the vector output by the heterogeneous selection network, and output the fusion result; then use the multimodal information fusion network to fuse the vector output by the heterogeneous selection network and the result output by the robust fusion network, and output the prediction vector;
[0013] 5) Dimensionally reduce and normalize the prediction vector to obtain the predicted value of the malware.
[0014] Preferably, the step of extracting the binary sequence and generating the grayscale image in step 1) includes:
[0015] Extract the binary byte stream of the binary application, and determine the width of the grayscale image according to the size of the binary byte stream;
[0016] Take 8-bit binary data as a group, convert the value of the 8-bit binary data into a grayscale value of 0-255, and convert the binary byte stream into grayscale image data in pixels to generate a grayscale image.
[0017] Preferably, the step of extracting the interface call sequence in step 1) includes:
[0018] Perform reverse analysis on the binary application and record the results of the reverse analysis;
[0019] Extract the instruction code sequence therefrom according to the results of the reverse analysis, and retain the interface call instructions in the instruction code sequence to obtain an interface call sequence;
[0020] For each interface call instruction in the interface call sequence, extract its call information, including the family name and package name to which the interface belongs in the interface call, and the call content of the interface call.
[0021] Preferably, the steps of extracting the function call relationship and generating the function call graph and function control flow graph in step 1) include:
[0022] Extract the function call sequence of the binary application, and generate a function call graph according to the function call relationship in the function call sequence;
[0023] Determine whether the function type of each node in the function call graph is an external function or a local function; if it is an external function, extract the function name in the external function node to form a set of external function names; if it is a local function, for the instruction sequence contained in the local function node, extract the basic block;
[0024] Generate the function control flow graph of each local function according to the jump relationship between the basic blocks in each local function.
[0025] Preferably, the steps of extracting image features for the grayscale image in step 2) include:
[0026] Use the Grad-CAM network to generate a heat map according to the gradient information of the grayscale image;
[0027] Overlay the generated heat map with the grayscale image to generate a new heat map;
[0028] Process each pixel in the new heat map using a brightness threshold to extract the highlighted image in the new heat map;
[0029] Record the position of the highlighted image pixels relative to the pixel space as the image features.
[0030] Preferably, the steps of extracting call sequence features for the interface call sequence in step 2) include:
[0031] Count the family names and package names that appear in all interface call instructions of the interface call sequence to generate a set of family names and a set of package names;
[0032] According to the set of family names, for each pair of adjacent interface call instructions in the interface call sequence, count the number of times the family to which the interface call instruction belongs after the call is the same as the family to which the interface call instruction belongs before the call, and the number of times the family to which other interface call instructions belong after the family to which the interface call instruction belongs before the call;
[0033] According to the set of package names, for each pair of adjacent interface call instructions in the interface call sequence, count the number of times the package to which the interface call instruction belongs after the call is the same as the package to which the interface call instruction belongs before the call, and the number of times the package to which other interface call instructions belong after the package to which the interface call instruction belongs before the call;
[0034] According to the set of family names, the set of package names, and the above two statistical counts, construct a Markov chain representing the interface call sequence. This Markov chain consists of a state set and state transition probabilities. Each state in this state set represents a family name or a package name, and the state transition probability refers to the probability of transitioning from one state to another;
[0035] Extract the top pre-set number of nodes with the most occurrences in the Markov chain to form a filtered Markov chain as the call sequence feature.
[0036] Preferably, the steps of extracting global graph features for the function call graph and the function control flow graph in step 2) include:
[0037] For all local functions included in the function call graph, use a multi-layer GraphSAGE model to generate feature vectors for each basic block node in the function control flow graph within each local function. Each layer of GraphSAGE uses a self-transfer function and a message-passing function to learn and update the feature vectors of each layer from the feature vectors generated in the previous layer and the feature vectors of other neighbor nodes; after being processed by the entire layer of the GraphSAGE model, obtain the total feature vectors;
[0038] Use an aggregation model to perform a max pooling operation on the total feature vectors of all nodes in the function control flow graph corresponding to each local function to generate graph vectors;
[0039] For all external functions included in the function call graph, use a one-hot encoding method to generate the encoding corresponding to the external function name and map it into a vector space to obtain the initial feature vectors of the external functions;
[0040] Input the above graph vectors and initial feature vectors into the downstream encoding layer of the GraphSAGE model. For each node in the function call graph, use a multi-layer GraphSAGE model to update the feature vectors of each node to obtain a new total feature vector;
[0041] The maximum pooling operation is performed on the new total feature vectors of all nodes in the function call graph using an aggregation model to obtain the global graph features.
[0042] Preferably, in step 3), the step of inputting the feature set into the heterogeneous selection network and outputting the vector of the probability of being attacked for each modality includes:
[0043] Input the feature set into the heterogeneous selection network. By minimizing the cross-entropy loss, the output is a vector of several terms. The last term represents the probability that none of the modalities are attacked, and the other terms except the last one represent the probability that each modality is attacked.
[0044] Preferably, in step 4), the fusion network is used to perform fusion based on the shallow neural network NN, which includes several fusion operations, and each fusion operation is used to exclude a certain attacked modality.
[0045] Preferably, in step 5), the prediction vector is passed through a multi-layer standard fully connected layer to reduce the dimension to one dimension; then the sigmoid activation function is used to limit the output scalar value range within [0, 1], and finally the predicted value of the malware is obtained after normalization.
[0046] A mobile platform malware detection device using the above method includes:
[0047] A modality generation module, which is used to extract the binary sequence of the binary application to be detected and generate a grayscale image, extract its interface call sequence, and extract its function call relationship and generate a function call graph and a function control flow graph. These three extraction results correspond to the grayscale image modality, the interface call sequence modality, and the graph structure modality in sequence;
[0048] A feature extraction module, which is used to extract image features for the grayscale image, extract call sequence features for the interface call sequence, and extract global graph features for the function call graph and the function control flow graph;
[0049] A heterogeneous selection module, which is used to use the heterogeneous selection network to output the results of the probability of being attacked for each modality according to the feature set composed of image features, call sequence features, and global graph features;
[0050] A robust fusion module, which is used to use the fusion network of the robust fusion network to fuse the image features, call sequence features, global graph features, and the vector output by the heterogeneous selection network, and output the fusion result; and use the multi-modal information fusion network to fuse the vector output by the heterogeneous selection network and the result output by the robust fusion network, and output the prediction vector;
[0051] A malware prediction module, which is used to reduce the dimension and normalize the prediction vector to obtain the predicted value of the malware.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0053] 1) The present invention represents mobile platform applications through multiple modalities, and uses different feature representation methods to represent the features of mobile platform applications, effectively capturing multi-faceted information of application files;
[0054] 2) The present invention uses a malware detection device for multi-modal information fusion, establishes a network and a modal fusion method capable of processing and correlating information from multiple modalities, provides more information for model decision-making, and can effectively improve the accuracy of the overall decision result. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a schematic diagram of the mobile platform malware detection method based on multi-modal information fusion of the present invention;
[0056] Figure 2 is a schematic diagram of the architecture of the mobile platform malware detection method based on multi-modal information fusion of the present invention;
[0057] Figure 3 is an example diagram of the grayscale image modality after extraction;
[0058] Figure 4 is an example diagram of the interface call sequence modality after extraction;
[0059] Figure 5 is an example diagram of the function call graph after extraction;
[0060] Figure 6 is a schematic diagram of the grayscale image modality generation and feature extraction process;
[0061] Figure 7 is a schematic diagram of the interface call sequence modality generation and feature extraction process;
[0062] Figure 8 is a schematic diagram of the function call graph and control flow graph modality generation and feature extraction process;
[0063] Figure 9 is a schematic diagram of the malware prediction process based on heterogeneous selection network and robust fusion strategy;
[0064] Figure 10 is a schematic diagram of the structure of the mobile platform malware detection device based on multi-modal information fusion of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0065] The present invention will be further described below with reference to the accompanying drawings through embodiments, but the scope of the present invention is not limited in any way.
[0066] A malware detection method for mobile platforms based on multi-modal information fusion proposed by the present invention takes the binary program package of the mobile platform as the analysis object, and includes the following steps:
[0067] 1) Modal generation.
[0068] The modal formal representation of the binary application p is <M1, M2, …, M k >>, where each M i (i ∈ [1, k]) represents a modal representation of the binary application. In the present invention, a mode represents a method of presenting a binary application, which can be used for an independent task or a group of closely related tasks. Specifically, the present invention selects three of them to represent the binary application, and specifically represents the binary application p as where I represents the grayscale image modality, based on the grayscale image; represents the interface call sequence modality, based on the interface call sequence; represents the graph structure modality, based on the function call graph and the function control flow graph.
[0069] For a binary file, its most intuitive representation is its binary content, presented in the form of a binary data sequence. The grayscale image modality can fully represent the binary data and facilitate the processing of the modality to adapt to machine learning tasks. The application program realizes the program function by sequentially calling different functions. The interface call sequence modality can fully represent the order in which the application program sequentially calls functions during execution. Through the sequence pattern in machine learning, the function call order rule in the malware can be found, and further, the binary application can be predicted for malware. The function call graph is an extended representation of the interface call sequence. Through a richer graph structure, on the basis of the function call order, the call relationship between functions is represented; at the same time, for the local functions whose code can be obtained from the binary application file, the function control flow graph is further used for representation, which can fully describe the function call relationship and branch process during the execution of the application program.
[0070] The information between the modalities is mainly complementary, but not independent of each other. There is redundant information between any two modalities, that is The redundant information between the modalities helps to detect the modalities with inconsistent manifestations due to attacks based on the existing modal information, so as to counter the single-modal attack against malware detection.
[0071] 1-1) Extract the binary sequence in the binary application package and convert it into the grayscale image modality I.
[0072] 1-1a) Read the binary byte stream of the binary application p; determine the width of the grayscale image based on the size of the binary byte stream; the criteria are as shown in the following table:
[0073] Table 1 Comparison Table of Grayscale Image Width and Binary File Size
[0074] File size <10k 10-30k 30-60k 60-100k 100-200k 200-500k 500-1000k >1000k Image width 32px 64px 128px 256px 384px 512px 768px 1024px
[0075] 1-1b) Take 8-bit binary data as a group, discard the data at the end of the binary data that is less than 8 bits, convert the value of the 8-bit binary data to a grayscale value between 0 and 255, and store it in pixels in the generated grayscale image Ip.
[0076] 1-1c) If the generated pixel is the last pixel of the current pixel row, that is, the horizontal position of the generated pixel is the same as the image width, then save the next generated grayscale pixel in the next row of the current pixel row; otherwise, continue to execute 1-1b).
[0077] 1-1d) If the generated pixel is the last 8 bits of the current binary byte stream, then determine whether the generated grayscale pixel is the last pixel of the current pixel row. If it is, save the data of the current pixel row; otherwise, discard the current pixel row; finally, save the generated grayscale image as I p 。
[0078] 1-2) Extract the interface call sequence from the binary application package and convert it into an interface call sequence mode
[0079] 1-2a) Use the IDA reverse engineering tool to perform reverse analysis on the binary application package, and record the reverse analysis result of the application p as p r 。
[0080] 1-2b) According to the reverse analysis result, extract the instruction code sequence therein, retain the interface call instructions in the instruction code sequence, and represent the extracted interface call sequence as where l is the number of interface calls in the call sequence.
[0081] 1-2c) For each interface call instruction C in the interface call sequence i (i ∈ l), extract the call information and formally represent it as <f i , p i , c i >, where f i is the name of the family to which the interface belongs in the interface call, p i is the name of the package to which the interface belongs in the interface call, and c iFor the call content of the interface call, obtain the interface call sequence mode
[0082] 1-3) Extract the function call relationships in the binary application package and convert them into a graph structure mode
[0083] Among them, represent the binary application p as a function call graph between functions (Steps 1-3a) to 1-3d)), represent the two different types of functions included in the binary application p using the function name (Step 1-3f)) and the function control flow graph within the function (Steps 1-3g) to 1-3j)) respectively
[0084] 1-3a) According to the call content c extracted from the instruction C in Step 1-2c i Analyze and record the calling function of the call content as the callee, and the function that executes the call instruction as the caller i
[0085] 1-3b) Obtain the function call sequence: Summarize the callers and callees extracted in Step 1-3a), and record each caller or callee as a function node. The function node set is represented by where m is the number of nodes, and each node represents a function in the application p
[0086] 1-3c) Obtain the function call relationships: According to the call relationships of the callers and callees extracted in Step 1-3a), record each call relationship as an edge connecting two function nodes. The edge set is represented by ε p Each edge E in the edge set ε p can be represented as E = (N p , N i )(1 ≤ i, j ≤ m), representing the call relationship between two functions, that is, N j function calls N i function j
[0087] 1-3d) Combine the function node set extracted in Step 1-3b) and the edge set ε p extracted in Step 1-3d) to obtain the function call graph between functions of the application p, and denote this function call graph as
[0088] 1-3e) According to the function node set extracted in Step (1-3b) Function types in the function can be divided into two different types: external functions and local functions Wherein, an external function is a system function or library function provided by the mobile system platform; a local function is a function written and designed by a software developer. If the function type of the node is an external function, execute step 1-3f); if the function type of the node is a local function, execute steps 1-3g) to 1-3j).
[0089] 1-3f) Extract the function name in the external function node and record it as the information representation of the node. Finally, all the external function names are summarized and recorded as the external function name set
[0090] 1-3g) Extract the instruction sequence contained in the local function node. According to the specific content and execution order of the instructions in the function, the function can be divided into multiple code basic blocks. The code in the basic block is a sequential execution structure, and the jump relationship between the basic blocks represents the execution relationship between the codes.
[0091] 1-3h) Summarize the code basic blocks extracted in step 1-3g), record each basic block as a node, and use the basic block node set as express, Where n is the number of basic blocks, and each node represents a basic block of code in the function.
[0092] 1-3i) According to the jump relationship of the basic block extracted in step 1-3g), each jump relationship is recorded as an edge connecting two basic block nodes, and the basic block edge set is used express, Basic Block Edge Set Each edge in Can be expressed as K = (V i ,V j )(1≤i,j≤n), represents the control flow path between two basic blocks, namely V i Basic block jumps to V j Basic blocks.
[0093] 1-3j) The basic block node set extracted in step 1-3h) and the basic block edge set extracted in step 1-3i) Combined, we get the function control flow graph within the function, which is recorded as Function call graph And the function control flow graph Graph structure mode
[0094] 2) Feature extraction.
[0095] According to the three modalities generated in step 1), different features are used to represent the modalities respectively:
[0096] 2-1) For the binary grayscale image I extracted in step 1-1) p , generate an image feature representation The specific implementation method is as follows:
[0097] 2-1a) Use Grad-CAM (Gradient-weighted Class Activation Mapping, Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh and Dhruv Batra, "Grad-cam: Visual explanations from deep networks via gradient-based localization", Proc. of the IEEE international conference on computer vision, pp.618-626, 2017.) to generate a heatmap H corresponding to the grayscale image through the gradient information on the convolutional layer I , which can highlight the image regions that have a greater impact on the malware detection task.
[0098] 2-1b) Superimpose the heatmap generated in step 2-1a) on the original image to form a new heatmap
[0099] 2-1c) Process each pixel in the new heatmap using a brightness threshold to extract the highlighted image in the new heatmap.
[0100] 2-1d) Record the position of the highlighted pixels relative to the pixel space as the feature representation of the grayscale image, and record this feature representation as
[0101] 2-2) For the interface call sequence generated in step 1-2) Generate a call sequence feature representation
[0102] 2-2a) According to each interface call instruction node C of the interface call sequence in step 1-2c) i (i∈n) = <f i , p i , c i >, count the family names f that appear in all nodesi , record the set composed of family names as where o is the statistical number of family names, represents a specific family name, indicating the same family name f that appears in all nodes i . At the same time, according to each node C in step 1-2c) i (i ∈ n) = <f i , p i , c i >, count the package names p that appear in all nodes i , and record the set composed of package names as where m is the statistical number of package names, represents a specific package name, indicating the same package name p that appears in all nodes i .
[0103] 2-2b) According to the interface call sequence extracted in step 1-2b) For each pair of adjacent interface calls C i →C j (1 ≤ i < n; j = i + 1), count the number of times O that the family f i to which the called interface C i (or ) belongs is called after the family f j to which the called interface C j (or ), and the number of times O that the family f ij to which the called interface C i belongs is called after the family f i (or ) and the family f j to which other interfaces except the interface C k (or ) belongs is called ik . Similarly, according to the set of package names generated in 2-2a), count the number of times O that another package name p i (or ) appears after a certain package name p j (or ), and the number of times O that other package names p ij appear after the package name p i (or ) k (or ) is called ik .
[0104] 2-2c) According to the name set in step 2-2a) Based on the statistical count O in step 2 - 2b), a Markov chain representing the interface call sequence can be constructed. Where is the state set in the Markov chain, where q is the statistical quantity of the state names. Each of these states represents a family name or a package name. Denotes state Converting to The probability, by calculating the number of times O that state S i is followed by state S j appears, and then dividing by all states, that is ij Divided by all states, that is
[0105] 2 - 2d) Since the package names and family names of the functions defined by developers may be irregular, resulting in an excessive number of nodes in the Markov chain, it is necessary to extract the top 50 nodes with the highest occurrence times in the Markov chain and discard the other nodes. The filtered Markov chain is denoted as As the final feature representation.
[0106] 2 - 3) For the function call graph extracted in step 1 - 3d) and the function control flow graph extracted in step 1 - 3j) The global graph feature vectors are obtained through training with a graph neural network (GNN).
[0107] 2 - 3a) Adopt a simplified GraphSAGE model (Will Hamilton, Zhitao Ying, Jure Leskovec. Inductive representation learning on large graphs[C]. Advances in Neural Information Processing Systems. Long Beach, CA, USA, 2017: 1024–1034.) to generate the feature vectors of each basic block (node) in the function control flow graph Specifically, as described in step 1 - 3e), the function call graph of the binary program The nodes in Contain Local functions, denoted by For Each node in The feature vector generated by the t - th layer of the GraphSAGE model can be expressed as Where d Trepresents the dimension of the node features output by the last layer, i.e., layer T. Each layer of GraphSAGE uses the self-transition function f node and the message passing function f message to learn and update the node feature vectors of each layer from the feature vectors in the previous layer and the feature vectors of other neighbor nodes. The node feature vector generated by the t-th layer of GraphSAGE can be expressed as:
[0108]
[0109] where σ represents the activation function ReLu(x) = max(0, x); represents the set of all neighbor nodes of node v in the function control flow graph ; and are the model parameters in the two functions f node and f message respectively; after passing through T layers of GraphSAGE, the obtained feature vector is
[0110] 2-3b) To learn the graph vector of the function control flow graph corresponding to the local function, a pooling model is used to perform max-pooling operation on the feature vectors of all nodes in each graph:
[0111] 2-3c) For the external functions included in , the encoding corresponding to the external function name is generated using the one-hot encoding method and further mapped into a vector space with dimension d, so as to obtain the initial feature vectors of all external functions
[0112] 2-3d) Taking the graph vector obtained in step 2-3b) and the feature vector corresponding to the one-hot encoding obtained in step 2-3c) as the initial feature vectors and inputting them into the downstream encoding layer. For each node in the function call graph K layers of GraphSAGE are used to update the feature vectors of each node. Among them, the feature vector of node n obtained in the k-th layer of GraphSAGE is as follows:
[0113]
[0114] where σ represents the activation function ReLu(x) = max(0, x); represents the function call graph The set of all neighbor nodes of node n; and They are f node and f message Model parameters in both functions.
[0115] 2-3e) Use the aggregation model to perform the maximum pooling operation on the feature vectors of all nodes in the function call graph and calculate Global graph feature vector where d K Represents the dimension of the node features output by the last layer, i.e., the K layer.
[0116] 3) Using a heterogeneous selection network to identify the attacked feature representation that is inconsistent with other features, the specific steps include:
[0117] 3-1) The feature set extracted from multiple modalities is defined as z = [z1, z2, …, z k ], where z i =g i (x i ), g i Represents the feature extraction function. Specifically, g = 3, where k1 represents step 2-1), and z1 represents the feature representation extracted by the grayscale image modality. g2 represents step 2-2), z2 represents the feature representation extracted from the interface call sequence g3 represents steps 2-3), z3 represents the function call graph And the function control flow graph Generated feature representation
[0118] 3-2) Using a heterogeneous selection network o to identify inconsistent elements under attack, including the following steps:
[0119] The feature set z extracted in step 3-1) is used as the input of the heterogeneous selection network o, and the heterogeneous selection prediction is performed by minimizing the following cross entropy loss:
[0120]
[0121] in, Indicates that from x i After being attacked The output o(z) of the heterogeneous selection network is a vector of size k+1, where the i-th item (i∈k) represents the probability of the i-th mode being attacked, i.e., z i The probability that the features of the other modalities are inconsistent; the k+1th term represents the probability that all modalities are not attacked.
[0122] 4) Use a robust fusion strategy to fuse multiple modalities, which specifically includes the following steps:
[0123] Fusion requires a multi-modal information fusion network, which can be expressed as f: where x = {x1,..., x k} represents the k input modalities; y represents the output prediction result; specifically, k = 3, where x1 represents the grayscale image I p I generated in step 1-1); x2 represents the interface call sequence extracted in step 1-2 x3 represents the function call graph extracted in step 1-3 and the function control flow graph
[0124] Traditionally, using the multi-modal information fusion network f: The defense performance P against a single-modal attack is predicted * is where E is the expected value; Indicates that the input x and output y are sampled from the distribution sampled; is the loss function; Indicates the attack on a certain modality i ∈ k; f(x i + δ, x -i ) is the prediction result with the δ attack behavior imposed on the i-th modality and the other modalities not under attack as the input; x -i is the set of x after removing the i-th one. However, this prediction result is not ideal. Therefore, the present invention makes further improvements by fusing the output results of the heterogeneous selection network o and the robust fusion network f robust to be able to give an ideal prediction result, and the specific steps are as follows.
[0125] 4-1) Fuse the features extracted in step 2), and use the robust fusion network f robust , where, represents the fusion network; the fusion network Integrates the vector output by the heterogeneous selection network o in step 3-2) into the robust fusion network and consists of k + 1 fusion operations, e = e1, e2,..., e k+1 , and each fusion operation is used to exclude a certain modality under attack:
[0126]
[0127] where, represents the concatenation operation; NN represents a shallow neural network.
[0128] 4-2) Use a multi-modal information fusion network to fuse the results of the heterogeneous selection network o and the robust fusion network f robust to obtain the output After improvement, the defense performance P of the multi-modal information fusion network f against single-modal attacks can be updated and expressed as where the output of the heterogeneous selection network o serves as a parameter of the robust fusion network
[0129] 5) Malware prediction, the steps of which include:
[0130] 5-1) Pass the prediction vector z output in 4-2 output through a multi-layer standard fully connected layer to gradually reduce the output dimension to 1
[0131] 5-2) Adopt the sigmoid activation function to limit the output scalar value range within [0,1], and finally obtain the normalized malware prediction value m = sigmoid(MLP(z output ))
[0132] Based on the same inventive concept, the present invention also provides a mobile platform malware detection device adopting the above method, which includes:
[0133] A modality generation module for converting the mobile platform application to be detected into multiple modality representations; specifically, for the binary application to be detected, extract its binary sequence and generate a grayscale image, extract its interface call sequence, and extract its function call relationship and generate a function call graph and a function control flow graph, and these three extraction results correspond to the grayscale image modality, the interface call sequence modality, and the graph structure modality in sequence;
[0134] A feature extraction module for extracting features in multiple modality representation methods and converting them into machine-recognizable feature representations; specifically, for the grayscale image, extract image features, for the interface call sequence, extract call sequence features, and for the function call graph and the function control flow graph, extract global graph features;
[0135] A heterogeneous selection module that uses a heterogeneous selection network to identify feature representations that are attacked and inconsistent with other features; specifically, use a heterogeneous selection network to output the results of the attacked probability of each modality according to the feature set composed of image features, call sequence features, and global graph features;
[0136] A robust fusion module, according to the output of the heterogeneous selection module, performs feature fusion on the features extracted from different modalities in the feature extraction module; specifically, it uses the fusion network of the robust fusion network to fuse the image features, call sequence features, global graph features, and the vector output by the heterogeneous selection network, and outputs a fusion result; and uses a multi-modal information fusion network to fuse the output result of the heterogeneous selection network with the result output by the robust fusion network, and outputs a prediction vector.
[0137] A malware prediction module calculates the probability that the software to be detected is malware, specifically by reducing the dimension and normalizing the prediction vector to obtain a prediction value of malware.
[0138] As Figure 2 shown, the mobile platform malware detection method and device based on multi-modal information fusion takes a binary program file as input, and includes five modules: a modality generation module, a feature extraction module, a heterogeneous selection module, a robust fusion module, and a malware prediction module. The method and device provided by the present invention will be described below in combination with each module.
[0139] The modality generation module includes three modality generation steps, where:
[0140] The grayscale image modality represents the binary sequence of the application program file using grayscale image features. The specific method is:
[0141] (1) Through the input binary application p, using the binary sequence contained therein, generate a grayscale image Ip;
[0142] The interface call sequence modality analyzes the assembly instruction sequence after reverse analysis of the binary file, and performs reverse analysis on the program file through the IDA reverse tool. The specific method includes:
[0143] (2) According to the binary application p in step (1), perform reverse analysis on the program file through the IDA reverse tool to obtain an analysis result p r ;
[0144] (3) Extract the interface call sequence from the analysis result
[0145] The function call graph and control flow graph modality further analyzes the interface call functions in the interface call sequence and represents the binary application p using a graph structure. The specific method includes:
[0146] (4) According to the interface call sequence in step (3) Extract the call relationship of the interface calls in the analysis result, and generate a function call graph according to the extracted call relationship
[0147] (5) Analyze each node in the function call graph respectively and use one-hot encoding to represent external function nodes and use the function control flow graph to represent local function nodes;
[0148] (6) Correlate each node in the interface call graph generated in step (4) with the function control flow graph and one-hot encoding generated in step (5), and combine them to generate a new graphical representation.
[0149] Figure 3 An example of the grayscale image generated in step (1) is given; Figure 4 An example of the interface call sequence extracted in step (3) is given; Figure 5 An example of the interface call graph extracted in step (4) is given.
[0150] The feature extraction module includes feature extraction steps corresponding to three modalities, where:
[0151] The grayscale image features are based on the generated grayscale image I p , and the feature vectors therein are extracted in the form of a heatmap. Its workflow is as Figure 6 shown, and the specific steps are as follows:
[0152] (1) Given a grayscale image I generated from a binary application p p ;
[0153] (2) Use Grad-CAM to generate the heatmap H p corresponding to the grayscale image I I ;
[0154] (3) Superimpose the generated grayscale image I p and the heatmap H I to obtain a new superimposed heatmap
[0155] (4) Update each pixel in the new heatmap using a brightness threshold;
[0156] (5) Extract the highlighted positions relative to the pixel space as the feature representation F I of the grayscale image.
[0157] The interface call sequence features are characterized by analyzing the extracted interface call sequence using a Markov chain. Its workflow is as Figure 7 shown, and the specific steps are as follows:
[0158] (1) Given the interface call sequence
[0159] (2) Extract each interface call function C in the interface call sequence i The package name p to which it belongs i And the family name f i ;
[0160] (3) Generate a Markov chain of the conversion relationship between package names and family names according to the order in the interface call sequence
[0161] (4) Extract the top 50 nodes with the most occurrences in the Markov chain as the feature representation of the interface call sequence mode
[0162] The features of the function call graph and the control flow graph are extracted as the feature representation of the graph structure through the graph embedding method. Its working process is as follows Figure 8 As shown below, the specific steps are as follows:
[0163] (1) Given the generated function call graph And the function control flow graph And the one-hot encoding of the function names of external functions
[0164] (2) Use the GraphSAGE model to generate the feature vectors of each node in the function control flow graph
[0165] (3) Use the encoding layer to generate the global graph vector combining the function call graph and the function control flow graph As the feature representation of the mode.
[0166] Input the feature representations of the three modes into the heterogeneous selection module and the robust fusion module, and use the malware prediction module to detect the software to be detected; specifically, identify the inconsistent modes under attack through the heterogeneous selection network, use the robust fusion strategy to fuse the features of the three modes, and finally predict the malware through the neural network. The specific working process is as follows Figure 9 As shown below, the steps are as follows:
[0167] (1) Given the feature vectors of the three modes
[0168] (2) Use the feature fusion network to fuse the given three modes respectively, and at the same time use the heterogeneous selection network to identify the inconsistent modes under attack;
[0169] (3) Use the robust fusion strategy to fuse the features of the three modes to obtain the feature vector of the application p;
[0170] (4) Initialize a three-layer standard fully connected neural network, gradually reduce the output dimension, and set the output dimension of the last layer to 1;
[0171] (5) Perform a three-layer fully connected neural network operation on the connected graph vector;
[0172] (6) Use the sigmoid activation function to limit the output scalar value range within [0, 1] to obtain the probability that the software to be detected is malicious software.
[0173] Although the present invention has been disclosed above in embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention should be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to that defined by the claims.
Claims
1. A malicious software detection method for mobile platforms based on multi-modal information fusion, characterized in that, The steps include: 1) For the binary application to be detected, extract its binary sequence and generate a grayscale image, extract its interface call sequence, and extract its function call relationship and generate a function call graph and a function control flow graph. These three extraction results correspond to the grayscale image modality, the interface call sequence modality, and the graph structure modality in sequence; 2) Extract image features for the grayscale image, extract call sequence features for the interface call sequence, and extract global graph features for the function call graph and the function control flow graph; 3) A feature set is composed of the image features, the call sequence features, and the global graph features. Input this feature set into the heterogeneous selection network to output a vector of the attacked probabilities of each modality; 4) Construct a robust fusion network based on the fusion network. Use the fusion network to fuse the image features, the call sequence features, the global graph features, and the vector output by the heterogeneous selection network to output a fusion result; then use the multi-modal information fusion network to fuse the vector output by the heterogeneous selection network and the result output by the robust fusion network to output a prediction vector; 5) Reduce the dimension of the prediction vector and normalize it to obtain the predicted value of the malware.
2. The method according to claim 1, wherein The steps of extracting the binary sequence and generating the grayscale image in step 1) include: Extract the binary byte stream of the binary application, and determine the width of the grayscale image based on the size of the binary byte stream; Take 8-bit binary data as a group, convert the value of the 8-bit binary data into a grayscale value of 0-255, and convert the binary byte stream into grayscale image data in units of pixels to generate a grayscale image.
3. The method according to claim 1, wherein The steps of extracting the interface call sequence in step 1) include: Perform reverse analysis on the binary application and record the reverse analysis results; Extract the instruction code sequence therefrom according to the reverse analysis results, and retain the interface call instructions in the instruction code sequence to obtain the interface call sequence; For each interface call instruction in the interface call sequence, extract its call information, including the family name and package name of the interface to which the interface belongs in the interface call, and the call content of the interface call.
4. The method according to claim 1, wherein The steps of extracting the function call relationship and generating the function call graph and the function control flow graph in step 1) include: Extract the function call sequence of the binary application, and generate a function call graph according to the function call relationship in the function call sequence; Determine whether the function type of each node in the function call graph is an external function or a local function; if it is an external function, extract the function name in the external function node to form a set of external function names; if it is a local function, for the instruction sequence included in the local function node, extract the basic block; Generate the function control flow graph of each local function according to the jump relationship between the basic blocks within each local function.
5. The method according to claim 1, characterized in that, The steps of extracting image features for the grayscale image in step 2) include: Use the Grad-CAM network to generate a heat map according to the gradient information of the grayscale image; Overlay the generated heat map with the grayscale image to generate a new heat map; Process each pixel in the new heat map using a brightness threshold to extract the highlighted image in the new heat map; Record the position of the highlighted image pixels relative to the pixel space as the image features.
6. The method according to claim 1, characterized in that The steps of extracting call sequence features for the interface call sequence in step 2) include: Count the family names and package names that appear in all interface call instructions of the interface call sequence to generate a family name set and a package name set; According to the family name set, for each pair of adjacent interface call instructions in the interface call sequence, count the number of times the family to which the post - call interface call instruction belongs after the family to which the pre - call interface call instruction belongs, and the number of times the family to which other interface call instructions belong after the family to which the pre - call interface call instruction belongs; According to the package name set, for each pair of adjacent interface call instructions in the interface call sequence, count the number of times the package to which the post - call interface call instruction belongs after the package to which the pre - call interface call instruction belongs, and the number of times the package to which other interface call instructions belong after the package to which the pre - call interface call instruction belongs; According to the family name set, the package name set, and the above two statistical counts, construct a Markov chain representing the interface call sequence. This Markov chain consists of a state set and state transition probabilities. Each state in this state set represents a family name or a package name, and the state transition probability refers to the probability of transitioning from one state to another; Extract the top pre - set number of nodes with the most occurrences in the Markov chain to form a filtered Markov chain as the call sequence feature.
7. The method according to claim 1, wherein The steps of extracting global graph features for the function call graph and the function control flow graph in step 2) include: For all local functions included in the function call graph, use a multi - layer GraphSAGE model to generate feature vectors for each basic block node in the function control flow graph within each local function. Each layer of GraphSAGE uses a self - passing function and a message - passing function to learn and update the feature vectors of each layer from the feature vectors generated in the previous layer and the feature vectors of other neighbor nodes; after being processed by all layers of the GraphSAGE model, a total feature vector is obtained; Use an aggregation model to perform a max - pooling operation on the total feature vectors of all nodes in the function control flow graph corresponding to each local function to generate a graph vector; For all external functions included in the function call graph, use the one - hot encoding method to generate the encoding corresponding to the external function name and map it into the vector space to obtain the initial feature vector of the external function; Input the above graph vector and initial feature vector into the downstream encoding layer of the GraphSAGE model. For each node in the function call graph, use a multi - layer GraphSAGE model to update the feature vector of each node to obtain a new total feature vector; Use an aggregation model to perform a max - pooling operation on the new total feature vectors of all nodes in the function call graph to obtain the global graph feature.
8. The method according to claim 1, wherein The steps of inputting the feature set into the heterogeneous selection network and outputting a vector of modal attack probabilities in step 3) include: Input the feature set into the heterogeneous selection network. By minimizing the cross - entropy loss, an output vector with several items is obtained. The last item represents the probability that all modalities are not attacked, and the other items except the last one represent the probability that each modality is attacked.
9. The method according to claim 1, wherein In step 4), the fusion network is used to perform fusion based on the shallow neural network NN, which includes several fusion operations, and each fusion operation is used to exclude a certain attacked modality; In step 5), the prediction vector is passed through a multi-layer standard fully connected layer to reduce the dimension to one dimension; Then, the sigmoid activation function is used to limit the output scalar value range within [0,1], and finally, the predicted value of the malware after normalization is obtained.
10. A mobile platform malware detection device using the method according to any one of claims 1-9, characterized in that, It includes: A modality generation module, which is used to extract the binary sequence of the binary application to be detected and generate a grayscale image, extract its interface call sequence, and extract its function call relationship and generate a function call graph and a function control flow graph. These three extraction results correspond to the grayscale image modality, the interface call sequence modality, and the graph structure modality in sequence; A feature extraction module, which is used to extract image features for the grayscale image, extract call sequence features for the interface call sequence, and extract global graph features for the function call graph and the function control flow graph; An outlier selection module, which is used to use the outlier selection network to output a vector of the attacked probability of each modality according to the feature set composed of image features, call sequence features, and global graph features; A robust fusion module, which is used to use the fusion network of the robust fusion network to fuse the image features, call sequence features, global graph features, and the vector output by the outlier selection network, and output the fusion result; and use the multi-modal information fusion network to fuse the vector output by the outlier selection network with the result output by the robust fusion network, and output the prediction vector; A malware prediction module, which is used to reduce the dimension and normalize the prediction vector to obtain the predicted value of the malware.
Citation Information
Patent Citations
Malicious software detection method based on multi-modal deep learning
CN111382439A
Malicious software detection method and device, equipment and storage medium
CN113360912A