Malicious APP family identification method based on multi-modal feature fusion
Through multimodal feature fusion and detection classification model, the shortcomings of traditional malicious APP detection methods in identifying new malicious APPs and reducing false alarm rates are solved, and higher recognition accuracy and generalization capabilities are achieved, which are suitable for real-time detection.
Patent Information
- Application Number
- CN202510170736.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Traditional malicious APP detection methods are difficult to deal with new malicious APPs, with high computing resources, high false alarm rate, insufficient recognition accuracy and generalization capabilities.
The recognition method based on multimodal feature fusion is adopted, and the low-rank multimodal fusion of bytecode image features, opcode sequence features and control flow chart features is achieved by combining the CNN-BiLSTM-Attention detection classification model to realize the identification of the malicious APP family.
It improves the accuracy and generalization ability of malicious APP recognition, reduces the false positive rate, can more effectively capture the potential behavioral characteristics and patterns of malicious APP, and quickly process a large number of malicious APP samples to achieve real-time detection.
Smart Images

Figure CN119989350A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information security technology and relates to a method for identifying a malicious APP family based on multimodal feature fusion. Background Art
[0002] In recent years, the number and types of malicious apps have exploded, and their spread speed and degree of harm have also been increasing. These malicious apps not only steal user privacy information and spread malicious code, but may also carry out network attacks, seriously threatening the privacy and property security of users. Traditional malicious app detection methods mainly rely on signature matching and behavior analysis, but they face problems such as difficulty in dealing with new malicious apps, high consumption of computing resources, high false alarm rate, and insufficient recognition accuracy and generalization ability.
[0003] In order to solve these problems, researchers began to explore the application of multimodal feature fusion technology in malicious APP detection. Multimodal feature fusion technology can integrate feature information from different modalities to more comprehensively characterize the essential characteristics of things. In the field of malicious APP detection, multimodal feature fusion technology can effectively improve recognition accuracy and generalization ability, and reduce false alarm rate. At present, some researchers have tried to combine features of multiple different modalities to build a more complex and comprehensive malicious APP identification model. These multimodal features cover multiple levels of information from static features to dynamic features, and can more deeply explore the potential behavioral characteristics and patterns of malicious APPs. However, the existing multimodal feature fusion methods still have some problems, such as the feature fusion method is not effective enough and the model complexity is high, which requires further research and improvement. Summary of the invention
[0004] In view of this, the object of the present invention is to provide a method for identifying malicious APP families based on multimodal feature fusion.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for identifying malicious APP families based on multimodal feature fusion, the method comprising the following steps:
[0007] S1: Data preprocessing, including using tools to unpack Android APK files and decompile DEX bytecode files;
[0008] S2: Bytecode image feature extraction: The bytecode file classes.dex obtained by unpacking the Android APK file is converted into an RGB image, and the EfficientNetV2L convolutional neural network is used for feature extraction to obtain the bytecode image features;
[0009] S3: Extract opcode sequence features. Decompile the bytecode file classes.dex to obtain all smali files in the file directory, extract instruction opcodes from the smali files, and map the instructions according to the instruction opcode comparison table to obtain the opcode sequence of the method code block; concatenate the opcode sequences extracted from all method code blocks to obtain the original opcode sequence of the Android application, and then use the k-Shingles algorithm to convert the original sequence into a shingle set and the SimHash algorithm to map the shingle set into a signature vector to obtain the opcode sequence features;
[0010] S4: Extract control flow graph features. After decompiling classes.dex, we get all the smali files in the file directory. We use static analysis tools to extract control flow graphs from the Smali code in the smali files. According to all the Dalvik instructions provided by the Android open source project team, we convert their nodes and edges into vector representations that can be used as neural network inputs. Then, we pass the capsule graph neural network to get the control flow graph feature vector.
[0011] S5: Multimodal feature fusion, bytecode image features, opcode sequence features, and control flow graph features are fused in low rank multimodal mode, and the low rank decomposition factor L between any two vectors is calculated for the feature vectors of the three modes, bytecode image features F1, opcode sequence features F2, and control flow graph features F3. 12 , L 13 , L 23 , get the fused feature vector F f =F1+F2+F3+L 12 +L 13 +L 23 ;
[0012] S6: Detection classification model,Using the public malicious APK dataset, by obtaining the labeled screening data as the model dataset, a CNN-BiLSTM-Attention detection classification model is constructed.
[0013] Further, the S2 is specifically:
[0014] Read the contents of the unpacked classes.dex file into a byte array byte by byte, and divide it into three-byte blocks, each corresponding to a pixel;
[0015] According to the number of data blocks, the byte array is converted into an RGB image with the same length and width, and the redundant bytes are discarded;
[0016] The obtained RGB image is resized to a uniform size and stored as a picture file in PNG format;
[0017] The pre-trained EfficientNetV2L convolutional neural network is used to extract features from RGB images, and the extracted image features are reduced to a 512-dimensional feature vector through a fully connected layer to obtain the bytecode image feature F1.
[0018] Further, the S3 is specifically:
[0019] Traverse the smali files obtained by decompiling classes.dex, extract all instruction opcodes in the method code block in each smali file, and obtain the original opcode sequence;
[0020] The k-Shingles algorithm is used to map the extracted original opcode sequence into a k-Shingles set, and then the SimHash algorithm is used to convert the k-Shingles set into a 512-dimensional signature vector.
[0021] Further, in S3, extracting the original operation code sequence is specifically:
[0022] For the method instructions in the smali file, the unnecessary operand part is ignored, the opcode part of the instruction is retained, and all the instructions therein are extracted; the instructions are mapped according to the comparison table of all Dalvik instructions provided by the Android open source project team to obtain the opcode sequence of the method code block, and finally the opcode sequences extracted from all method code blocks are spliced to obtain the opcode sequence of the Android application.
[0023] Furthermore, in S3, the original opcode sequence is converted into a 512-dimensional signature vector as follows:
[0024] According to all Dalvik instructions given by the Android open source project team, the original opcode sequence of the Android application is regarded as a document composed of 256 opcodes, and the document character set size is 256; a document is regarded as a string, and k-Shingles is used to represent the document by the document substring set; suppose that the document d consists of N words, D = (w1, w2, ..., w N ), where w n represents the nth word in document D, n∈[1,N]; the k-shingle of a document is defined as any substring of length k in the document, that is, the k-shingle is expressed as:
[0025] k-shingle=(w n , w n+1 ,…,w n+k-1 )
[0026] Where n∈[1, N-k+1], the k-shingle set is used to represent the document; k is selected as 4 to convert the opcode document into an opcode shingle substring set; a hash function is used to map all 4-shingle strings in the opcode document to a five-byte hash bucket, that is, the 4-shingle substring is mapped to [0, 2 40 -1]; then use the SimHash algorithm on the k-Shingles set, let S = {S1, S2, ..., S n} is the set of k-Shingles obtained from the original opcode sequence, where each S i is an opcode Shingle of length 4; for each S i , randomly generate a high-dimensional vector a i ∈R d , where d is a larger dimension, 1024; define the accumulator vector A∈R d And initialized to zero vector; for each S in the set S i , the corresponding vector a i Add to A:
[0027]
[0028] To obtain the SimHash signature, each dimension of A is quantized into a binary value by comparing each dimension with a threshold of 0, as follows:
[0029] SimHash(A)=(sign(A1),sign(A2),…,sign(A d ))
[0030] where sign(x i ) is the sign function. When x ≥ 0, sign(x i )=1,when x<0,sign(x i )=0;
[0031] Then use 512 random projection vectors d1, d2, ..., d 512 , each d j ∈R d , each d j Perform a dot product with the accumulator vector A and assign 0 or 1 depending on the sign of the dot product result:
[0032] h j =sign(d j A)
[0033] where h j is the jth dimension of the 512-dimensional feature vector; the final 512-dimensional feature vector F2 is given by:
[0034] F2=(h1,h2,…,h 512 ).
[0035] Further, the S4 is specifically:
[0036] Use static analysis tools to extract the control flow graph from the smali code and convert the nodes and edges of the control flow graph into vector representations. The node type is encoded as a one-hot vector.
[0037] The control flow graph is processed by the capsule graph neural network to obtain the graph embedding vector, and the final 512-dimensional control flow graph feature vector is obtained through the fully connected layer.
[0038] Furthermore, in S4, the control flow graph is extracted from the smali code and converted into a vector as follows:
[0039] Use static analysis tools to extract control flow graphs from smali code files. The control flow graph consists of nodes and edges, where each node represents a basic block, that is, a linear code without any jump target. It includes an API or method call. The edges between nodes indicate that the program jumps from one basic block to another. The jump statements are conditional jumps and unconditional jumps. The nodes of the control flow graph depict the entry and exit blocks where the control flow enters and leaves in addition to the basic blocks. According to all Dalvik instructions given by the Android open source project team, a 256-dimensional instruction vector is designed. For each node, its type is encoded as a one-hot vector, C is the set of nodes in the graph, and |C| is the number of nodes. Each node v i ∈C is represented as a vector y i .
[0040] Furthermore, in S4, the feature vector obtained through the capsule graph neural network is specifically:
[0041] The graph convolution layer uses the adjacency matrix G and the node feature vector y i To calculate the hidden state of the node; let W g is the weight matrix of the graph convolutional layer; node v i The hidden state r i It is expressed as:
[0042] r i =σ(Gy i W g )
[0043] Where σ is the Relu nonlinear activation function; the capsule layer uses a dynamic routing mechanism to aggregate the hidden states of the nodes; each capsule c j Indicates the existence of a specific feature; let p iis node v i Primary capsules, g j is a graph capsule; dynamic routing is represented as:
[0044]
[0045] where λ ij is from node v i To capsule c j The routing coefficient is learned through an iterative process:
[0046]
[0047] g′ j =p i W j
[0048] Where W j is the weight matrix from the primary capsule to the graph capsule, g' j is the mapped feature vector; to obtain the final feature vector, a fully connected layer is applied to the output of the graph capsule layer; the output vector V of the graph capsule layer c From all graph capsules g j Splicing, let W f and b f are the weights and biases of the fully connected layer; the final 512-dimensional feature vector F3 is expressed as:
[0049] F3=W f V c +b f .
[0050] Further, the S5 is specifically:
[0051] The low-rank multimodal fusion method is used for feature fusion. F1, F2, and F3 are feature vectors from three modalities respectively; F1 corresponds to the bytecode image feature vector, F2 corresponds to the opcode sequence feature vector, and F3 corresponds to the control flow graph feature vector; for two modalities F1 and F2, L 12 is the decomposition of F1×F2, expressed as:
[0052]
[0053] Among them, U1 and V2 are low-rank factors; between the feature vectors of the three modes, F1 and F2, F2 and F3, F1 and F3 obtain low-rank decomposition factors, and the fused feature vector is expressed as:
[0054] F f =F1+F2+F3+L 12 +L 13 +L 23
[0055] Where L 12 , L 13 , L 23 is the corresponding low-rank decomposition factor; after low-rank multimodal feature fusion, a 512-dimensional fusion feature is obtained.
[0056] Further, the S6 is specifically:
[0057] We collected open source data from VirusShare and other third-party data to build a malicious APK dataset, uploaded it to an online security scanning website to obtain analysis reports of various antivirus products on each malware, and used an automatic malware tagging tool to classify malware samples into families based on the malware category and behavior to obtain labels. We also removed malware that could not be found by various antivirus software and families that contained less than 20 malicious applications, and screened out labeled malware as a model dataset.
[0058] The input layer of the model corresponds to the sample feature vector 512 dimensions; a BiLSTM layer is added after the convolution layer and the pooling layer, and the Attention layer is added after the BiLSTM layer. The output feature sequence of the BiLSTM layer is assumed to be the matrix B = (b1, b2, ..., b N )∈R N×D , B is used as the input of the Attention layer, where the query vector Q a , the key vector K a and the input vector V a are all equal, which is the self-attention mechanism; its calculation process is expressed as:
[0059]
[0060] O=SV a
[0061] Where Q a =K a =V a =B, M is the attention score matrix, O is the output of the attention layer; the extracted bytecode image feature dimension is 512, the opcode sequence feature and the control flow graph feature dimension are both 512, the multimodal feature fusion part sets the output dimension to 512, and the model input layer neurons are set to 512; the software sample family category of the dataset is 10, then the number of neurons in the model output layer is 10.
[0062] The beneficial effects of the present invention are:
[0063] (1) Through multimodal feature fusion, we can understand the behavioral characteristics of the malicious APP family more comprehensively and deeply, thereby improving the accuracy of identification. Compared with using only a single modal feature, fusing multiple features can more effectively capture the potential behavioral characteristics and patterns of malicious APPs and reduce the false positive and false negative rates.
[0064] (2) Multimodal feature fusion can enhance the generalization ability of the model, enabling it to better cope with new malicious app families. As malicious app families continue to evolve, single modal features may be difficult to cope with new threats, while multimodal feature fusion can provide richer information and make the model more robust.
[0065] (3) By performing low-rank decomposition on the features, the impact of noise can be effectively reduced, thereby reducing the false alarm rate. False alarms can cause users trouble and reduce their trust in malicious app detection tools. Therefore, reducing the false alarm rate is crucial to improving user experience and the practicality of malicious app detection tools.
[0066] (4) The feature extraction and fusion method used in the present invention has high efficiency, can quickly process a large number of malicious APP samples, and realize real-time detection. This is of great significance for timely identification and prevention of the spread of malicious APPs.
[0067] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0069] Figure 1 It is a flow chart of the present invention;
[0070] Figure 2 It is a schematic diagram of the system architecture of the present invention;
[0071] Figure 3 This is a schematic diagram of the feature extraction fusion architecture of the present invention;
[0072] Figure 4 Schematic diagram of the bytecode image feature extraction process in the present invention. DETAILED DESCRIPTION
[0073] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0074] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0075] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0076] Please refer to the figure Figure 1 to Figure 4 The purpose of the present invention is to provide a method for identifying malicious APP families based on multimodal feature fusion, which realizes the identification and judgment of malicious APP families by fusing the bytecode image features, operation code sequence features and control flow chart features of malicious APPs and constructing a detection and classification model.
[0077] The technical solution of the present invention comprises the following steps:
[0078] 1. Data preprocessing: Unpack and decompile the Android APK file to obtain the bytecode file and smali code file.
[0079] 2. Feature extraction: extract bytecode image features, opcode sequence features and control flow graph features respectively.
[0080] 3. Multimodal feature fusion: The three extracted features are subjected to low-rank multimodal fusion to obtain fused features.
[0081] 4. Detection and classification: Build a CNN-BiLSTM-Attention model based on fusion features to classify malicious app families.
[0082] The specific implementation of the present invention is as follows:
[0083] S1: Data preprocessing module. Apktool and baksmali tools are used in the data preprocessing module. Apktool is used to unpack the APK file to obtain the bytecode file. Baksmali is used to decompile the DEX bytecode file to obtain the smali code file.
[0084] S2: Bytecode image feature extraction module. The bytecode file classes.dex obtained by unpacking the Android APK file is converted into an RGB image, and the EfficientNetV2L convolutional neural network is used to extract features to obtain bytecode image features.
[0085] S3: Operation code sequence feature extraction module. Extract instruction operation codes from the smali files in all file directories obtained after decompiling classes.dex to form the original operation code sequence, and use the k-Shingles algorithm and SimHash algorithm to map the operation code sequence into a signature vector.
[0086] S4: Control flow graph feature extraction module. Decompile from classes.dex to get all the smali files in the file directory. After extracting the control flow graph from the Smali code in the smali file, convert its nodes and edges into vector representations that can be used as neural network input. Then pass through the capsule graph neural network to get the graph embedding vector.
[0087] S5: Feature fusion module. The bytecode image features, opcode sequence features, and control flow graph features are multimodally fused using a low-rank multimodal fusion method.
[0088] S6: Detection and classification module. A malicious APK dataset is built based on VirusShare2020-2022 and other third-party data. The 7705 malicious Android APK files obtained by screening are used to build a CNN-BiLSTM-Attentin model.
[0089] S1: The data preprocessing module specifically includes the following steps:
[0090] S1-1: Unpacking and decompiling: First, use the Apktool tool to unpack the Android installation package file to obtain the global configuration file AndroidManifest.xml, the DEX bytecode file classes.dex and other resource files, read all the contents in the DEX file byte by byte and save them for subsequent bytecode image construction, and then use the baksmali tool to decompile the DEX bytecode file to obtain the smali code file.
[0091] S2: The bytecode image feature extraction module specifically includes the following steps:
[0092] S2-1: Bytecode image construction: First, read all the contents of the DEX file classes.dex into a byte array, and divide the byte array into three-byte data blocks so that each data block corresponds to one pixel. Then, according to the number of data blocks, convert the byte array into an RGB image with the same length and width, and discard the few extra bytes. Finally, store the resulting image as a PNG picture.
[0093] S2-2: Extract bytecode image features: Use the EfficientNetV2L pre-trained model on the ImageNet dataset to initialize the feature extraction model, and modify the model output layer for the bytecode image feature extraction task. The optimal size of the EfficientNetV2L input image is 480*480, and the number of channels is 1792. Therefore, the present invention uniformly scales the bytecode image to a size of 480*480 as the model input. In the model output layer, the 1792-channel feature map output by EfficientNetV2L is subjected to global average pooling to obtain the vector F oe ∈R 1792 , in order to reduce the feature dimension and reduce storage, the vector F oe Then, through a fully connected layer with 512 neurons, a 512-dimensional feature vector F1 is obtained, and this feature vector is used as the bytecode image feature.
[0094] S3: The operation code sequence feature extraction module specifically includes the following steps:
[0095] S3-1: Original opcode sequence: First, decompile the DEX bytecode file classes.dex to get the smali directory, then traverse each directory and its subdirectories, and divide the content of each smali file into method code blocks. For the method code block, ignore the operand part, extract all the instructions in it, and map the instructions according to the instruction opcode comparison table to get the opcode sequence of the method code block. Finally, splice the opcode sequences extracted from all method code blocks to get the opcode sequence of the Android application.
[0096] S3-2: Convert to shingle set: According to all Dalvik instructions provided by the Android open source project team, the original opcode sequence of Android applications can be regarded as a document consisting of 256 opcodes, and the document character set size is 256. A document can be regarded as a string, and k-Shingles are used to represent the document by a set of document substrings. Suppose document D consists of N words, D = (w1, w2, ..., w N ), where w n represents the nth word in document D, n∈[1,N]. The k-shingle of a document is defined as any substring of length k in the document, that is, the k-shingle can be expressed as:
[0097] k-shingle=(w n , w n+1 ,…,w n+k-1 )
[0098] Where n∈[1, N-k+1], the k-shingle set is used to represent the document. The present invention selects k as 4 to convert the opcode document into an opcode shingle substring set. A hash function is used to map all 4-shingle strings in the opcode document to a five-byte hash bucket, that is, the 4-shingle substring is mapped to [0, 2 40 -1].
[0099] S3-3: Calculate the signature vector: The present invention adopts the SimHash algorithm, assuming S = {S1, S2, ..., S n} is the set of k-Shingles obtained from the original opcode sequence, where each S i is an opcode Shingle of length 4. For each S i , randomly generate a high-dimensional vector a i ∈R d , where d is a larger dimension, 1024. Define the accumulator vector A∈R d And initialized to zero vector. For each S in the set S i , the corresponding vector a i Add to A:
[0100]
[0101] To obtain the SimHash signature, each dimension of A is quantized into a binary value by comparing each dimension with a threshold of 0. The formula is:
[0102] SimHash(A)=(sign(A1),sign(A2),…,sign(A d ))
[0103] where sign(x i ) is the sign function. When x ≥ 0, sign(x i )=1,when x<0,sign(x i )=0.
[0104] Then use 512 random projection vectors d1, d2, ..., d 512 , each d j ∈R d , each d j Perform a dot product with the accumulator vector A and assign 0 or 1 depending on the sign of the dot product result:
[0105] h j =sign(d j A)
[0106] where h j is the jth dimension of the 512-dimensional feature vector. The final 512-dimensional feature vector F2 is given by:
[0107] F2=(h1,h2,…,h 512 )
[0108] S4: The control flow graph feature extraction module specifically includes the following steps:
[0109] S4-1: Graph extraction: First, decompile the DEX bytecode file classes.dex to get the smali directory, then traverse each directory and its subdirectories to get all the method code blocks under the smali files, that is, the smali code. Use the AmanDroid tool to extract the control flow graph. The control flow graph consists of nodes and edges, where each node represents a basic block, that is, a linear code without any jump target, such as an API or method call. The edges between nodes indicate that the program jumps from one basic block to another. The jump statement can be a conditional jump or an unconditional jump, etc. The nodes of the control flow graph depict the entry block and exit block where the control flow enters and leaves in addition to the basic block.
[0110] S4-2: Graph vectorization: After converting the source code into a graph, the next step is to convert its nodes and edges into vector representations that can be used as neural network input. According to all Dalvik instructions given by the Android open source project team, a 256-dimensional instruction vector is designed. Therefore, for each node, its type is encoded as a one-hot vector, C is the set of nodes in the graph, and |C| is the number of nodes. Each node v i∈C is represented as a vector y i .
[0111] S4-3: Graph Embedding Vector: The graph convolution layer uses the adjacency matrix G and the node feature vector y i To calculate the hidden state of the node. Let W g is the weight matrix of the graph convolutional layer. i The hidden state r i It can be expressed as:
[0112] r i =σ(Gy i W g )
[0113] Where σ is the Relu nonlinear activation function. The capsule layer uses a dynamic routing mechanism to aggregate the hidden states of nodes. Each capsule c j Indicates the existence of a specific feature or concept. Let p i is node v i Primary capsules, g j is a graph capsule. Dynamic routing can be expressed as:
[0114]
[0115] where λ ij is from node v i To capsule c j The routing coefficient is learned through an iterative process:
[0116]
[0117] g′ j =p i W j
[0118] Where W j is the weight matrix from the primary capsule to the graph capsule, g' j is the mapped feature vector. To obtain the final feature vector, a fully connected layer needs to be applied to the output of the graph capsule layer. The output vector V of the graph capsule layer c From all graph capsules g j Splicing, let W f and b f are the weights and biases of the fully connected layer. The final 512-dimensional feature vector F3 can be expressed as:
[0119] F3=W f V c +b f
[0120] S5: The feature fusion recognition module specifically includes the following steps:
[0121] S5-1: Multimodal feature fusion: Use low-rank multimodal fusion method to perform feature fusion. It is known that F1, F2, and F3 are feature vectors from three modalities. F1 corresponds to the bytecode image feature vector, F2 corresponds to the opcode sequence feature vector, and F3 corresponds to the control flow graph feature vector. For two modalities F1 and F2. L 12 It is the decomposition of F1×F2 and can be expressed as:
[0122]
[0123] Among them, U1 and V2 are low-rank factors. For the feature vectors of the three modes, F1 and F2, F2 and F3, and F1 and F3 obtain low-rank decomposition factors, and the fused feature vector can be expressed as:
[0124] Ff=F1+F2+F3+L 12 +L 13 +L 23
[0125] Where L 12 , L 13 , L 23 is the corresponding low-rank decomposition factor. After low-rank multimodal feature fusion, a 512-dimensional fusion feature is obtained.
[0126] S6: The detection and classification module specifically includes the following steps:
[0127] S6-1: Obtain training data set: Collect VirusShare2020-2022 and other third-party data to build a malicious APK data set, upload it to the official website of VirusTotal to obtain analysis reports of various antivirus products on each malware, and use the automatic malware labeling tool AVclass2 to classify malware samples into families according to the malware category, behavior, etc. to obtain labels, and remove malware that cannot be found by various antivirus software and remove families that contain less than 20 malicious applications. A total of 7705 labeled malware from 10 families were screened. Specifically, there are 1518 Multi.Generic families, 992 SMS.Fakeapp families, 936 DroidKungFu families, 781 Agent.aas families, 772 Fakemoney families, 646 GriftHorse families, 558 Dropper families, 537 GenericML families, 490 Airpush families, and 475 RuMMS families.
[0128] S6-2: Detection model: Construct a CNN-BiLSTM-Attention model, where the input layer corresponds to a fused feature vector of 512 dimensions. Add a BiLSTM layer after the convolution layer and the pooling layer, and add the Attention layer after the BiLSTM layer. Suppose the output feature sequence of the BiLSTM layer is a matrix B = (b1, b2, ..., b N )∈R N×D , B is used as the input of the Attention layer, where the query vector Q a , the key vector K a and the input vector V a are equal, which is the self-attention mechanism. The calculation process is as follows:
[0129]
[0130] O=MV a
[0131] Where Q a =K a =V a =B, M is the attention score matrix, and O is the attention layer output. The bytecode image feature dimension extracted by the present invention is 512, the opcode sequence feature and the control flow chart feature dimension are both 512, and the multimodal feature fusion part sets the output dimension to 512, so the model input layer neurons should be set to 512. The family category of the data set software sample is 10, so the number of model output layer neurons is 10.
[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A method for identifying malicious APP families based on multimodal feature fusion, characterized by: The method comprises the following steps: S1: Data preprocessing, including using tools to unpack Android APK files and decompile DEX bytecode files; S2: Bytecode image feature extraction: The bytecode file classes.dex obtained by unpacking the Android APK file is converted into an RGB image, and the EfficientNetV2L convolutional neural network is used for feature extraction to obtain the bytecode image features; S3: Extract opcode sequence features. Decompile the bytecode file classes.dex to obtain all smali files in the file directory, extract instruction opcodes from the smali files, and map the instructions according to the instruction opcode comparison table to obtain the opcode sequence of the method code block; concatenate the opcode sequences extracted from all method code blocks to obtain the original opcode sequence of the Android application, and then use the k-Shingles algorithm to convert the original sequence into a shingle set and the SimHash algorithm to map the shingle set into a signature vector to obtain the opcode sequence features; S4: Extract control flow graph features. After decompiling classes.dex, we get all the smali files in the file directory. We use static analysis tools to extract control flow graphs from the Smali code in the smali files. According to all the Dalvik instructions provided by the Android open source project team, we convert their nodes and edges into vector representations that can be used as neural network inputs. Then, we pass the capsule graph neural network to get the control flow graph feature vector. S5: Multimodal feature fusion, bytecode image features, opcode sequence features, and control flow graph features are fused in low rank multimodal mode, and the low rank decomposition factor L between any two vectors is calculated for the feature vectors of the three modes, bytecode image features F1, opcode sequence features F2, and control flow graph features F3. 12 , L 13 , L 23 , get the fused feature vector f f =F1+F2+F3+L 12 +L 13 +L 23 ; S6: Detection classification model,Using the public malicious APK dataset, by obtaining the labeled screening data as the model dataset, a CNN-BiLSTM-Attention detection classification model is constructed.
2. The method for identifying malicious APP families based on multimodal feature fusion according to claim 1 is characterized in that: The S2 is specifically: Read the contents of the unpacked classes.dex file into a byte array byte by byte, and divide it into three-byte blocks, each corresponding to a pixel; According to the number of data blocks, the byte array is converted into an RGB image with the same length and width, and the redundant bytes are discarded; The obtained RGB image is resized to a uniform size and stored as a picture file in PNG format; The pre-trained EfficientNetV2L convolutional neural network is used to extract features from RGB images, and the extracted image features are reduced to a 512-dimensional feature vector through a fully connected layer to obtain the bytecode image feature F1.
3. The method for identifying malicious APP families based on multimodal feature fusion according to claim 2 is characterized in that: The S3 is specifically: Traverse the smali files obtained by decompiling classes.dex, extract all instruction opcodes in the method code block in each smali file, and obtain the original opcode sequence; The k-Shingles algorithm is used to map the extracted original opcode sequence into a k-Shingles set, and then the SimHash algorithm is used to convert the k-Shingles set into a 512-dimensional signature vector.
4. The method for identifying malicious APP families based on multimodal feature fusion according to claim 3 is characterized in that: In S3, the original operation code sequence is extracted as follows: For the method instructions in the smali file, the unnecessary operand part is ignored, the opcode part of the instruction is retained, and all the instructions therein are extracted; the instructions are mapped according to the comparison table of all Dalvik instructions provided by the Android open source project team to obtain the opcode sequence of the method code block, and finally the opcode sequences extracted from all method code blocks are spliced to obtain the opcode sequence of the Android application.
5. The method for identifying malicious APP families based on multimodal feature fusion according to claim 3 is characterized in that: In S3, the original opcode sequence is converted into a 512-dimensional signature vector as follows: According to all Dalvik instructions given by the Android open source project team, the original opcode sequence of the Android application is regarded as a document composed of 256 opcodes, and the document character set size is 256; a document is regarded as a string, and k-Shingles is used to represent the document by the document substring set; suppose that the document D consists of N words, D = (w1, w2, ..., w N ), where w n represents the nth word in document D, n∈[1,N]; the k-shingle of a document is defined as any substring of length k in the document, that is, the k-shingle is expressed as: k-shingle=(w n ,w n+1 ,…,w n+k-1 ) Where n∈[1, N-k+1], the k-shingle set is used to represent the document; k is selected as 4 to convert the opcode document into an opcode shingle substring set; a hash function is used to map all 4-shingle strings in the opcode document to a five-byte hash bucket, that is, the 4-shingle substring is mapped to [0, 2 40 -1]; then use the SimHash algorithm on the k-Shingles set, let S = {S1, S2, ..., S n } is the set of k-Shingles obtained from the original opcode sequence, where each S i is an opcode Shingle of length 4; for each S i , randomly generate a high-dimensional vector a i ∈R d , where d is a larger dimension, 1024; define the accumulator vector A∈R d And initialized to zero vector; for each S in the set S i , the corresponding vector a i Add to A: To obtain the SimHash signature, each dimension of A is quantized into a binary value by comparing each dimension with a threshold of 0, as follows: SimHash(A)=(sign(A1),sign(A2),…,sign(A d )) where sign(x i ) is the sign function. When x ≥ 0, sign(x i )=1,when x<0,sign(x i )=0; Then use 512 random projection vectors d1, d2, ..., d 512 , each d j ∈R d , each d j Perform a dot product with the accumulator vector A and assign 0 or 1 depending on the sign of the dot product result: h j =sign(d j ·A) where h j is the jth dimension of the 512-dimensional feature vector; the final 512-dimensional feature vector F2 is given by: <h2 style=";text-align:left;direction:ltr">F2=(h1,h2,…,h<h2 style=";text-align:left;direction:ltr"> 512 <h2 style=";text-align:left;direction:ltr"> )。 6. The method for identifying malicious APP families based on multimodal feature fusion according to claim 1 is characterized in that: The S4 is specifically: Use static analysis tools to extract the control flow graph from the smali code and convert the nodes and edges of the control flow graph into vector representations. The node type is encoded as a one-hot vector. The control flow graph is processed by the capsule graph neural network to obtain the graph embedding vector, and the final 512-dimensional control flow graph feature vector is obtained through the fully connected layer.
7. The method for identifying malicious APP families based on multimodal feature fusion according to claim 6 is characterized in that: In S4, the control flow graph is extracted from the smali code and converted into a vector as follows: Use static analysis tools on smali code files to extract control flow graphs; control flow graphs consist of nodes and edges, where each node represents a basic block, that is, a linear code without any jump targets; Includes an API or method call; the edge between nodes indicates that the program jumps from one basic block to another, and the jump statements are conditional jumps and unconditional jumps; the nodes of the control flow graph depict the entry block and exit block of the control flow in addition to the basic block; according to all Dalvik instructions given by the Android open source project team, a 256-dimensional instruction vector is designed; for each node, its type is encoded as a one-hot vector, C is the set of nodes in the graph, |C| is the number of nodes; each node v i ∈C is represented as a vector y i .
8. The method for identifying malicious APP families based on multimodal feature fusion according to claim 6 is characterized in that: In S4, the feature vector obtained by the capsule graph neural network is specifically: The graph convolution layer uses the adjacency matrix C and the node feature vector y i To calculate the hidden state of the node; let W g is the weight matrix of the graph convolutional layer; node v i The hidden state r i It is expressed as: r i =σ(Gy i W g ) Where σ is the Relu nonlinear activation function; the capsule layer uses a dynamic routing mechanism to aggregate the hidden states of the nodes; each capsule c j Indicates the existence of a specific feature; let p i is node v i Primary capsules, g j is a graph capsule; dynamic routing is represented as: where λ ij is from node v i To capsule c j The routing coefficient is learned through an iterative process: g′ j =p i W j Where W j is the weight matrix from the primary capsule to the graph capsule, g' j is the mapped feature vector; to obtain the final feature vector, a fully connected layer is applied to the output of the graph capsule layer; the output vector V of the graph capsule layer c From all graph capsules g j Splicing, let W f and b f are the weights and biases of the fully connected layer; the final 512-dimensional feature vector F3 is expressed as: F3=W f V c +b f 。 9. The method for identifying malicious APP families based on multimodal feature fusion according to claim 1 is characterized in that: The S5 is specifically: The low-rank multimodal fusion method is used for feature fusion. F1, F2, and F3 are feature vectors from three modalities respectively; F1 corresponds to the bytecode image feature vector, F2 corresponds to the opcode sequence feature vector, and F3 corresponds to the control flow graph feature vector; for two modalities F1 and F2, L 12 is the decomposition of F1×F2, expressed as: Among them, U1 and V2 are low-rank factors; between the feature vectors of the three modes, F1 and F2, F2 and F3, F1 and F3 obtain low-rank decomposition factors, and the fused feature vector is expressed as: F f =F1+F2+F3+L 12 +L 13 +L 23 Where L 12 , L 13 , L 23 is the corresponding low-rank decomposition factor; after low-rank multimodal feature fusion, a 512-dimensional fusion feature is obtained.
10. The method for identifying malicious APP families based on multimodal feature fusion according to claim 1 is characterized in that: The S6 is specifically: We collected open source data from VirusShare and other third-party data to build a malicious APK dataset, uploaded it to an online security scanning website to obtain analysis reports of various antivirus products on each malware, and used an automatic malware tagging tool to classify malware samples into families based on the malware category and behavior to obtain labels. We also removed malware that could not be found by various antivirus software and families that contained less than 20 malicious applications, and screened out labeled malware as a model dataset. The input layer of the model corresponds to the sample feature vector 512 dimensions; a BiLSTM layer is added after the convolution layer and the pooling layer, and the Attention layer is added after the BiLSTM layer. The output feature sequence of the BiLSTM layer is assumed to be the matrix B = (b1, b2, ..., b N )∈R N×D , B is used as the input of the Attention layer, where the query vector Q a , the key vector K a and the input vector V a are all equal, which is the self-attention mechanism; its calculation process is expressed as: O=SV a Where Q a =K a =V a =B, M is the attention score matrix, O is the output of the attention layer; the extracted bytecode image feature dimension is 512, the opcode sequence feature and the control flow graph feature dimension are both 512, the multimodal feature fusion part sets the output dimension to 512, and the model input layer neurons are set to 512; the software sample family category of the dataset is 10, then the number of neurons in the model output layer is 10.
Citation Information
Patent Citations
Android malicious application detection method and system based on multi-feature fusion
CN107180192A
Android malicious application program identification method based on multi-modal neural network
CN114491529A
Android malicious software automatic identification method based on feature processing
CN116956281A
Image-text multi-mode Android App fraud-related study and judgment classification method based on ViLT
CN117909972A
Multi-modal feature fusion Android malicious software detection method based on attention mechanism
CN118194288A
Cited By
Malware family classification method
CN120579008A