Android Malware Family Clustering Method Based on Multiple Features

Through the multi-feature-based Android malware family clustering method, using multiple feature information and clustering algorithms, the problem of difficult to identify unknown family malware in the existing technology is solved, and a more efficient and accurate malware family division is achieved.

CN114492586BActive Publication Date: 2025-06-03烟台北直网络科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111628481.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-06-03
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

It is difficult to effectively divide Android malware from unknown families in the existing technology, and based on supervised learning methods, it is difficult to effectively identify malware from unknown families.

Method used

A multi-feature-based clustering method of Android malware family is proposed. By extracting permission information, API sequence information and opcode sequence information from AndroidManifest.xml file and classes.dex file, feature extraction is combined with N-Gram model, and clustering is performed using Jaccard coefficient and InfoMap method.

Benefits of technology

Effectively partitioning malware into different clusters improves the accuracy of clustering, can identify malware from unknown families, and reduces the impact of a single feature on similarity calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492586B_ABST
    Figure CN114492586B_ABST
Patent Text Reader

Abstract

The present invention proposes an Android malware family clustering method based on multi-features. This method first extracts permission features, API features, and opcode features to represent information in different aspects of the samples; secondly, redundant features in the API features and opcode features are filtered in the preprocessing stage to reduce subsequent calculation time and improve the clustering effect; then the Jaccard coefficient is used to calculate the similarity of the permission features, API features, and opcode features between samples, and at the same time the sequence similarity of the APIs is calculated using the API sequence information, and all similarities are integrated to obtain the similarity between malware; finally, a malware network is constructed by setting a similarity threshold, and the InfoMap algorithm in community discovery is used to achieve clustering. Through the method of the present invention, the family clustering of malware can be effectively realized, which has very important significance for malware detection and cause analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a clustering method for Android malware families based on multiple features, and belongs to the field of software maintenance. Background Art

[0002] With the rapid development of mobile Internet technology, mobile terminal devices represented by smart phones have been widely popularized. Due to its characteristics such as open source, openness and flexibility, the Android system has become one of the most widely used operating systems in mobile terminals. However, at the same time, the Android system has also become the main target of malicious attackers, and the continuous emergence of a large number of Android malware seriously threatens people's information security. The malware named Godless discovered in 2016 can "infect" low-version Android systems to obtain root privileges. Similarly, the malware named Dvmap discovered in 2017 can inject malicious code into the running time library in low-version Android systems to obtain root privileges. When an attacker obtains the root privilege of a device, they can directly operate on the device without user authorization, including stealing the user's mobile phone data and installing other applications without permission.

[0003] Reports show that the number of newly added Android malware reaches the million level every year, and most of them are developed based on existing malware. Specifically, developers "inject" malicious components or malicious code into normal Apps and repackage them to generate new malware. Therefore, these malware have a certain family nature, that is, they contain the same malicious components or code. Currently, in the analysis related to malware, the family classification of Android malware is also an important research content. By clustering malware and dividing different malware into different families, the common malicious components and code within the same family can be identified, which can help developers better understand the properties and behaviors of malware, and thus more efficiently identify malware and judge its harm level. Therefore, Android malware family clustering has very important significance.

[0004] Regarding the problem of Android malware family clustering, many classification methods have been proposed by researchers. These methods often also select different software features and use machine learning algorithms to train classifiers to achieve the family classification of Android malware and divide malware into different families. However, most of the existing methods are based on supervised learning, highly dependent on labeled data sets, and unable to effectively divide Android malware of unknown families. To address these problems, the present invention proposes a new malware family clustering method. Summary of the Invention

[0005] To effectively achieve malware family clustering, the present invention proposes a multi-feature-based Android malware family clustering method, which can effectively divide malware into different clusters.

[0006] To achieve the above object, the technical solution of the present invention is:

[0007] A multi-feature-based Android malware family clustering method, comprising the following steps:

[0008] Step 1: Given a set A=(A 1 ,A 2 ,…A m ) of m Android application malware installation packages (apk files), for each malware installation package A i (i = 1, 2, …, m), use the Android analysis tool Androguard to obtain the AndroidManifest.xml file and the classes.dex file therein;

[0009] Step 2: Extract permission information from the AndroidManifest.xml file, and extract API sequence information and opcode sequence information from the classes.dex file; each malware can be represented as A i = (apkId, permission, APISequence, opcode), where apkId represents the number of the malware installation package, permission represents the permission set information, APISequence represents the API sequence information, and opcode represents the opcode sequence information;

[0010] Step 3: Filter permission and APISequence: According to the permission list defined in the official document, filter the third-party or custom permissions in permission; by identifying the packageName in the API, filter the non-official APIs in APISequence;

[0011] Step 4: Extract features from the opcode information: Use the N-Gram model to extract the byte segment features gram in the opcode. Each gram is composed of n opcodes. For example, "1a 71 39" represents a 3-gram opcode feature, and "1a71 39 45" represents a 4-gram opcode feature; N-Gram is a probability-based discriminant model. By sliding a window of size n over the text content by bytes, calculate the probability of each gram appearing; sort all the grams in descending order of probability, and select the top K% grams with the highest probability;

[0012] After the preprocessing in Steps 3 and 4, each malware package can be represented as A i =<apkId, prePermission, preAPISequence, preOpcode>;

[0013] Step 5: Represent the prePermission in all packages in the form of a bag of words and create a set where l i represents the number of permissions in the i-th package; for any packages A a and A b , calculate the similarity Sim per (A a , A b ) of prePermission using the Jaccard coefficient; the Jaccard coefficient mainly measures the similarity according to the ratio of the number of identical permissions contained in two packages to the number of all different permissions;

[0014] Step 6: Represent the preOpcode in all packages in the form of a bag of words and create a set where d i represents the number of opcodes in the i-th package; for any packages A a and A b , calculate the similarity Sim opc (A a , A b ) of preOpcode in the packages using the Jaccard coefficient;

[0015] Step 7: For the preAPISequence in the packages, form a sequence set of all APIs t i is the number of APIs in the i-th package; preAPISequence is actually the sequence information composed of APIs, so both its semantic similarity and sequence similarity should be considered. Among them, the semantic similarity is calculated using the Jaccard coefficient, and the sequence similarity is calculated using the corresponding relationship of the API sequence positions; combine the two similarity values to obtain the similarity Sim a and A b of preAPISequence in any packages A api (A a , A b );

[0016] Step 8: Integrate the three different types of similarities and calculate for any packages A a and Ab The final similarity Sim(A a , A b ):

[0017]

[0018] Step Nine: Software package clustering: Use the InfoMap method to cluster software packages. InfoMap is a very efficient community discovery method, and its core principle is to minimize entropy in information theory; all malware is divided into c categories through the InfoMap method.

[0019] Preferably, the fourth step includes the following steps:

[0020] Sub-step 4-1: Given a malicious Android software sample set A = (A 1 , A 2 , … A m ), extract the n-gram (n = 5) features of the opcodes of all software packages, and construct a set Ω = {o 1 , o 2 , …, o s}, where s is the number of all 5-gram opcodes;

[0021] Sub-step 4-2: Calculate the frequency of each o k (k = 1, 2, … s) in the Android malware sample set A, denoted as P(M) = {p(o 1 , A), p(o 2 , A), …, p(o s , A)}, and the calculation formula is as follows:

[0022]

[0023] where count(o k , A) represents the number of samples containing o k in the sample set A;

[0024] Sub-step 4-3: Sort all o k in descending order of probability, and select the top K% (K = 1) 5-gram opcodes as the key opcode feature list.

[0025] Preferably, the seventh step includes the following steps:

[0026] Sub-step 7-1: For any two software packages A a and A b , the semantic similarity Sim sem (A a , Ab )for:

[0027]

[0028] Sub-step 7-2, Assume One(Api a ,Api b ) indicates API a and API b The set of APIs that appear only once in all a ,Api b ) means One(Api a ,Api b ) in API a The vector of position numbers in Ps(Api a ,Api b ) indicates Pf(Api a ,Api b ) in the corresponding words of each component in API b The vector generated by the sequence sorting in Re(Api a ,Api b ) indicates Ps(A a ,A b ) is the reverse number of each adjacent component; then the sequence similarity Sim ord (APIS a ,APIS b )for:

[0029]

[0030] Sub-step 7-3, finally get A a and A b preAPISequence similarity Sim api (A a ,A b )=λ 1 ×Sim sem (A a , A b )+A 2 ×Sim ord (A a , A b ), where λ 1 =0.7,λ 2 =0.3.

[0031] Preferably, the step nine comprises the following steps:

[0032] Sub-step 9-1, initialization, each software package A i as independent communities;

[0033] Sub-step 9-2: Take the similarity Sim(A i , A j ) between nodes as the transition probability, denoted as Meanwhile, to avoid random walk entering isolated regions, a crossing probability τ (τ is a hyperparameter and 0 < τ < 1) is introduced; thus, the next transition probability of A j is where c is the number of classes;

[0034] Sub-step 9-3: Define as the optimization objective, where is the probability of the event "leaving class G k " occurring during the random walk process, and

[0035]

[0036] H(Q) represents the entropy of movement between communities, and

[0037]

[0038] where is the probability of modular movement occurring within community G k , and

[0039]

[0040] is the entropy of movement within the module, and

[0041]

[0042]

[0043] Sub-step 9-4: Randomly sample a sequence of nodes from the graph, and sequentially calculate the community L(M) where the point is assigned to a neighbor node. Find the community that causes the largest decrease in L(M), and assign the node to that community; if L(M) does not decrease for each one, the community to which the node belongs remains unchanged;

[0044] Sub-step 9-5: Repeat Sub-step 9-4 until the community to which each node belongs does not change, then stop the loop.

[0045] Compared with traditional classification methods, the beneficial effects of the present invention are as follows: The present invention constructs a multi-feature-based clustering method for Android malware families, which can effectively divide different malware into different clusters, improving the accuracy of clustering; The present invention combines various types of features to calculate the similarity between malware, effectively reducing the influence of a single feature on similarity calculation and making the similarity calculation more accurate; The present invention introduces the InfoMap method to implement malware clustering, which can effectively identify malware of unknown families. Description of the Drawings

[0046] Figure 1 It is a flowchart of the multi-feature-based clustering method for Android malware families of the present invention. Detailed Implementation Modes

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] Embodiment 1

[0049] Data source acquisition: The datasets used in this experiment include two projects, namely dataset-I and dataset-II. Among them, dataset-I contains 5,560 samples from 179 Android malware families, and dataset-II contains 8,407 samples from 36 malware families. The dataset-I involves more families, but the scale (i.e., the number of software packages) of most of the families is less than 10. The dataset-II involves a larger number of samples and fewer malware families, and the scale of each family exceeds 10.

[0050] To make the purpose, technical solutions and advantages of the present invention clearer, the following combines the attached Figure 1 drawings to provide a detailed description of the multi-feature-based clustering method for Android malware families provided by the present invention patent, including the following steps:

[0051] Step 1: Given a set A=(A 1 ,A 2 ,…A m ) of m Android application malware installation packages (apk files), for each malware installation package A i(i = 1, 2, …, m), use the Android analysis tool Androguard to obtain the AndroidManifest.xml file and the classes.dex file therein;

[0052] Step 2: Extract permission information from the AndroidManifest.xml file, and extract API sequence information and opcode sequence information from the classes.dex file; Each malware can be represented as A i = (apkId, permission, APISequence, opcode), where apkId represents the number of the malware installation package, permission represents the permission set information, APISequence represents the API sequence information, and opcode represents the opcode sequence information;

[0053] Step 3: Filter permission and APISequence: According to the permission list defined in the official document, filter the third-party or custom permissions in permission; By identifying the packageName in the API, filter the non-official APIs in APISequence;

[0054] Step 4: Extract features from the opcode information: Use the model to extract the byte fragment feature gram in the opcode. The detailed steps are as follows:

[0055] 4-1. In the given malicious Android software sample set A = (A 1 , A 2 , … A m ), extract the n-gram (n = 5) features of the opcodes of all software packages, and construct the set Ω = {o 1 , o 2 , …, o s}, where s is the number of all 5-gram opcodes;

[0056] 4-2. Calculate the frequency of each o k (k = 1, 2, … s) in the Android malware sample set A, denoted as P(M) = {p(o 1 , A), p(o 2 , A), …, p(o s , A)}, and the calculation formula is as follows:

[0057]

[0058] where count(o k , A) represents the number of samples containing o k in the sample set A;

[0059] 4-3. For all o k Sort them in descending order of probability, and select the top K% (K = 1) 5-gram opcodes as the key opcode feature list;

[0060] After the preprocessing in steps (3) and (4), each malware package is represented as A i =<apkId, prePermission, preAPISequence, preOpcode>, i = 1, 2…m.

[0061] Step Five: Represent the prePermission in all packages in the form of a bag of words and create a set where l i represents the number of permissions in the i-th package; for any packages A a and A b in prePermission, calculate the similarity Sim of prePermission using the Jaccard coefficient per (A a , A b ), and the formula is:

[0062]

[0063] where || represents the number of permissions included in the set;

[0064] Step Six: Also represent the perOpcode in all packages in the form of a bag of words and create a set where d i represents the number of opcodes in the i-th package; for any packages A a and A b in preOpcode, calculate the similarity Sim of preOpcode in the package using the Jaccard coefficient opc (A a , A b ), and the formula is:

[0065]

[0066] Step Seven: For the preAPISequence in the package, form a sequence set of all APIs where t i represents the number of APIs in the i-th package; calculate its semantic similarity using the Jaccard coefficient and its sequence similarity using the sequence position information of the APIs. The steps are as follows:

[0067] 7-1. For any two software packages A a and A b , the semantic similarity Sim sem (A a , A b ) is as follows:

[0068]

[0069] 7-2. Assume that One(Api a , Api b ) represents the set of APIs that appear and only appear once in both Api a and Api b . Pf(Api a , Api b ) represents the vector formed by the position numbers of the words in One(Api a , Api b ) in Api a . Ps(Api a , Api b ) represents the vector generated by sorting the sequences of the corresponding words in Pf(Api a , Api b ) in Api b . Re(Api a , Api b ) represents the number of inversions of adjacent components of Ps(A a , A b ); then the sequence similarity Sim ord (APIS a , APIS b ) is as follows:

[0070]

[0071] 7-3. Finally, the preAPISequenc similarity Sim a and A b in A api (A a , A b ) = λ 1 ×Sim sem (A a , A b ) + λ 2 ×Sim ord (A a , A b ), where λ 1 = 0.7, λ 2 = 0.3;

[0072] Step Eight: Integrate the three different types of similarities to calculate the final similarity between software packages; given any two software packages A a and A b , the final similarity Sim(A a ,A b ) is defined as:

[0073]

[0074] Step Nine: Software package clustering: Use the InfoMap method to cluster software packages. InfoMap is a very efficient community discovery method, and its core principle is to minimize entropy in information theory; if c is the number of expected clusters, the specific clustering process is as follows:

[0075] Sub-step 9-1: Initialization. Treat each software package A i as an independent community;

[0076] Sub-step 9-2: Take the similarity Sim(A i ,A j ) between nodes as the transition probability, denoted as At the same time, in order to avoid random walks entering isolated regions, a crossing probability τ (τ is a hyperparameter, and 0 < τ < 1) is introduced; thus, the next transition probability of A j is where c is the number of classes;

[0077] Sub-step 9-3: Define as the optimization objective, where is the probability of the event "leaving class G k " occurring during the random walk process, and H(Q) represents the entropy of movement between communities, and where is the probability of module movement occurring within community G k , and is the entropy of movement within the module, and

[0078] Sub-step 9-4: Randomly sample a sequence of nodes in the graph, and sequentially calculate the community L(M) where the node is assigned to its neighbor nodes. Find the community that causes the largest decrease in L(M), and assign the node to that community; if each L(M) does not decrease, the community to which the node belongs remains unchanged;

[0079] Sub-step 9-5: Repeat Sub-step 9-4 until the community to which each node belongs does not change, and then stop the loop.

[0080] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments still fall within the protection scope of the present invention.

Claims

1. A clustering method for Android malware families based on multiple features, characterized in that: It includes the following steps Step 1: Given a set \(A=(A 1 ,A 2 ,\cdots,A m )\) of \(m\) Android application malware installation packages, for each malware installation package \(A i \), where \(i = 1, 2,\cdots,m\), use the Android analysis tool Androguard to obtain the AndroidManifest.xml file and the classes.dex file therein; Step 2: Extract permission information from the AndroidManifest.xml file, and extract API sequence information and opcode sequence information from the classes.dex file; each malware is represented as A i =(apkId, permission, APISequence, opcode), where apkId represents the number of the malware installation package, permission represents the permission set information, APISequence represents the API sequence information, and opcode represents the opcode sequence information; Step 3: Filter permission and API Sequence: According to the permission list defined in the official document, filter the third-party or custom permissions in permission; Filter the unofficial APIs in sequence by identifying the packageName in the API; Step 4: Extract features from opcode information: Use the model to extract the byte segment feature gram in opcode; After the preprocessing in Steps 3 and 4, each malware package is represented as A i =<apkId, prePermission, preAPISequence, preOpcode>, where i = 1, 2... m; Step 5: Represent the prePermissions in all software packages in the form of a bag of words, and create a set Per = {per 1 , per 2 , …, perl i}, where l i represents the number of permissions in the i-th software package; for any software packages A a and A b 's prePermissions, use the Jaccard coefficient to calculate the similarity Sim per (A a , A b ), and the formula is: where || represents the number of permissions included in the set; Step 6: Represent the preOpcode in all software packages in the form of a bag of words, and create a set Opc = where d i represents the number of opcodes in the i-th software package; for any software packages A a and A b in the preOpcode, use the Jaccard coefficient to calculate the similarity Sim op (A a , A b ) of the preOpcode in the software package, and the formula is: Step 7: For the preAPISequence in the software package, form a sequence set of all the APIs where t i represents the number of operation codes in the i-th software package; for semantic similarity, calculate it using the Jaccard coefficient, and for sequence similarity, calculate it using the corresponding relationship of the API sequence positions; finally, comprehensively obtain the preAPISequence similarity Sim a between any software package A b and A api (A a , A b ); Step Eight: Integrate the three different types of similarity to calculate the final similarity between software packages; Given any two software packages A a and A b , the final similarity Sim(A a ,A b ) is defined as: Step 9: Software package clustering: Use the InfoMap method to cluster the software packages; Divide all malicious software into c categories through the InfoMap method.

2. A clustering method for Android malware families based on multiple features according to claim 1, characterized in that: The said Step 4 includes the following steps: Sub-step 4-1: Given a malicious Android software sample set A = (A 1 , A 2 , …, A m ), extract the n-gram (n = 5) features of the opcodes of all software packages, and construct a set Ω = {o 1 , o 2 , …, o s}, where s is the number of all 5-gram opcodes; Sub-step 4-2: Calculate the frequency of each o in Ω k appearing in the Android malware sample set A, denoted as P(M) = {p(o 1 , A), p(o 2 , A), …, p(o s , A)}, and the calculation formula is as follows: where k = 1, 2, …, s, count(o k , A) represents the number of samples containing o k in the sample set A; Sub-step 4-3: For all o k Sort them in descending order of probability, and select the top K% of the 5-gram operation codes as the key operation code feature list, where K = 1.

3. A clustering method for Android malware families based on multiple features according to claim 1, characterized in that: The said Step 7 includes the following steps: Sub-step 7-1: For any two software packages A a and A b , the semantics Sim sem (A a , A b ) is as follows: Sub-step 7-2: Assume One(Api a , Api b ) represents the set of APIs that appear and only appear once in both Api a and Api b . Pf(Api a , Api b ) represents the vector composed of the position numbers of the words in One(Api a , Api b ) in Api a . Ps(Api a , Api b ) represents the vector generated by sorting the sequences of the corresponding words in each component in Pf(Api a , Api b ) in Api b . Re(Api a , Api b ) represents the number of inversions of adjacent components of Ps(A a , A b ). Then the sequence similarity Sim ord (APIS a , APIS b ) is as follows: Sub-step 7-3, finally obtain A a and A b The similarity Sim of preAPISequenc in api (A a , A b ) = λ 1 × Sim sem (A a , A b ) + λ 2 × Sim ord (A a , A b ), where λ 1 = 0.7, λ 2 = 0.

3.

4. A clustering method for Android malware families based on multiple features according to claim 1, characterized in that: The said Step 9 includes the following steps: Sub-step 9-1, initialization: Treat each software package A i as an independent community; Sub-step 9-2: Take the similarity Sim(A i , A j ) between nodes as the transition probability, denoted as Meanwhile, to avoid random walk from entering isolated regions, a traversal probability τ is introduced, where τ is a hyperparameter and 0 < τ < 1; thus, the next transition probability of A j is where c is the number of classes; Sub-step 9-3: Define as the optimization objective, where is the probability of the event "leaving class G k " occurring during the random walk process, and H(Q) represents the entropy of movement between communities, and Among them is the probability of in-module movement occurring k inside community G, and is the motion entropy within the module, and Sub-step 9-4: Randomly sample a sequence of nodes in the graph, and sequentially calculate the community L(M) where the point is assigned to the neighbor node. Find the community that makes L(M) drop the most, and assign the node to this community; If each L(M) does not drop, the community to which the node belongs remains unchanged; Sub-step 9-5: Repeat sub-step 9-4 until there is no change in the community to which each node belongs, and stop the loop.

Citation Information

Patent Citations

  • Multi-feature detection method for mobile network terminal malware for Android

    CN108280350A

  • Android malicious software family clustering method based on method call graph

    CN111814148A