Malicious APP detection method based on API call graph centrality

Through the malicious APP detection method based on the API call graph center, the problem of difficulty in detecting unknown or variant malicious APPs in the existing technology is solved, and higher detection accuracy and system stability are achieved.

CN120012083APending Publication Date: 2025-05-16NANJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510069770.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing malicious APP detection methods are difficult to effectively detect unknown or variant malicious APPs, and their performance is degraded or battery life is shortened on resource-constrained mobile devices.

Method used

A malicious APP detection method based on the centrality of the API call graph is adopted. By analyzing the API call relationship from the APK file, an API call graph is constructed, the centrality of sensitive API nodes is calculated, and features with high correlation are selected through chi-square test, the feature dimension is reduced, and a classification model is constructed for detection.

Benefits of technology

It improves the detection accuracy of malicious behavior patterns, optimizes feature dimensions, improves the scalability of the system and the stability and accuracy of processing large-scale data sets, and enhances the distinction and generalization capabilities of malicious APP detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012083A_ABST
    Figure CN120012083A_ABST
Patent Text Reader

Abstract

The invention relates to a malicious APP detection method based on API call graph centrality. The method comprises the following steps: analyzing an API call relationship from an APK file and constructing an API call graph; carrying out centrality calculation and normalization processing on the sensitive API nodes; selecting features with high correlation with tags based on chi-square test, and reducing feature dimensions; a classification model is constructed, and malicious APPs and benign APPs are distinguished through sample features; and optimizing the detection model based on the voting mechanism. According to the method, the chi-square test method is adopted to carry out significant dimension reduction on the original feature dimension, and redundant data and noise are effectively removed. Compared with a MalScan prototype system, the feature dimension reduction processing not only improves the data processing efficiency, but also retains high-relevance features having significant contributions to classification tasks, further improves the expandability of the system and the adaptability of the system in a high-dimensional data environment, greatly reduces the calculation burden through the optimization, and reduces the calculation cost. And the stability and the accuracy of the system during large-scale data set processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer and information security technology, and in particular to a malicious APP detection method based on API call graph centrality. Background Art

[0002] With the continuous development of science and technology, smart devices are becoming more and more popular. The Android system has become the main target of malicious APP attacks due to its open source characteristics and wide application. According to security reports, the number of malicious Android APPs continues to grow, and its main manifestations include malicious APPs disguised as legitimate applications and advertising APPs. These malicious APPs are usually spread through third-party application stores, phishing websites or malicious links, inducing users to accidentally install malicious APPs, which in turn causes privacy leakage, data loss and even economic losses to users and enterprises.

[0003] Traditional Android malicious app detection methods are mainly divided into the following two categories: (i) Signature-based detection: Detection is performed by matching the signature codes of known malicious apps. Each malicious app has its own unique signature code, which can be extracted and stored in a database. When a new app is installed or run, the detection system will compare its signature code with the malicious signature code in the database to determine whether it is a malicious app. Although this method has fast detection speed and high efficiency, since signature-based detection relies on known signature codes, it cannot detect unknown or newly emerging malicious apps, as well as malicious apps that have been mutated through code obfuscation, encryption and other technologies.

[0004] (ii) Behavior-based detection: Identify potential threats by monitoring various behaviors of apps during operation, such as network communication, file reading and writing, sensitive API calls, etc. However, this method requires a lot of computing resources and is not suitable for resource-constrained mobile devices, where it may result in performance degradation or shortened battery life.

[0005] To this end, machine learning and deep learning technologies have been introduced into malicious app detection. By extracting static and dynamic features of applications, such as API call graphs, permission requests, and the frequency of sensitive API usage, classification models are constructed to achieve automated detection of malicious apps. Although such methods have strong generalization capabilities, they still face challenges in dealing with complex and diverse malicious app attacks, such as weak resistance to adversarial samples and low efficiency in processing high-dimensional data. Summary of the invention

[0006] In view of the above-mentioned shortcomings of the prior art, an object of the present invention is to provide a malicious APP detection method based on API call graph centrality to solve one or more problems in the prior art.

[0007] To achieve the above object, the technical solution of the present invention is as follows: A malicious APP detection method based on API call graph centrality includes the following steps: Parse API call relationships from APK files and build API call graphs; Calculate the centrality of sensitive API nodes and normalize them; Select features with high correlation with labels based on the chi-square test to reduce feature dimensions; Build a classification model to distinguish malicious apps from benign apps based on sample features; Optimization of detection model based on voting mechanism.

[0008] Furthermore, the centrality calculation and normalization processing of sensitive API nodes includes the following steps: Selecting centrality indicators, including Degree centrality, K-Shell centrality, Katz centrality, Closeness centrality, Harmonic centrality and PageRank centrality; Centrality calculation, obtain the centrality eigenvalues ​​of sensitive API nodes and generate the eigenvectors of each APK sample; The sample data is shuffled; The centrality eigenvalues ​​are normalized to ensure the consistency of the eigenvalues ​​in scale.

[0009] Furthermore, the centrality calculation is to calculate each of the sensitive API nodes one by one for the Degree centrality, the K-Shell centrality, the Katz centrality, the Closeness centrality, the Harmonic centrality and the PageRank centrality.

[0010] Furthermore, the step of selecting features with high correlation with labels based on the chi-square test includes the following steps: Normalize the feature input and use the different centrality feature values ​​of all APK samples in the dataset to form a feature matrix as input; Main feature extraction; The Selector object is retained, that is, the selection tool used to extract the main features is retained.

[0011] Furthermore, the main feature extraction includes the following steps: Calculate the correlation between each feature and the label in the feature matrix; A threshold is set and features within the threshold are retained to reduce the dimension of the feature matrix.

[0012] Furthermore, the construction of the classification model includes the following steps: Feature preparation, forming a data set based on each feature in the feature matrix after dimensionality reduction, and the data set is classified into a training set and a validation set; Classification model construction; Model training and hyperparameter optimization; Model testing.

[0013] Furthermore, the data set corresponds one-to-one to the APK samples, including the feature values ​​after dimensionality reduction and the corresponding labels.

[0014] Furthermore, the classification models are constructed based on different classification algorithms, including XGBoost algorithm, LightGBM algorithm and 3NN algorithm.

[0015] Furthermore, the model training is implemented based on supervised learning, the classification model performance needs to be evaluated before the hyperparameter optimization, and the evaluation method is implemented based on ten-fold cross validation.

[0016] Furthermore, the optimization of the detection model based on the voting mechanism includes the following steps: Model-level voting, predicting each centrality metric based on the classification model and making decisions based on the threshold; Feature-level voting, setting decision rules to define the sample judgment criteria and output the final prediction results; Result comparison, based on the final prediction results and the test set labels.

[0017] Compared with the prior art, the beneficial technical effects of the present invention are as follows: (i) The present invention introduces two new centrality indicators - K-Shell centrality and PageRank centrality. Compared with the four basic centrality indicators used by the MalScan prototype system, these two new indicators not only have low time complexity, but can also be combined with the original method to quantify the importance of sensitive API nodes in the call graph from a more comprehensive perspective, thereby improving the detection accuracy of malicious behavior patterns. Especially in complex call relationships, K-Shell centrality enhances the model's ability to identify structured attacks by quantifying the core degree of nodes, while PageRank centrality can further optimize malicious APP detection performance based on the global influence of nodes.

[0018] (ii) Feature dimension optimization: This paper uses the chi-square test method to significantly reduce the original feature dimension, effectively removing redundant data and noise. Compared with the MalScan prototype system, feature dimension reduction processing not only improves data processing efficiency, but also retains highly correlated features that contribute significantly to classification tasks, further improving the system's scalability and adaptability in high-dimensional data environments. This optimization greatly reduces the computational burden and improves the stability and accuracy of the system when processing large-scale data sets.

[0019] (III) Improved classifier selection: Compared with the 1NN algorithm, 3NN algorithm and random forest algorithm used in the MalScan prototype system, the present invention adopts three advanced classification models, XGBoost, LightGBM and 3NN. These models perform well in large-scale data sets and complex tasks, especially in classification accuracy and model performance. Thanks to the effective implementation of feature dimensionality reduction, the problem of long model training time has been significantly alleviated, the model training process has become more efficient, and it is more adaptable to data sets of different sizes.

[0020] (IV) Model-level voting: Based on the feature-level voting of the MalScan prototype system, the present invention innovatively introduces a model-level voting mechanism. By integrating the detection results of multiple machine learning models, a comprehensive assessment of whether a sample is a malicious app is made in the form of voting, further enhancing the discrimination ability and detection performance of malicious app detection. Compared with the single model method, model-level voting fully considers the detection advantages of different models and can effectively alleviate the detection error problem caused by feature evolution. Through this mechanism, the present invention achieves dual optimization of the accuracy and generalization ability of malicious app detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flow chart of a malicious APP detection method based on API call graph centrality according to an embodiment of the present invention is shown.

[0022] Figure 2 A system architecture diagram of APK parsing based on API call graph centrality according to an embodiment of the present invention is shown.

[0023] Figure 3 A system architecture diagram of a malicious APP detection method based on API call graph centrality according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical scheme and advantages of the present invention clearer, the following is a further detailed description of a malicious APP detection method based on the centrality of the API call graph proposed by the present invention in combination with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are in a very simplified form and use non-precise proportions, which are only used to conveniently and clearly assist in explaining the purpose of the embodiments of the present invention. In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, please refer to the accompanying drawings. It should be noted that the structure, proportion, size, etc. illustrated in the accompanying drawings of this specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention, so they have no technical substantive significance. Any structural modification, change in proportional relationship or adjustment of size, without affecting the effect that the present invention can produce and the purpose that can be achieved, should still fall within the scope of the technical content disclosed by the present invention.

[0025] See also Figures 1 to 3 The malicious APP detection method based on API call graph centrality of this embodiment includes the following steps: Step 1: Parse the API call relationship from the APK file and build an API call graph.

[0026] Parse the API call relationship from the APK file and build an API call graph, as follows Figure 2 As shown in the figure, first, the APK file is statically analyzed through the Androguard analysis toolkit to parse out the API call relationship, then a complete API call graph is constructed, and the GEXF graph file of the sample is generated. The process includes detailed analysis of method references (MethodReferences), class definitions (ClassDefinitions) and call instructions to ensure that the results cover all possible call paths.

[0027] Each node represents an API. The API can be either an API call provided by the system or a user API defined in the application. Figure 3 The arrows in step 1 in Figure 1 indicate the calling relationship between APIs. The directed edge from one API to another reveals the calling path and logical dependency.

[0028] For example, the data set provided by the MalScan prototype system contains 30,715 APK samples, including 15,285 benign apps and 15,430 malicious apps. The data set covers the period from January 2011 to December 2018.

[0029] Step 2: Calculate the centrality of sensitive API nodes and normalize them.

[0030] Step 2.1: Select the centrality index.

[0031] Based on the API call graph, we select multiple centrality indicators including Degree centrality, K-Shell centrality, Katz centrality, Closeness centrality, Harmonic centrality and PageRank centrality to evaluate the importance of sensitive API nodes. These indicators can analyze the role and status of nodes in the graph from the local and global perspectives.

[0032] Step 2.2: Centrality calculation, obtain the centrality feature value of the sensitive API node and generate the feature vector of each APK sample.

[0033] For each selected centrality index, the generated sample GEXF graph file is loaded into memory, and the Degree centrality, K-Shell centrality, Katz centrality, Closeness centrality, Harmonic centrality and PageRank centrality are calculated one by one for each sensitive API node using the known technical calculation formula to form the feature vector of each APK sample. This feature vector reflects the usage pattern and call dependency of sensitive APIs in malicious apps.

[0034] Here, the prototype system provides 21986 sensitive APIs, that is, the feature vectors are all 21986-dimensional.

[0035] Step 2.3: Shuffle the sample data.

[0036] Then, based on the existing well-known technology, the sample order is shuffled by the Shuffle Function. The sample refers to the feature vector obtained by calculating the centrality of each APK in the above data set. This can ensure the randomness of the sample data and prevent the sample arrangement order from affecting the training. Here, the random seed is set to seed=42 so that the experimental results can be reproduced.

[0037] Step 2.4: Normalize the centrality eigenvalues ​​to ensure the consistency of the eigenvalues ​​in scale.

[0038] Next, in order to unify the data range and adapt to subsequent model processing, each centrality eigenvalue is normalized and adjusted to the range of [0,1] to have a unified measurement scale, ensuring the consistency of the magnitude of different centrality features, avoiding the impact of differences in eigenvalue ranges on subsequent analysis, and facilitating subsequent chi-square test analysis. At the same time, it provides a basis for unifying the representation of different features and enhancing the comparability and robustness of the algorithm.

[0039] Step 3: Select features with high correlation with labels based on the chi-square test to reduce feature dimensions.

[0040] Step 3.1: Normalize the feature input and use the different centrality feature values ​​of all APK samples in the dataset to construct a feature matrix as input.

[0041] The normalized sensitive API node centrality feature values ​​of each APK sample obtained in the above step 2.4 are used as input features to construct a feature matrix.

[0042] Step 3.2: Main feature extraction.

[0043] Step 3.2.1: Calculate the correlation between each feature and the label in the feature matrix.

[0044] Perform a chi-square test on the normalized feature matrix. Each APK sample in the data set is labeled "malicious" or "benign" in advance. By calculating the correlation between each feature and the label, that is, under the premise of assuming that the feature and the label are independent, by constructing independent feature data and comparing the difference with the actual feature data, the correlation between the feature and the label is evaluated, and then the possible contribution to the model prediction is quantified. The main steps are as follows: Calculate observation frequency: Use the well-known calculation formula in the field of statistics to calculate the joint frequency of each feature and the target label to obtain the observation frequency matrix.

[0045] Calculate expected frequencies: Use well-known calculation formulas in the field of statistics to calculate the expected frequency matrix based on the total frequency of the features and the probability of the categories.

[0046] Calculate the chi-square statistic: Based on the observed frequency and expected frequency, the chi-square value is calculated as shown below:

[0047] In the formula, is the observation frequency, is the expected frequency, and the chi-square value measures the correlation with the label.

[0048] Step 3.2.2: Set a threshold and retain the features within the threshold to reduce the dimension of the feature matrix.

[0049] In this embodiment, the threshold percentile=30 is set, that is, the features with the top 30% relevance are selected to retain the main information in the original data, while effectively reducing the feature dimension. This process not only optimizes data storage and computing efficiency, but also speeds up the training and reasoning speed of the model, avoiding the dimensionality disaster problem caused by high-dimensional data. By simplifying the feature structure, the chi-square test also improves the scalability of the model, enabling it to better handle large-scale data sets and improve classification accuracy.

[0050] Step 3.3: The Selector object is retained, that is, the selection tool used to extract the main features is retained.

[0051] The Selector object used for the main feature extraction is retained for feature dimensionality reduction of subsequent new samples, i.e., test set samples, to ensure that the fitted chi-square test transformation can be applied to new test samples, i.e., for test set testing, the test set samples need to be transformed equally to maintain consistency. For example, if the first and third features are selected as the main features in the samples of the training set and other features are filtered out, then the first and third features should also be selected for the new samples of the test set to maintain consistency, which can ensure that the features of the test samples are consistent with the feature dimensions of the training samples, so that the classifier can make correct predictions.

[0052] Step 4: Build a classification model to distinguish malicious apps from benign apps based on sample features.

[0053] Step 4.1: Feature preparation: A new data set is formed based on each feature in the feature matrix after dimensionality reduction. The new data set is classified into a training set and a validation set.

[0054] Use the feature matrix after dimensionality reduction in step 3.2 as the input data of the classification model. The new data set corresponds to the APK sample one by one, that is, each piece of data in the new data set corresponds to the feature vector of each APK sample, including the centrality eigenvalue in the feature vector and the corresponding label, such as "malicious" or "benign", to form the training set and validation set, providing a basis for model training.

[0055] Step 4.2: Classification model construction.

[0056] According to the feature distribution and sample characteristics of the input data, machine learning algorithms that can process complex features are preferred, that is, the classification models are constructed based on different classification algorithms, including XGBoost algorithm, LightGBM algorithm and 3NN algorithm. These three algorithms are suitable for processing complex feature data due to their robustness and tolerance to noise samples.

[0057] Specifically, the XGBoost algorithm is based on the gradient boosting tree GBT, which improves model performance by integrating multiple trees. It has excellent computational efficiency and strong generalization ability, and is suitable for large-scale data sets and complex models. The LightGBM algorithm uses a gradient boosting framework, but optimizes memory usage and training speed through a histogram algorithm, which is particularly suitable for large data volumes and category feature processing. The 3NN algorithm determines the classification by selecting the three nearest neighbors, which is suitable for scenarios with simple data distribution and small sample size.

[0058] Step 4.3: Model training and hyperparameter optimization.

[0059] First, the model training is based on supervised learning. During the training process, the model learns the mapping relationship between features and sample categories, and gradually optimizes the classification boundaries to maximize the ability to distinguish samples of different categories. The classification model performance needs to be evaluated before the hyperparameter optimization. The evaluation method is based on the well-known technology of ten-fold cross validation. Through ten-fold cross validation, the model hyperparameters are continuously adjusted to evaluate the model performance, and then the model hyperparameters are adjusted to obtain the best results.

[0060] For example, in the main hyperparameter settings of XGBoost and LightGBM, the maximum depth max_depth=8, to prevent model overfitting, the estimator n_estimators=400, the learning rate learning_rate=0.05, the binary classification task target is enabled, the random seed is fixed to 42, and GPU acceleration is enabled. Refer to Tables 1 and 2 below, which report the results of the ten-fold cross validation of the present invention and the MalScan prototype system on the data set from 2011 to 2015 obtained in this embodiment, and all seeds are fixed to 42. By comparison, it is obvious that the model performance of the present invention is much better than that of the MalScan prototype system.

[0061] Table 1 Results of 10-fold cross validation of the present invention on data sets of each year (unit: %) Table 2 Results of 10-fold cross validation of the MalScan prototype system on the datasets of each year (unit: %) Step 4.4: Model testing.

[0062] The performance of the optimized model is evaluated on the test data set to verify its generalization ability. The prediction accuracy and robustness of the model on unknown data are ensured by statistically analyzing indicators in the confusion matrix of well-known technologies such as TPR and FPR.

[0063] Table 3 reports the robustness results of the two new centrality metrics obtained in this embodiment compared to the prototype system in terms of the evolution of Android APK over time. The experiment used the training set of 2011, and then tested the test sets from 2012 to 2015, with all seeds fixed at 42. As can be seen from the table, except for the fact that the two new centrality methods were slightly inferior to the Katz centrality of the prototype system on the test set of 2012, the two new methods performed very well on the test sets of other years. Table 3 Robustness of various centralities in the present invention to the evolution of Android APK over time (unit: %) Step 5: Optimization of detection model based on voting mechanism.

[0064] Step 5.1: Model-level voting, predict each centrality metric based on the classification model and make a decision based on the threshold.

[0065] For each centrality index, three trained classification models are used for prediction, and then model-level voting is performed. This voting method is based on the prediction results of each classifier, and by defining a set threshold such as N ≥ 1, which means that as long as one of the three models defines a sample as malicious, then the sample is characterized as malicious, and then a comprehensive decision is output, that is, the decision output after the judgment of the three models, to improve the overall performance of the system.

[0066] Step 5.2: Feature-level voting, setting decision rules to define the judgment criteria of samples and output the final prediction results.

[0067] Reasonable decision-making rules are set. For example, a sample is considered malicious if at least N centrality features are defined. That is, in this embodiment, six centrality indicators are equivalent to six standards. These six standards are used to judge and review the same sample. Finally, the final prediction result is output based on the independent voting results of all features.

[0068] Step 5.3: Compare the results based on the final prediction results and the test set labels.

[0069] Since feature-level voting is further implemented based on the results of model-level voting, the final prediction results of step 5.2 are compared with the test set labels. The model performance is comprehensively evaluated by analyzing the consistency between the predicted values ​​and the actual labels. Here, the actual labels are the labels previously calibrated in the test set. Combined with the key indicators in the confusion matrix, the statistical data such as accuracy, precision, recall and F1 value are calculated to intuitively demonstrate the detection capability of the voting-based classification model.

[0070] As shown in Table 4 and Table 5, the results obtained in this embodiment are reported on the robustness of the voting mechanism of the present invention and the voting mechanism of the MalScan prototype system to the evolution of Android APK over time. The experiment used the training set of 2011, and tested the test sets from 2012 to 2014. All seeds were fixed to 42, and Vote in the table refers to the set feature-level voting threshold. By comparison, it can be concluded that the performance of the present invention and the prototype system is better than the prototype system in the comparison of the test sets of the first two years. When the time is further extended and the test set of 2014 is tested, the F1 score of the present invention is close to the F1 score of the prototype system, but the TPR is 95.66%, which is much higher than the 81.99% of the prototype system, which reveals the strong robustness of the present invention in the evolution of Android APK over time.

[0071] Table 4 Robustness of the voting mechanism of the present invention to the evolution of Android APK over time (unit: %) Table 5 Robustness of the MalScan prototype system voting mechanism to the evolution of Android APK over time (unit: %) This multi-model, multi-feature comprehensive judgment mechanism can improve the robustness, accuracy and recall of the detection model, especially when dealing with diverse malicious APP variants and APP evolution problems.

[0072] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0073] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A malicious APP detection method based on API call graph centrality, characterized by: The steps include: Parse API call relationships from APK files and build API call graphs; Calculate the centrality of sensitive API nodes and normalize them; Select features with high correlation with labels based on the chi-square test to reduce feature dimensions; Build a classification model to distinguish malicious apps from benign apps through sample features; optimize the detection model based on the voting mechanism.

2. A malicious APP detection method based on API call graph centrality as claimed in claim 1, characterized in that: The steps of calculating the centrality of sensitive API nodes and normalizing them include the following: Selecting centrality indicators, including Degree centrality, K-Shell centrality, Katz centrality, Closeness centrality, Harmonic centrality and PageRank centrality; Centrality calculation, obtain the centrality eigenvalues ​​of sensitive API nodes and generate the eigenvectors of each APK sample; The sample data is shuffled; the centrality eigenvalues ​​are normalized to ensure the consistency of the eigenvalues ​​in scale.

3. A malicious APP detection method based on API call graph centrality as claimed in claim 2, characterized in that: The centrality calculation is to calculate each of the sensitive API nodes one by one for the Degree centrality, the K-Shell centrality, the Katz centrality, the Closeness centrality, the Harmonic centrality and the PageRank centrality.

4. A malicious APP detection method based on API call graph centrality as claimed in claim 3, characterized in that: The method of selecting features with high correlation with labels based on the chi-square test includes the following steps: normalizing feature input, and using different centrality feature values ​​of all APK samples in the data set to form a feature matrix as input; Extraction of main features; the Selector object is retained, that is, the selection tool used to extract the main features is retained.

5. A malicious APP detection method based on API call graph centrality as claimed in claim 4, characterized in that: The main feature extraction comprises the following steps: The correlation between each feature and the label in the feature matrix is ​​calculated; a threshold is set and features within the threshold are retained to reduce the dimension of the feature matrix.

6. A malicious APP detection method based on API call graph centrality as claimed in claim 5, characterized in that: The construction of the classification model comprises the following steps: Feature preparation, forming a new data set based on each feature in the feature matrix after dimensionality reduction, wherein the new data set is classified into a training set and a validation set; Classification model construction; Model training and hyperparameter optimization; model testing.

7. A malicious APP detection method based on API call graph centrality as claimed in claim 6, characterized in that: The new data set corresponds one-to-one to the APK samples, including the feature values ​​after dimensionality reduction and the corresponding labels.

8. A malicious APP detection method based on API call graph centrality as claimed in claim 7, characterized in that: The classification models are constructed based on different classification algorithms, including XGBoost algorithm, LightGBM algorithm and 3NN algorithm.

9. A malicious APP detection method based on API call graph centrality as claimed in claim 8, characterized in that: The model training is implemented based on supervised learning. The classification model performance needs to be evaluated before the hyperparameter optimization. The evaluation method is implemented based on ten-fold cross validation.

10. A malicious APP detection method based on API call graph centrality as claimed in claim 9, characterized in that: The optimization of the detection model based on the voting mechanism includes the following steps: Model-level voting, predicting each centrality metric based on the classification model and making decisions based on the threshold; Feature-level voting, setting decision rules to define the sample judgment criteria and output the final prediction results; Result comparison, based on the final prediction results and the test set labels.

Citation Information

Patent Citations

  • APT organization identification method and system based on stacking integration and storage medium

    CN111797394A

  • Mobile application malicious behavior pattern detection method based on API call graph extraction and recording medium and device for performing the same

    US20220164447A1

  • Malicious VBA detection using graph representation

    US20240211596A1

  • Systems and methods for detecting unknown portable executables malware

    US20240370558A1