Malware Detector Training Method, Detector, Electronic Device, and Storage Medium
By selecting representative sample data sets and training malware detectors, the problem of high difficulty in model training in rules-based Android malware detection methods is solved, and high-accuracy malware detection is achieved, reducing labor costs.
Patent Information
- Application Number
- CN202210201495.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-03
AI Technical Summary
Rules-based Android malware detection methods rely on a large number of manual analysis, which makes model training difficult and cannot effectively reduce the difficulty of model training, while ensuring the accuracy of the model after training.
By obtaining the characteristic parameters of the original sample dataset, selecting the representative sample dataset with a proportion of α in the total sample, and inputting it to the preset training model for training, obtaining a malware detector. This method reduces the difficulty of model training while ensuring the accuracy of malware detection rates.
It realizes that while reducing the difficulty of model training, it ensures the accuracy of the model after training, reduces the need for manual analysis and reduces labor costs.
Smart Images

Figure CN114598443B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software security, and in particular to a method for training a malware detector, a detector, an electronic device, and a storage medium. Background Art
[0002] Currently, a huge number of malware poses a great threat to the security of the Android system and the rights and interests of users. Therefore, researching methods for detecting Android malware is one of the important contents in the field of security protection for mobile operating systems.
[0003] Interpretable Android malware detection methods are mainly rule-based Android malware detection methods. This method mainly extracts the permissions that malware frequently requests but benign software rarely requests, and uses these as rules for detecting Android malware, and then uses this rule set to detect malware.
[0004] However, the inventors found that: The rule-based Android malware detection method can reflect the causal relationship between features and detection results, but this method is based on a large amount of manual analysis, and the training difficulty of the model is high. Summary of the Invention
[0005] The present invention provides a method for training a malware detector, a detector, an electronic device, and a storage medium, which can reduce the training difficulty of the model while ensuring the accuracy of the trained model.
[0006] According to one aspect of the present invention, there is provided a method for training a malware detector, including: obtaining an original sample data set, and obtaining the original malware detection rate of the original sample data set, where the original sample data set includes a plurality of original samples; obtaining the feature parameters of each original sample, where the feature parameters are used to characterize the uncertainty degree of the original sample being malware; according to the feature parameters, selecting a representative sample data set with a proportion of α in the total samples from the original sample data set, and obtaining the malware detection rate of the representative sample data set, where α is greater than 0 and less than 1, and the difference between the malware detection rate and the original malware detection rate is within a first preset range; inputting the representative sample data set into a preset training model for training to obtain a malware detector.
[0007] According to another aspect of the present invention, there is provided a malware detector, comprising: an original sample detection rate acquisition module, configured to acquire an original sample data set and obtain the original malware detection rate of the original sample data set, wherein the original sample data set includes a plurality of original samples; a feature parameter acquisition module, configured to acquire the feature parameters of each of the original samples, the feature parameters being used to characterize the uncertainty degree of the original sample being a malware; a representative sample detection rate acquisition module, configured to select a representative sample data set with a proportion of α in the total samples from the original sample data set according to the feature parameters, and obtain the malware detection rate of the representative samples, wherein α is greater than 0 and less than 1, and the difference between the malware detection rate and the original software detection rate is within a first preset range; a detector training module, configured to input the representative sample data set into a preset training model for training to obtain a malware detector.
[0008] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:
[0009] at least one processor; and
[0010] a memory communicatively connected to the at least one processor; wherein,
[0011] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the malware detector training method according to any embodiment of the present invention.
[0012] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the malware detector training method according to any embodiment of the present invention when executed.
[0013] Compared with the related art, the embodiments of the present invention have at least the following advantages:
[0014] By selecting a representative sample data set from the original sample data set according to the feature parameters of the original samples, on the one hand, the number of sample data input into the preset model for training can be reduced, so that the model does not need to train a large amount of data, thereby reducing the difficulty of model training; on the other hand, it can ensure that the difference between the malware detection rate and the original malware detection rate is within a first preset range, so that training with the representative sample data set can achieve the same training effect as training with the original sample data set, thereby ensuring the accuracy of the preset model after training; in addition, the above model training method does not require manual analysis, which can greatly reduce the labor cost.
[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 is a flowchart of a method for training a malware detector according to Embodiment 1 of the present invention;
[0018] Figure 2 is a flowchart of a method for training a malware detector according to Embodiment 2 of the present invention;
[0019] Figure 3 is a flowchart of a method for training a malware detector according to Embodiment 3 of the present invention;
[0020] Figure 4 is a flowchart of a method for training a malware detector according to Embodiment 4 of the present invention;
[0021] Figure 5 is a schematic block diagram of the principle of a method for training a malware detector according to Embodiment 4 of the present invention;
[0022] Figure 6 Schematic structural diagram of a malware detector according to Embodiment 5 of the present invention;
[0023] Figure 7 is a schematic structural diagram of an electronic device for implementing the method for training a malware detector according to Embodiment 6 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] Embodiment 1
[0027] Figure 1 The following is a flowchart of a method for training a malware detector provided in Embodiment 1 of the present invention. As Figure 1 shown, the method includes:
[0028] S110. Obtain the original sample dataset and obtain the original malware detection rate of the original sample dataset.
[0029] Specifically, the original sample dataset includes multiple original samples. When obtaining the original sample dataset, the types of the original samples (malware or benign software) can be known at the same time. That is to say, the specific number of malware in the original sample dataset can be known at this time; input the original samples into the classifier trained by the feature set, and the probability that each sample is classified as malware or benign software can be obtained (for example, if the probability that a sample is malware is 0.6 and the probability of benign software is 0.4, then the sample is determined to be malware); assume that the specific number of malware in the original sample dataset is 500, and the number of malware detected by the classifier trained by the feature set is 450, then the original malware detection rate of the original sample dataset is 450 / 500 = 90%.
[0030] S120. Obtain the feature parameters of each original sample.
[0031] Specifically, the feature parameters are used to characterize the uncertainty degree of the original sample being malware. In this embodiment, the feature parameter is the information entropy, and the information entropy of the original sample can be obtained in the following way:
[0032] Input multiple original samples into the preset trained classifier to obtain the probability that each original sample is classified as malware or benign software; obtain the information entropy according to the following formula:
[0033]
[0034] Where n is the number of original samples, i is the serial number of the original sample, p(y i ) is the probability that the original sample is classified as malware or benign software, and H(Y) is the information entropy.
[0035] It should be noted that the preset training classifier can be the feature set training classifier mentioned above. This embodiment does not specifically limit the type of the preset training classifier, and only needs to be able to distinguish whether the original sample is malware or benign software.
[0036] S130. According to the feature parameters, select a representative sample data set with a proportion of α in the total samples from the original sample data set, and obtain the malware detection rate of the representative sample data set.
[0037] Specifically, α is greater than 0 and less than 1, and the difference between the malware detection rate and the original malware detection rate is within the first preset range.
[0038] It is worth mentioning that, in order to ensure that the difference between the malware detection rate and the original malware detection rate is within the first preset range, before selecting a representative sample data set with a proportion of α in the total samples from the original sample data set, the multiple original samples in the original sample data set will be sorted in descending order according to the size of the information entropy, and then the samples with the largest information entropy and a proportion of α in the total samples among the multiple original samples will be selected as the representative sample data set. Since the larger the information entropy, the higher the uncertainty that the original sample is malware, that is, the more difficult it is to determine whether the original sample is malware. Through the above method, the selection of the representative sample data set is more targeted and has a wider coverage range, so as to meet the requirement that the difference between the malware detection rate and the original malware detection rate is within the first preset range, and further ensure the accuracy of subsequent training.
[0039] S140. Input the representative sample data set into the preset training model for training to obtain a malware detector.
[0040] Specifically, the preset training model in this embodiment can be a detection model constructed based on the AdaBoost algorithm.
[0041] Compared with the related art, the embodiments of the present invention have at least the following advantages: By selecting a representative sample data set from the original sample data set according to the characteristic parameters of the original samples, on the one hand, the number of sample data input for training the preset model can be reduced, so that the model does not need to train a large amount of data, thereby reducing the difficulty of model training; on the other hand, it can ensure that the difference between the malware detection rate and the original malware detection rate is within the first preset range, so that training with the representative sample data set can achieve the same training effect as training with the original sample data set, thereby ensuring the accuracy of the preset model after training; in addition, the above model training method does not require manual analysis, which can greatly reduce the labor cost.
[0042] Embodiment 2
[0043] Figure 2 It is a flowchart of a malware detector training method provided by Embodiment 2 of the present invention. This embodiment is an example of the foregoing embodiment, and specifically illustrates how to ensure that the difference between the malware detection rate and the original malware detection rate is within the first preset range.
[0044] Specifically, as Figure 2 shown, the method includes:
[0045] S210. Obtain the original sample data set and obtain the original malware detection rate of the original sample data set.
[0046] S220. Obtain the characteristic parameters of each original sample.
[0047] S230. According to the characteristic parameters, select a representative sample data set with a proportion of α in the total samples from the original sample data set, and obtain the malware detection rate of the representative sample data set.
[0048] S240. Determine whether the difference between the malware detection rate and the original malware detection rate is within the first preset range. If so, execute step S260; if not, execute step S250.
[0049] Specifically, the first preset range can be set according to actual needs. For example, it can be set to 0-0.3%. This embodiment does not specifically limit the size of the first preset range.
[0050] S250. Adjust the size of α and execute step S230.
[0051] Specifically, if the difference between the malware detection rate and the original malware detection rate is not within the first preset range, α will be continuously increased until the difference between the new malware detection rate and the original malware detection rate is within the first preset range; if the initially selected α can satisfy that the difference between the new malware detection rate and the original malware detection rate is within the first preset range, α can be appropriately reduced to minimize the number of representative samples as much as possible, thereby minimizing the difficulty of model training.
[0052] S260. Input the representative sample dataset into a preset training model for training to obtain a malware detector.
[0053] It is not difficult to find that steps S210 to S230 and S260 in this embodiment are the same as steps S110 to S140 in the foregoing embodiment. To avoid repetition, they will not be elaborated here.
[0054] Embodiment III
[0055] Figure 3 It is a flowchart of a method for training a malware detector provided in Embodiment III of the present invention. This embodiment is an example illustration of the foregoing embodiment, specifically illustrating how to obtain a malware detector.
[0056] Specifically, as Figure 3 shown, the method includes:
[0057] S310. Obtain an original sample dataset and obtain the original malware detection rate of the original sample dataset.
[0058] S320. Obtain the feature parameters of each original sample.
[0059] S330. According to the feature parameters, select a representative sample dataset with a proportion of α in the total samples from the original sample dataset, and obtain the malware detection rate of the representative sample dataset.
[0060] S340. Input the representative sample dataset into a detection model based on the AdaBoost algorithm to extract initial detection rules.
[0061] Specifically, in this embodiment, the detection model based on the AdaBoost algorithm constructs multiple interdependent decision trees, and finally makes a decision on the classification result by weighting all the decision trees. Sample selection and attribute selection, as two random processes for generating decision trees, can effectively reduce the problem of overfitting. At the same time, by combining multiple trees to detect malware, the problem of underfitting caused by single-tree discrimination is avoided, and the detection effect can be significantly improved.
[0062] For ease of understanding, the method for extracting initial rules in this embodiment will be specifically described below:
[0063] Rule r = {if f 1 ∩ f 2 ∩ f 3 … ∩ f n then result}, which consists of a rule body and a detection result. Among them, the rule body C = {f 1 ∩ f 2 ∩ f 3 … ∩ f n} is a Boolean expression between logical connectives extracted from multiple nodes of a tree model, and the detection result is that the application is malicious or benign. When the leaf nodes of the random tree exist, all leaf nodes and the root node form a rule. For example, "if android.intent.category.test > 0.5 Then malware". Among them, "android.intent.category.test > 0.5" is a logical connective, "&" is a Boolean expression, and "malware" is the detection result.
[0064] S350. Remove the redundant logical connectives in each initial detection rule, and use the initial detection rule after removing the redundant logical connectives as the refined detection rule.
[0065] Specifically, in this embodiment, the initial detection rule can be pruned by the leave-one-out pruning method to remove the redundant logical connectives in the initial detection rule.
[0066] The leave-one-out pruning method is specifically as follows: Obtain the initial error rate of the initial detection rule and the multiple error rates after removing each logical connective from the initial detection rule; Determine whether the difference between each error rate and the initial error rate is within a second preset range; When it is determined that it is not within the preset range, regard the logical connective corresponding to this error rate as a redundant logical connective; Remove the redundant logical connectives.
[0067] It should be noted that the second preset range can be set according to actual needs. For example, it can be set to 0 - 0.1%. This embodiment does not specifically limit the size of the second preset range.
[0068] Please refer to Table 1, which is the code for the refined rule extraction method in this embodiment:
[0069]
[0070] Table 1
[0071] For ease of understanding, the following is a specific example of how this embodiment removes redundant logical connectives:
[0072] Assume that the second preset range is 0.05%, the initial detection rule is A and B and C or D (A, B, C, and D are all rules, and and or are logical connectives), and the initial error rate of the initial rule is 2%. Then, remove A, B, C, and D separately to obtain the first rule: B and C or D; the second rule: A and C or D; the third rule: A and B or D; the fourth rule: A and B and C. The error rate of the first rule is 2.1%, the error rate of the second rule is 2.02%, the error rate of the third rule is 2.08%, and the error rate of the fourth rule is 2.2%. Then, the logical connective corresponding to rule B is a redundant logical connective and is removed. And so on, until all redundant logical connectives in the initial detection rule are removed.
[0073] S360. Construct a malware detector according to the refined detection rule.
[0074] It is not difficult to find that steps S310 to S330 and S360 in this embodiment are the same as steps S210 to S230 and S260 in the foregoing embodiment. To avoid repetition, they will not be elaborated here.
[0075] Embodiment 4
[0076] Figure 4 The figure is a flowchart of a method for training a malware detector provided in Embodiment 4 of the present invention. This embodiment makes further improvements on the basis of the foregoing embodiment. The specific improvement lies in: after obtaining the refined rule, redundant rules in the refined rule are also removed, and then a malware detector is constructed according to the refined detection rule after removing the redundant rules. In this way, effective detection of malware can be achieved and the interpretability of the malware detector can be improved.
[0077] Specifically, as Figure 4 shown, the method includes:
[0078] S410. Obtain the original sample data set and obtain the original malware detection rate of the original sample data set.
[0079] S420. Obtain the feature parameters of each original sample.
[0080] S430. According to the feature parameters, select a representative sample data set with a proportion of α in the total samples from the original sample data set, and obtain the malware detection rate of the representative sample data set.
[0081] S440. Input the representative sample data set into the detection model based on the AdaBoost algorithm to extract the initial detection rule.
[0082] S450. Remove the redundant logical connectives in each initial detection rule, and use the initial detection rule after removing the redundant logical connectives as the refined detection rule.
[0083] S460. Obtain the evaluation index parameters of the refined detection rule.
[0084] Specifically, the evaluation index parameters in this embodiment include one of the following or any combination thereof: rule occurrence frequency, rule error rate, and rule length.
[0085] More specifically, the rule occurrence frequency is the proportion of the number of samples that satisfy the rule to the total number of samples; the rule error rate is the error rate caused by using the rule for classification, and the smaller the error rate, the better the ability of the rule to detect malware; the rule length is the number of logical connectives in the rule, and the smaller the number, the higher the readability and the stronger the interpretability of the rule.
[0086] S470. Remove the redundant rules in the refined detection rule according to the evaluation index parameters.
[0087] Specifically, first, among the detection rules that generate conflicts, select the rules with low error rate, high frequency, and short length. Then construct a rule matrix including sample sequences, rules, and sample labels as the rule training set, as shown in Table 2. Among them, the sample serial number refers to malware or benign software. The value of 1 in the matrix represents that the sample conforms to the corresponding rule, and the value of 0 in the matrix represents that the sample does not conform to the corresponding rule. Then use the rule training set to train the Adaboost classifier to obtain the importance of the rules. Finally, by sorting the importance of the rules in descending order, the most important rules can be obtained, and the redundant rules can be removed.
[0088]
[0089] Table 2
[0090] S480. Construct a malware detector according to the refined detection rule after removing the redundant rules.
[0091] Please refer to Figure 5 , for the sake of easy understanding, the following is a specific description of the malware detector training method of this example:
[0092] The principle framework of the malware detector training method mainly includes two modules, namely representative sample selection and rule detection model construction. In the representative sample selection module, information entropy is used to select samples, redundant samples are removed, and representative samples are selected as a new data set; in the detection rule extraction module, the Adaboost model is trained using the new data set, and then operations such as initial rule extraction, rule measurement, rule pruning, and redundant rule removal are performed on the trained random forest model to obtain a refined rule set. Finally, a malware rule detector is constructed using the refined rule set to detect Android malware, and the detection result of whether the application is malicious or benign, as well as the explanation result of whether the application is malicious or benign, are obtained.
[0093] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this patent.
[0094] Embodiment 5
[0095] Figure 6 The following is a schematic structural diagram of a malware detector provided in Embodiment 5 of the present invention. As Figure 6 shown, the malware detector includes:
[0096] An original sample detection rate acquisition module 1, configured to obtain an original sample data set and obtain the original malware detection rate of the original sample data set, where the original sample data set includes a plurality of original samples;
[0097] A feature parameter acquisition module 2, configured to obtain the feature parameters of each of the original samples, where the feature parameters are used to characterize the uncertainty degree of the original sample being a malware;
[0098] A representative sample detection rate acquisition module 3, configured to select a representative sample data set with a proportion of α in the total samples from the original sample data set according to the feature parameters, and obtain the malware detection rate of the representative samples, where α is greater than 0 and less than 1, and the difference between the malware detection rate and the original software detection rate is within a first preset range;
[0099] A detector training module 4, configured to input the representative sample data set into a preset training model for training to obtain a malware detector.
[0100] It is not difficult to find that this embodiment is a device embodiment corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and in order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.
[0101] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of the present invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other units in this embodiment.
[0102] The following is a specific experimental analysis of the malware detector provided in this embodiment:
[0103] First, 80% of the data is used as the training set and 20% of the data is used as the test set.
[0104] Secondly, the training set is used to select the parameters of the malware detection model based on rule extraction. When selecting representative samples, the input features are 274 dimensions and the representative sample ratio is 0.5. When training the random forest model, the ten-fold cross-validation method is used to adjust the parameters. 80% of the data in the training set is used as training data and 20% of the data is used as validation data to optimize the model parameters. The final number of trees is 100, and the number of features randomly selected for each decision tree is 241. When extracting rules from the detection model based on the AdaBoost algorithm, the ten-fold cross-validation method is used to adjust the parameters. 80% of the data in the training set is used as training data and 20% of the data is used as validation data to optimize the model parameters. The minimum frequency threshold of the rule is 5e-04, the error rate threshold is 0.04, and the rule length threshold is 3.
[0105] Finally, the test set is used to calculate the evaluation indicators of the explainable malware detection rule extraction method (RBE) and 9 comparison methods, including accuracy, precision, recall and F value. Among them, the 9 comparison methods include: decision tree (DT), gradient boosting tree method (GDBT), neural network method (MLP), logistic regression method (LR), boosting tree method (ADABOOST) and Bayesian algorithm (NB), RFRULES method, EBBAM method and Sigpid method. The 6 detection tools include: AntiVir, AVG, BitDefender, ClamAV, ESET, F-Secure.
[0106]
[0107] Table 3 Experimental Results of Malware Detection
[0108]
[0109] Table 4 Comparative Experimental Results of Malware Detection Tools
[0110] It can be seen from the experimental results that:
[0111] (1) The method for extracting interpretable malware detection rules is superior to the comparative method. Comparing the malware detector training method of the foregoing embodiment with other basic classification algorithms, the experimental results show that the precision rate (0.974), recall value (0.982), and F1 value (0.978) of the malware detection algorithm based on rule extraction are all higher than those of other algorithms. This method effectively detects malware by automatically extracting a small number of decision rules with low complexity and high accuracy from the tree model, mining the Boolean logical relationship between features and the relationship between detection results, and has a better detection effect compared to the single model in other comparative methods.
[0112] (2) The method for extracting interpretable malware detection rules is superior to the comparative malware detection tools. Comparing the malware detector of this embodiment with malware detection tools, the comparative analysis tools are AntiVir, AVG, BitDefender, ClamAV, ESET, and F-Secure. The detection rate (recall rate) of the interpretable Android malware detection rule extraction method is 98.2%, which is higher than that of 6 detection tools. The detection tools use the rules summarized by experts for detection, and the timeliness of the rules is insufficient, and the ability to detect malware is insufficient. For example, the F-Secre tool can only reach a detection rate (recall rate) of 64.16%. Compared with the detection tools, the interpretable Android malware detection rule extraction method extracts rules from the machine learning model, has timeliness, and can effectively detect malware.
[0113] Taking case analysis, the number of rules, and the degree of interpretability as evaluation indicators of the experimental results, where, for example, the explanation is to use the explanation result to explain a specific malware family and evaluate whether the explanation result conforms to the family characteristics, and the number of rules is the number of rules in the rule set. The formula for calculating the interpretability of the rule set is as follows:
[0114]
[0115] where RuleSet is the rule set, weight i is the weight of a single rule, is the interpretability of a single rule, and i is the number of rules.
[0116]
[0117] Among them, the value of maxAttribute is the number of attributes, and the value of curCondition is the number of conditions for a single rule.
[0118] First, 80% of the data in the DREBIN dataset is used as the training set, and 20% of the data is used as the test set. Then, the training set is used to select the parameters of the Android malware detection model extracted based on rules. The 274-dimensional features selected from the DREBIN dataset are used as the input feature set.
[0119] When selecting representative samples, the random forest algorithm is used to calculate the classification probability of the samples. The number of trees is 100, and the best parameter for the sample ratio is 0.5. When training the random forest model, the ten-fold cross-validation method is used for parameter tuning. 80% of the data in the training set is alternately used as the training data, and 20% of the data is used as the validation data to optimize the model parameters. The number of trees is 100, and the number of features randomly selected for each decision tree is 241.
[0120] When extracting rules from the detection model based on the AdaBoost algorithm, the ten-fold cross-validation method is used for parameter tuning. 80% of the data in the training set is alternately used as the training data, and 20% of the data is used as the validation data to optimize the model parameters. Among them, the minimum frequency threshold for rules is 5e-04, the error rate threshold is 0.04, the rule length threshold is 3, and the number of trees in the AdaBoost algorithm is 100.
[0121] Secondly, the test set is used to calculate the evaluation metrics, namely the interpretability and the number of rules, of the interpretable Android malware detection rule extraction method (RBE) and the RFRULES (2019) method.
[0122] Finally, the FakeInstall family (925 malware) and Dowgin (3385 malware) in the DREBIN dataset are respectively selected as new datasets, and the above process is repeated to output the interpretation results of the interpretable Android malware detection rule extraction method (RBE) and the EBBAM method (2018).
[0123]
[0124] Table 5 Results of Comparative Experiments on Malware Interpretability
[0125]
[0126] Table 6 Results of Comparative Experiments on Malware Interpretability (Fake Installer Family)
[0127]
[0128] Table 7 Comparison Experiment Results of Malware Interpretability (Dowgin Family)
[0129] It can be seen from the experimental results that:
[0130] The interpretability of the rule extraction method for interpretable malware detection (RBE method) is better than that of the RFRULES method. The number of rules of the RBE method (34) is 221 less than that of the RFRULES method (255), and the interpretability (99.73%) is 1.04% higher than that of the comparative method. Therefore, the method in this paper has better interpretability than the RFRULES method.
[0131] The interpretability of the rule extraction method for interpretable malware detection (RBE method) is better than that of the EBBAM method. Taking the FakeInstaller malware family and the Dowgin malware family as examples for comparative analysis.
[0132] The Fake Installer malware family has two malicious behaviors: (1) Charging without the user's permission; (2) The program backdoor remotely controls the user's mobile phone. It can be seen from the experimental results in Table 6 that the rule "If Permission:SEND_SMS>0.5∩android.intent.action.VIEW<=0.5" shows that this malicious family obtains the SMS permission without the user's consent, while the EBBAM method simply uses the SEND_SMS permission to explain this behavior. This permission is also commonly used by benign software, which only shows that this feature contributes greatly to the classifier, but it is impossible to directly use this feature to distinguish whether the software is malicious or benign. The rule "android.intent.Test>0.5∩android.app.KeyguardManager.exitKeyguardSecurely≤0.5" shows that when the Android system is in the test mode and the unlock mode, the remote server may control the Android system, while the EBBAM method only uses the READ PHONE STATE permission to explain this behavior, and it is difficult to directly establish a logical relationship between this feature and the detection result.
[0133] The Dowgin malware family is an advertising malware bundled with some applications. This family has two malicious behaviors: (1) The device continuously downloads and installs other malware, causing the user's mobile phone to freeze and affecting the normal use of the mobile phone; (2) Sending the user's device information to the remote end. It can be seen from the experimental results in Table 7 that
[0134] The rule "android.app.DownloadManager.addCompletedDownload > 1.5 ∩ android.intent.action.ACTION_PACKAGE_ADDED <= 0.5" indicates that the application downloads malware without the user's consent.
[0135] "android.app.NotificationManager.notify > 4" shows that the malware displays messages in the notification bar more than 4 times. Additionally,
[0136] "android.telephony.TelephonyManager.getSimSerialNumber > 1.5 ∩ java.net.URL.openStream <= 0.5" shows that the malware obtains the user's number and sends it to a certain link. In the EBBAM method, the most important 4 features have no direct relation to these two behaviors and are features that normal software would use.
[0137] In summary, the malware detector training method in the foregoing embodiments can reflect the interaction between features and the causal relationship between detection results, has a lower rule complexity, and has higher interpretability.
[0138] Embodiment Six
[0139] Figure 7 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0140] As Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0141] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0142] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the malware detector training method.
[0143] In some embodiments, the malware detector training method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the malware detector training method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the malware detector training method by any other appropriate means (e.g., by means of firmware).
[0144] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0145] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0146] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0148] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0149] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0150] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0151] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for training a malware detector, characterized in that, it includes: Obtain an original sample data set and obtain the original malware detection rate of the original sample data set, where the original sample data set includes multiple original samples; Obtain the feature parameters of each of the original samples, where the feature parameters are used to characterize the uncertainty degree of the original sample being malware; According to the feature parameters, select a representative sample data set with a proportion of α in the total samples from the original sample data set, and obtain the malware detection rate of the representative sample data set, where α is greater than 0 and less than 1, and the difference between the malware detection rate and the original malware detection rate is within a first preset range; Input the representative sample data set into a preset training model for training to obtain a malware detector; Among them, the step of inputting the representative sample data set into a preset training model for training to obtain a malware detector includes: Input the representative sample data set into a detection model based on the AdaBoost algorithm to extract initial detection rules, where the initial detection rules are feature expressions connected by multiple logical connectives; Remove the redundant logical connectives in each of the initial detection rules, and use the initial detection rules after removing the redundant logical connectives as refined detection rules; Construct the malware detector according to the refined detection rules.
2. The method for training a malware detector according to claim 1, characterized in that, the feature parameter is information entropy; The step of obtaining the feature parameters of each of the original samples includes: Input multiple of the original samples into a preset training classifier to obtain the probability of each of the original samples being classified as malware or benign software; Obtain the information entropy according to the following formula: where n is the number of original samples, i is the serial number of the original sample, p(y i ) is the probability that the original sample is classified as malware or benign software, and H(Y) is the information entropy.
3. The method for training a malware detector according to claim 2, characterized in that, The step of selecting a representative sample data set with a proportion of α in the total samples from the original sample data set according to the feature parameters includes: Arrange the multiple original samples in the original sample data set in descending order according to the magnitude of the information entropy; Select the samples with the largest information entropy and a proportion of α in the total samples among the multiple original samples as the representative sample data set.
4. The method for training a malware detector according to any one of claims 1-3, characterized in that, after obtaining the malware detection rate of the representative sample data set, it further includes: Judge whether the difference between the malware detection rate and the original malware detection rate is within the first preset range; When it is determined that it is within the first preset range, then execute the step of inputting the representative sample data set into a preset training model; When it is determined that it is not within the first preset range, adjust the magnitude of α to obtain a new malware detection rate until the difference between the new malware detection rate and the original malware detection rate is within the first preset range.
5. The method for training a malware detector according to claim 1, characterized in that, The step of removing the redundant logical connectives in each of the initial detection rules includes: Obtain the initial error rate of the initial detection rule and multiple error rates after removing each logical connective from the initial detection rule; Respectively determine whether the difference between each of the error rates and the initial error rate is within a second preset range; when it is determined that it is not within the preset range, use the logical connective corresponding to the error rate as the redundant logical connective; Remove the redundant logical connectives.
6. The malware detector training method according to claim 1, wherein, before constructing the malware detector according to the refined detection rule, further comprising: Obtain the evaluation index parameters of the refined detection rule; Remove redundant rules from the refined detection rule according to the evaluation index parameters; The constructing the malware detector according to the refined detection rule includes: Construct the malware detector according to the refined detection rule after removing the redundant rules.
7. The malware detector training method according to claim 6, wherein, The evaluation index parameters include one or any combination of the following: Rule occurrence frequency, rule error rate, and rule length.
8. A malware detector, wherein, comprising: An original sample detection rate acquisition module, configured to acquire an original sample data set and obtain the original malware detection rate of the original sample data set, wherein the original sample data set includes a plurality of original samples; A feature parameter acquisition module, configured to acquire the feature parameters of each of the original samples, and the feature parameters are used to characterize the uncertainty degree of the original sample being a malware; A representative sample detection rate acquisition module, configured to select a representative sample data set with a proportion of α in the total samples from the original sample data set according to the feature parameters, and obtain the malware detection rate of the representative samples, wherein α is greater than 0 and less than 1, and the difference between the malware detection rate and the original malware detection rate is within a first preset range; A detector training module, configured to input the representative sample data set into a preset training model for training to obtain a malware detector; The detector training module is specifically configured to input the representative sample data set into a detection model based on the AdaBoost algorithm, extract an initial detection rule, wherein the initial detection rule is a feature expression connected by a plurality of logical connectives; remove redundant logical connectives in each of the initial detection rules, and use the initial detection rule after removing the redundant logical connectives as the refined detection rule; construct the malware detector according to the refined detection rule.
9. An electronic device, wherein, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the malware detector training method according to any one of claims 1-7.
10. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer instructions for causing a processor to implement the malware detector training method according to any one of claims 1-7 when executed.
Citation Information
Patent Citations
Model training method and device and computer readable storage medium
CN112529210A
Adaboost-based Android malicious software detection method and system and storage medium
CN113704759A