Android malicious software classification method based on multi-path feature fusion and active learning
By introducing multipath feature fusion and active learning strategies into the Android malware detection model, the problem of excessive parameters and concept drift in the deep learning model is solved, and efficient malware recognition and generalization capabilities are improved.
Patent Information
- Application Number
- CN202411773470.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-06
AI Technical Summary
In Android malware detection, the existing deep learning models have high computing and storage costs, increased risk of overfitting, and faced the problem of malware concept drift, which reduces the generalization ability of the model.
A detection and classification framework for multi-path feature fusion is proposed. By increasing the depth and width of the network layer, the amount of model parameters is reduced, and the active learning strategy is introduced, and the concept drift problem is dealt with by timely updating the model and sample selection strategy based on uncertainty metrics.
While reducing the amount of model parameters, it improves the accuracy of malware recognition, effectively alleviates the impact of concept drift on model performance, and reduces sample marking costs.
Smart Images

Figure CN119942166A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Android malware detection and classification, proposes a detection and classification framework of multi-path feature fusion, and introduces active learning on this basis to detect and analyze Android malware. Background Art
[0002] Deep learning is a branch of machine learning that achieves learning and understanding of complex tasks by simulating the neural networks and functions of the human brain. The core idea is to build a deep neural network, abstract and hierarchically represent information through a multi-layer neural network, and learn features and patterns from a large amount of data. Compared with traditional machine learning methods, deep learning can automatically learn the features in the input data without manually designing feature extractors. The multi-level structure of the neural network allows the model to automatically learn and represent the hierarchical features of the data. Several common deep learning models in the field of Android malware detection include Convolutional Neural Network (CNN), VGG (Visual Geometry Group), RseNet (Residual Networks), Multilayer Perceptron (MLP), etc.
[0003] With the popularity of smartphones and the rapid development of mobile applications, the Android platform has become the main target of malware attacks. In order to effectively identify and defend against Android malware, researchers have developed a variety of detection techniques, including static analysis, dynamic analysis, hybrid analysis methods, and image analysis-based methods. Good results have been achieved using deep learning models for detection and analysis. However, while pursuing the effect, the number of parameters has also increased due to the continuous expansion of the model size, which will lead to expensive computing and storage costs, increased risk of overfitting, and other problems. The large number of parameters requires a lot of computing resources in the training and inference stages, and too many parameters may make the model more likely to be overly dependent on training data, reducing generalization performance.
[0004] Active learning is a strategy in machine learning that actively selects a portion of samples for labeling to improve the performance of the model on the sample set. Traditional supervised learning usually requires a large number of labeled samples, while active learning attempts to minimize the labeling cost and only labels those samples that are most beneficial to model learning. There are two different strategies in the field of active learning: flow-based active learning and pool-based active learning.
[0005] Due to the ever-changing and evolving characteristics of current malware, the field of Android malware classification faces the problem of concept drift, which is a very severe challenge for the security field. The existence of concept drift reduces the detection and classification performance of most models, because most models usually rely on known sample features to build classification models, but as time changes, when the feature distribution of samples in certain categories changes, the classification effect of the model for these newly appeared samples will be greatly affected, resulting in a decrease in the generalization ability of the model. At this time, the newly appeared samples need to be added to the training set and the model needs to be retrained to adapt to the new samples, but this often consumes a lot of sample labeling costs. The active learning strategy effectively makes up for the shortcomings. As time goes by, the model is dynamically updated to improve the generalization ability of the model while minimizing the labeling cost of unlabeled samples. Summary of the invention
[0006] In the field of medical image segmentation, Yamanakkanava et al. proposed an MF2-Net framework with an encoder-decoder deep learning structure, which includes multiple encoder paths integrated with the SGC module to achieve efficient segmentation performance by extracting multi-scale features. Inspired by this, the present invention proposes a detection and classification framework with multi-path feature fusion. By increasing the depth and width of the network layer, the number of parameters of the deep learning detection model is greatly reduced, while further improving the accuracy of malware identification. On this basis, active learning is introduced for detection and analysis to deal with the problem of malware concept drift in the real environment. It not only improves the generalization ability by updating the model, but also effectively reduces the labeling requirements of samples.
[0007] The specific steps are as follows:
[0008] S1 converts the executable file of the application into a grayscale image;
[0009] S2 designs multiple different feature extraction paths, decomposing the k×k convolution operation into k×1 and 1×k double-layer convolution operations to reduce the number of model parameters, while using multi-scale convolution operations to allow the model to learn features;
[0010] S3 fuses the features extracted from multiple paths to train the classification model;
[0011] S4 updates the model regularly by setting an active learning iteration cycle and adopts a sample selection strategy based on uncertainty measurement. Each time the model is updated, only high-value samples are selected to update the model.
[0012] Furthermore, step S1 specifically includes:
[0013] S11 uses the APK object method provided by the Androguard library to find the classes.dex file contained in the APK file, and then extracts the byte stream in the dex file through the get_bytes function. Each byte in the byte stream corresponds to a pixel value in the grayscale image. Finally, the byte stream in each application dex file is converted into a corresponding grayscale image. The specific steps are as follows:
[0014] S101 extracts byte stream: obtains all byte data in the dex file through the get_bytes function, and then appends each byte data one by one to the bytes type stream to generate the byte stream of the dex file;
[0015] S102 extracts grayscale values: each byte in the byte stream represents a pixel value, ranging from 0 to 255, and these pixel values constitute the pixel information of the original image;
[0016] S103 constructs the original image: constructs the obtained pixel information into a 1×n grayscale image, where n is the length of the byte stream;
[0017] S104 calculates the coordinate mapping of the target pixel in the target image: the original image is further adjusted to a target image of size m×m, and the coordinates of each pixel in the target image in the original image are first calculated to determine the position of the target pixel in the original image for subsequent interpolation calculation;
[0018] S105 performs interpolation calculation: according to the position of the target pixel in the original image, a bilinear interpolation algorithm is used to perform weighted average of the grayscale values of four adjacent pixels in the original image to obtain an interpolation result corresponding to the target pixel;
[0019] S106 Obtaining the grayscale value of the target pixel: taking the interpolation result obtained for the target pixel as its corresponding grayscale value;
[0020] S107 repeats steps S104 to S106 for all pixels in the target image until the pixel values of the entire target image are interpolated and calculated, and finally an image of size m×m is obtained, and the grayscale image conversion is completed for downstream tasks;
[0021] S12 bilinear interpolation is based on the values of four adjacent points in the matrix area. Assume that the coordinates of a pixel point P in the target image are (x, y), and the four adjacent points corresponding to this point in the original image are Q 11 , Q 12 , Q 21 , Q 22 , then we have the following formula:
[0022]
[0023] Among them, u represents the relative position of the target pixel (x, y) in the horizontal direction in the original image, v represents the relative position in the vertical direction, (x1, y1) and (x2, y2) are Q 11 and Q 22 The coordinates of the target point are then calculated using the bilinear interpolation algorithm, namely:
[0024] f(x,y)=(1-u)(1-v)Q 11 +u(1-v)Q 21 +(1-u)vQ 12 +uvQ 22 (2)
[0025] Where f represents the bilinear interpolation function, Q 11 , Q 21 , Q 12 and Q 22 Represents Q 11 , Q 12 , Q 21 and Q 22 The value of the pixel.
[0026] Furthermore, step S2 specifically includes:
[0027] S21 introduces the context extraction module, decomposing the convolution kernel of size k×k into two smaller one-dimensional convolution kernels, k×1 and 1×k. CEM also uses a stack of multiple convolution kernels of different sizes (k=1, 3, 5, 7) to increase the width of the network layer. The number of convolution kernels in all convolution layers is 32, and RuLU is used as the activation function. Zero padding is used to ensure that the input and output sizes are consistent. The output of the last convolution layer corresponding to different convolution kernels is added as the output of CEM. Specifically, the output of CEM CEM o Calculated by the following formula:
[0028]
[0029] Among them, x l is the input grayscale image, w l k×1 is the weight associated with the convolution kernel of size k×1, b l is the bias term, * is the convolution operation with RuLU as the activation function;
[0030] The output of the context extraction module S22 first passes through the maximum pooling layer, and then the output of the pooling layer CEM mThen it is passed to the intermediate module of the next layer. IM uses convolution kernels of size 3×3 and 5×5. The operation of 3×3 convolution kernel is decomposed into convolution operations of size 3×1 and 1×3. Similarly, the operation of 5×5 convolution kernel is also decomposed into convolution operations of size 5×1 and 1×5. Each convolution layer uses RuLU as the activation function and contains 64 convolution kernels. Their outputs are added element by element to form the output of IM. o , the specific formula is as follows:
[0031]
[0032] Among them, CEM m The output of the context extraction module is pooled and used as the input of IM. l 1×k is the weight associated with the convolution kernel of size 1×k;
[0033] The output of the S23 intermediate module is passed to the local extraction module through the pooling layer. The LEM contains smaller 1×1 and 3×3 convolution kernels, where the 3×3 convolution operation is decomposed into a double-layer convolution operation of 3×1 and 1×3. Each convolution layer has 128 convolution kernels, and RuLU is used as the activation function. The feature maps of the convolution layers corresponding to different convolution kernels are then added element by element to finally obtain the output of the LEM. The output of the LEM o Calculated by the following formula:
[0034]
[0035] Among them, IM m is the output of the intermediate module and serves as the input of LEM.
[0036] Furthermore, step S3 specifically includes:
[0037] S31 extracts features from the same image through three paths respectively, and finally fuses these three different features to obtain image features. Here, the three paths are defined as P1, P2 and P3, which correspond to the three paths from left to right of the multi-path feature extraction module. Then P1 contains three sub-modules and three pooling layers, P2 contains two sub-modules, CEM and IM, and two pooling layers, and P3 contains a CEM sub-module and a pooling layer. Let the output of CEM be CEM o , the output of IM is IM o , the output of LEM is LEM o , where P 1o , P 2o and P 3oRepresent the outputs of P1, P2 and P3 respectively, MP is the maximum pooling operation function, and the pooling window size is 2×2. Then the outputs of P1, P2 and P3 are calculated by the following formula:
[0038] P 10 =MP(LEM0,2) (6)
[0039] P 20 =MP(IM0,2) (7)
[0040] P 30 =MP(CEM0,2) (8)
[0041] For the feature fusion method, splicing fusion is adopted, and the convolution layer and pooling layer are used to adjust the size and channel of the feature map. In order to achieve the final feature splicing fusion, P 2o and P 3o The dimensions are resized to match P 1o Same size, first, P 2o and P 3o The convolution operation with the convolution kernel size of 1×1 and RuLU as the activation function is performed to adjust the number of channels to 128, and then the pooling operation is performed again. 2o , the output of the convolutional layer is subjected to a maximum pooling operation with a pooling window size of 2×2, P 3o Then a maximum pooling operation of size 4×4 is performed, and finally the three sets of feature maps of the same size are spliced together according to a specific axis. The specific formula is as follows:
[0042] Fusion=Conc(P 10 , P 20 , P 30 ) (9)
[0043] Among them, Fusion is the feature obtained after splicing and fusion, and Conc is the feature splicing function;
[0044] S32 performs convolution and pooling operations on the fused features again with RuLU as the activation function. For binary classification tasks, the number of neurons in the output layer is 1, Sigmoid is used as the activation function, and the final output value range is (0, 1); in multi-classification tasks, the number of neurons in the output layer depends on the number of categories, and Softmax is used as the activation function. The model is trained with the cross entropy loss function, and the final output is the probability of each category, and the sum of these probabilities is 1.
[0045] Furthermore, step S4 specifically includes:
[0046] S41 first trains the classification model with labeled samples. When new samples appear and meet the time threshold, the trained model is used to predict these samples. Then, through the designed sample selection strategy, a part of high-value samples that have the greatest impact on the model classification performance are selected for manual labeling. Finally, these samples are used to update the model. The specific steps are as follows:
[0047] S401: Build an initial training set and mark it to obtain a marked sample set D;
[0048] S402 trains a classification model using a labeled sample set D;
[0049] S403 determines whether the time threshold is reached. If the condition is met, the newly appeared unlabeled sample set D n Make predictions using the trained model;
[0050] S404 selects a portion of high-value samples through a sample selection strategy based on the prediction results to form a sample set D h ;
[0051] S405 for sample set D h All samples in the dataset are manually labeled to determine their categories;
[0052] S406: The labeled sample set D h Add to the initial training D to form a new sample set D s , retrain the model with this sample set;
[0053] The above steps S403 to S406 are a process of updating the model once. As time goes by, when the newly appeared unlabeled samples reach the time threshold again, the next model update will be performed.
[0054] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0055] 1. Aiming at the problem of too many parameters in deep learning models, an Android malware detection and classification scheme based on multi-path feature fusion is proposed. The binary executable file of the application is converted into a grayscale image, and multiple different paths are designed for feature extraction. During the extraction process, the k×k convolution operation is decomposed into k×1 and 1×k double-layer convolution operations, and multi-scale convolution operations are used to increase the depth and width of the network, thereby reducing the number of model parameters while ensuring the overall performance of the network. After multiple sets of comparative experiments, it is proved that the method of the present invention can achieve higher classification performance in a resource-constrained environment.
[0056] 2. Aiming at the concept drift problem caused by the continuous evolution of malware, an Android malware classification scheme based on active learning is proposed. By setting a reasonable active learning iteration cycle to complete the regular update of the model, the purpose of updating the model is to let the model learn the data and feature distribution of newly emerging malware to improve the generalization ability of the model on new data sets. Experiments have proved that the active learning strategy proposed in this invention can effectively alleviate the problem of model performance degradation caused by concept drift.
[0057] 3. In response to the newly emerged problem of high sample labeling costs, the present invention adopts a sample selection strategy based on uncertainty measurement. When updating the model, the sample selection strategy is used to select high-value samples that are most beneficial to the model, that is, samples with the highest uncertainty. After multiple sets of comparative experiments, it is proved that this strategy can achieve effective model updates and higher classification performance under the realistic conditions of lower labeled sample requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is the sample distribution diagram of the AMD dataset in each year.
[0059] Figure 2 This is a graph showing the performance changes of the model on the 2014 dataset. The changes in classification accuracy have been on an upward trend.
[0060] Figure 3 This is a comparison chart of model performance of different sampling strategies on the 2014 dataset. After four model updates, the uncertainty sampling strategy and the full sampling strategy finally achieved comparable detection results on the 2014 dataset, while the results of the random sampling strategy were significantly different from the first two.
[0061] Figure 4 This is a graph showing the performance changes of the model on the real dataset. DETAILED DESCRIPTION
[0062] To verify that the model proposed in this invention has certain advantages in performance and resource consumption, the effectiveness of model optimization in Android malware detection binary classification and malicious family multi-classification tasks, and the Android malware classification method based on active learning can improve the generalization ability of the model on the newly appeared sample set, and achieve effective model update and higher classification performance under the condition of low labeled sample requirements. The specific steps are as follows:
[0063] (1) Dataset:
[0064] (101) Detection and classification framework based on multi-path feature fusion: There are two target datasets, namely the S1 dataset and the Drebin dataset. The S1 dataset in Table 1 was collected by ourselves and contains a total of 13,322 samples. Among them, there are 6,979 benign samples and 6,343 malicious samples. All malicious samples come from the VirusShare dataset, which is a very large malware dataset that includes various malicious samples such as Android and Windows. All benign samples come from the 360 App Store, and each sample is tested for maliciousness through Virus Total to obtain a purer benign sample dataset.
[0065] Table 1S1 Dataset distribution
[0066]
[0067] Table 2 shows the distribution of each family in the Drebin dataset, which is widely used as a benchmark dataset.
[0068] Table 2 Distribution of families in the Drebin dataset
[0069]
[0070]
[0071] The Drebin dataset contains 5560 malicious samples from 179 different malicious families, and the year and month labels of the samples are between August 2010 and October 2012. This experiment will select malicious samples from 20 different families from the Drebin data. Due to the diverse design of applications, especially malicious applications designed to circumvent certain features, some of the applications cannot be successfully converted into images. Therefore, the Drebin dataset of this experiment contains a total of 4603 samples.
[0072] (102) Android malware classification based on active learning: The dataset comes from two public datasets, namely the AMD dataset and the VirusShare dataset. The AMD dataset contains 71 malware families and a total of 24,553 malicious samples. The sample year labels are from 2010 to 2016. In order to verify the effectiveness of the method proposed in this chapter, according to the experimental requirements, this section selects 2,602 samples from the AMD dataset for experiments, which include 6 malicious families with year labels from 2010 to 2014. Each year contains samples from these 6 families, constructing a relatively evenly distributed dataset. The sample distribution of each year is as follows: Figure 1 shown.
[0073] For each malware family, due to the limited AMD dataset, the sample distribution in the six families selected in this section is relatively unbalanced, and the details are shown in Table 3.
[0074] Table 3 Details of samples of each family in the AMD dataset
[0075]
[0076]
[0077] According to the experimental requirements, 466 malicious samples in the VirusShare dataset were finally selected as the real dataset of the experiment. The sample year labels are from 2017 to 2021. Since no samples of the "Lotoor" family were collected during this period, the real dataset only contains 5 families in the AMD dataset. The details are shown in Table 4.
[0078] Table 4 Details of samples from each family in the VirusShare dataset
[0079]
[0080] (2) Comparison system:
[0081] (201) The Android malware detection and classification model with multi-path feature fusion proposed in the present invention is adopted, and the S1 dataset and the Drebin dataset are used as the experimental datasets for binary classification and multi-classification respectively, of which 80% are used as training sets and 20% are used as test sets. Comparative experiments are carried out with four common deep learning models, namely CNN, VGG, ResNet and MLP. Since these four models have a high influence in the classification field, they are used to verify the effectiveness of the method of the present invention.
[0082] CNN: It uses the classic convolutional neural network architecture, which includes two convolutional layers, two pooling layers and one fully connected layer.
[0083] VGG: Contains two convolutional blocks and one fully connected layer, where each convolutional block contains two convolutional layers and one pooling layer.
[0084] ResNet: Contains two convolutional layers, two residual blocks, and one fully connected layer.
[0085] MLP: consists of only three fully connected layers.
[0086] Each fully connected layer in the above four models uses Dropout regularization technology to prevent overfitting.
[0087] (3) Experimental results and analysis:
[0088] (301) During the experiment, each model was tested ten times, and the one with the best performance was taken as the experimental result. The specific results of each model are shown in Tables 5 and 6.
[0089] Table 5S1 Model performance comparison details on the dataset
[0090]
[0091] As can be seen from Table 5, the method proposed in the present invention performs well in the binary classification task, and all evaluation indicators have great advantages over the other four models. As for the number of model parameters, the number of parameters of the model proposed in the present invention is less than 5% of that of VGG. This is because the k×k convolution operation in the network layer is decomposed into k×1 and 1×k double-layer convolution operations, which increases the depth of the network layer to reduce the number of model parameters, thereby reducing the computational cost of the experiment, reducing the training time and the required memory; in addition, the method uses multi-scale convolution operations to increase the width of the network layer so that the model can learn richer features, thereby improving the model classification performance, so all evaluation indicators are the highest. Experiments have proved that the method of the present invention is very beneficial for restricted experimental environments.
[0092] Table 6. Comparison of model performance on the Drebin dataset
[0093]
[0094] As can be seen from Table 6, the method proposed in the present invention is the best in terms of the three evaluation indicators of Accuracy, Recall and F1-Score, and the number of parameters is only 647,000, which has a great advantage over other models.
[0095] (302) Active learning was introduced to conduct experiments. The specific experimental results are shown in Table 7, where the evaluation indicator for each year is the classification accuracy of the model on the test set of that year.
[0096] Table 7 Comparison of classification performance in each year during model update process
[0097]
[0098] As can be seen from Table 7, after the model was updated once, the classification accuracy in 2011 increased. This is because the training set used for the model update includes the training set of that year. In the following years, the classification accuracy has increased compared to when the model was not updated. It is worth noting that the 2014 dataset was never involved in the training in the first three updates of the model, but with the continuous update of the model, the classification accuracy of that year has been on an upward trend.
[0099] from Figure 2It can be seen that the active learning strategy of the present invention is used to regularly update the model, which can effectively improve the classification accuracy of new samples. This also shows that the method of the present invention can effectively alleviate the problem of decreased model generalization ability caused by concept drift.
[0100] (303) The experiment will be compared with two methods, random sampling and full sampling. The number of samples selected for random sampling is 250, and full sampling means that during the model update process, all training sets of each year are added to the original training set to train the model. The specific experimental results are shown in Table 8, where US represents the uncertainty sampling used in this method, RS represents random sampling, FS represents full sampling, and the evaluation indicator is the classification accuracy of the model on the test set of each year.
[0101] Table 8 Effect of different sample selection strategies on model performance
[0102]
[0103]
[0104] It can be seen from Table 8 that when updating the model, if the same number of samples are selected using random sampling and uncertainty sampling sample selection strategies, the performance of the updated model is significantly different. This shows that random sampling requires more samples to improve the detection effect of the model, thereby increasing the sample labeling cost. For uncertainty sampling, only about 65% of the samples need to be selected to achieve an effect similar to that of all sampling, and the accuracy rate differs by less than 1%, which greatly reduces the cost of manual labeling.
[0105] from Figure 3 This proves once again that the method of the present invention can use an active learning strategy to regularly update the model when fewer samples are required, thereby improving the generalization ability of the model and reducing the required sample labeling cost.
[0106] (304) Experiments were conducted on a real dataset from VirusShare, which was used as the new malware after four model updates. Specifically, after each model update, the classification accuracy on the real dataset at the current moment was recorded to observe the changing trend of the classification performance on the real dataset over time and with model updates. At the same time, the active learning strategy of the present invention was used to update the model for the fifth time using samples from the real dataset.
[0107] Figure 4The performance of the model on the real dataset changes over time, with the horizontal axis representing the number of model updates and the vertical axis representing the classification accuracy. It can be seen that the performance of the model on the real dataset shows an upward trend over time, which indicates that the use of active learning strategies to regularly update the model can effectively improve the classification performance of the model on the newly appeared sample set. In addition, the year labels of the samples in the real dataset are from 2017 to 2021, and the first four updates only updated the data to 2014, so the classification accuracy on the real dataset is slightly lower, but this is enough to prove that effective model updates can be achieved at a lower labeling cost to alleviate the problems caused by concept drift.
Claims
1. An Android malware classification method based on multi-path feature fusion and active learning, characterized in that: The following steps are involved: S1 converts the executable file of the application into a grayscale image; S2 designs multiple different feature extraction paths, decomposing the k×k convolution operation into k×1 and 1×k double-layer convolution operations to reduce the number of model parameters, while using multi-scale convolution operations to allow the model to learn features; S3 fuses the features extracted from multiple paths to train the classification model; S4 updates the model regularly by setting an active learning iteration cycle and adopts a sample selection strategy based on uncertainty measurement. Each time the model is updated, only high-value samples are selected to update the model.
2. The Android malware classification method based on multi-path feature fusion and active learning according to claim 1 is characterized in that: Step S1 specifically includes: S11 uses the APK object method provided by the Androguard library to find the classes.dex file contained in the APK file, and then extracts the byte stream in the dex file through the get_bytes function. Each byte in the byte stream corresponds to a pixel value in the grayscale image. Finally, the byte stream in each application dex file is converted into a corresponding grayscale image. The specific steps are as follows: S101 extracts byte stream: obtains all byte data in the dex file through the get_bytes function, and then appends each byte data one by one to the bytes type stream to generate the byte stream of the dex file; S102 extracts grayscale values: each byte in the byte stream represents a pixel value, ranging from 0 to 255, and these pixel values constitute the pixel information of the original image; S103 constructs the original image: constructs the obtained pixel information into a 1×n grayscale image, where n is the length of the byte stream; S104 calculates the coordinate mapping of the target pixel in the target image: the original image is further adjusted to a target image of size m×m, and the coordinates of each pixel in the target image in the original image are first calculated to determine the position of the target pixel in the original image for subsequent interpolation calculation; S105 performs interpolation calculation: according to the position of the target pixel in the original image, a bilinear interpolation algorithm is used to perform weighted average of the grayscale values of four adjacent pixels in the original image to obtain an interpolation result corresponding to the target pixel; S106 Obtaining the grayscale value of the target pixel: taking the interpolation result obtained for the target pixel as its corresponding grayscale value; S107 repeats steps S104 to S106 for all pixels in the target image until the pixel values of the entire target image are interpolated and calculated, and finally an image of size m×m is obtained, and the grayscale image conversion is completed for downstream tasks; S12 bilinear interpolation is based on the values of four adjacent points in the matrix area. Assume that the coordinates of a pixel point P in the target image are (x, y), and the four adjacent points corresponding to this point in the original image are Q 11 , Q 12 , Q 21 , Q 22 , then we have the following formula: Among them, u represents the relative position of the target pixel (x, y) in the horizontal direction in the original image, v represents the relative position in the vertical direction, (x1, y1) and (x2, y2) are Q 11 and Q 22 The coordinates of the target point are then calculated using the bilinear interpolation algorithm, namely: f(x,y)=(1-u)(1-v)Q 11 +u(1-v)Q 21 +(1-u)vQ 12 +uvQ 22 (2) Where f represents the bilinear interpolation function, Q 11 , Q 21 , Q 12 and Q 22 Represents Q 11 , Q 12 , Q 21 and Q 22 The value of the pixel.
3. The Android malware classification method based on multi-path feature fusion and active learning according to claim 1 is characterized in that: Step S2 specifically includes: S21 introduces the context extraction module, decomposing the convolution kernel of size k×k into two smaller one-dimensional convolution kernels, k×1 and 1×k. CEM also uses a stack of multiple convolution kernels of different sizes (k=1, 3, 5, 7) to increase the width of the network layer. The number of convolution kernels in all convolution layers is 32, and RuLU is used as the activation function. Zero padding is used to ensure that the input and output sizes are consistent. The output of the last convolution layer corresponding to different convolution kernels is added as the output of CEM. Specifically, the output of CEM CEM o Calculated by the following formula: Among them, x l is the input grayscale image, w l k×1 is the weight associated with the convolution kernel of size k×1, b l is the bias term, * is the convolution operation with RuLU as the activation function; The output of the context extraction module S22 first passes through the maximum pooling layer, and then the output of the pooling layer CEM m Then it is passed to the intermediate module of the next layer. IM uses convolution kernels of size 3×3 and 5×5. The operation of 3×3 convolution kernel is decomposed into convolution operations of size 3×1 and 1×3. Similarly, the operation of 5×5 convolution kernel is also decomposed into convolution operations of size 5×1 and 1×5. Each convolution layer uses RuLU as the activation function and contains 64 convolution kernels. Their outputs are added element by element to form the output of IM. o , the specific formula is as follows: Among them, CEM m The output of the context extraction module is pooled and used as the input of IM. l 1×k are the weights associated with the convolution kernel of size 1×k; The output of the S23 intermediate module is passed to the local extraction module through the pooling layer. The LEM contains smaller 1×1 and 3×3 convolution kernels, where the 3×3 convolution operation is decomposed into a double-layer convolution operation of 3×1 and 1×3. Each convolution layer has 128 convolution kernels, and RuLU is used as the activation function. The feature maps of the convolution layers corresponding to different convolution kernels are then added element by element to finally obtain the output of the LEM. The output of the LEM o Calculated by the following formula: Among them, IM m is the output of the intermediate module and serves as the input of LEM.
4. The Android malware classification method based on multi-path feature fusion and active learning according to claim 1 is characterized in that: Step S3 specifically includes: S31 extracts features from the same image through three paths respectively, and finally fuses these three different features to obtain image features. Here, the three paths are defined as P1, P2 and P3, which correspond to the three paths from left to right of the multi-path feature extraction module. Then P1 contains three sub-modules and three pooling layers, P2 contains two sub-modules, CEM and IM, and two pooling layers, and P3 contains a CEM sub-module and a pooling layer. Let the output of CEM be CEM o , the output of IM is IM o , the output of LEM is LEM o , where P 1o , P 2o and P 3o Represent the outputs of P1, P2 and P3 respectively, MP is the maximum pooling operation function, and the pooling window size is 2×2. Then the outputs of P1, P2 and P3 are calculated by the following formula: P 10 =MP(LEM0.2)(6) P 20 =MP(IM0.2)(7) P 50 =MP(CEM0,2)(8) For the feature fusion method, splicing fusion is adopted, and the convolution layer and pooling layer are used to adjust the size and channel of the feature map. In order to achieve the final feature splicing fusion, P 2o and P 3o The dimensions are resized to match P 1o Same size, first, change P 2o and P 3o The convolution operation with the convolution kernel size of 1×1 and RuLU as the activation function is performed to adjust the number of channels to 128, and then the pooling operation is performed again. 2o , the output of the convolutional layer is subjected to a maximum pooling operation with a pooling window size of 2×2, P 3o Then a maximum pooling operation of size 4×4 is performed, and finally the three sets of feature maps of the same size are spliced together according to a specific axis. The specific formula is as follows: Fusion=Conc(P 10 ,P 20 ,P 50 )(9) Among them, Fusion is the feature obtained after splicing and fusion, and Conc is the feature splicing function; S32 performs convolution and pooling operations on the fused features again with RuLU as the activation function. For binary classification tasks, the number of neurons in the output layer is 1, Sigmoid is used as the activation function, and the final output value range is (0,1); in multi-classification tasks, the number of neurons in the output layer depends on the number of categories, and Softmax is used as the activation function. The model is trained with the cross entropy loss function, and the final output is the probability of each category, and the sum of these probabilities is 1.
5. The Android malware classification method based on multi-path feature fusion and active learning according to claim 1 is characterized in that: Step S4 specifically includes: S41 first trains the classification model with labeled samples. When new samples appear and meet the time threshold, the trained model is used to predict these samples. Then, through the designed sample selection strategy, a part of high-value samples that have the greatest impact on the model classification performance are selected for manual labeling. Finally, these samples are used to update the model. The specific steps are as follows: S401: Build an initial training set and mark it to obtain a marked sample set D; S402 trains a classification model using a labeled sample set D; S403 determines whether the time threshold is reached. If the condition is met, the newly appeared unlabeled sample set D n Make predictions using the trained model; S404 selects a portion of high-value samples through a sample selection strategy based on the prediction results to form a sample set D h ; S405 for sample set D h All samples in the dataset are manually labeled to determine their categories; S406: The labeled sample set D h Add to the initial training D to form a new sample set D s , retrain the model with this sample set; The above steps S403 to S406 are a process of updating the model once. As time goes by, when the newly appeared unlabeled samples reach the time threshold again, the next model update will be performed.