Malicious code classification detection method and system based on improved MobileViT

Through the improved MobileViT network structure, combined with deep convolution, channel shuffling and Transformer mechanisms, the problems of large computing resource consumption and limited feature expression capabilities of malicious code detection in the existing technology are solved, and efficient and accurate detection of malicious code is achieved, which is suitable for real-time detection in resource-constrained environments.

CN119989347AActive Publication Date: 2025-05-13STATE GRID LIAONING ELECTRIC POWER CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510068951.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

Smart Images

  • Figure CN119989347A_ABST
    Figure CN119989347A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious code classification detection method and system based on improved MobileViT, and belongs to the technical field of malicious code classification detection. The method specifically comprises the following steps: collecting a binary file containing a malicious code sample, and converting the binary file into an image; preprocessing the image to obtain preprocessed malicious code image data; a malicious code classification detection model based on the improved MobileViT is constructed; the preprocessed malicious code image data are input into the model for training, and a trained malicious code classification detection model based on the improved MobileViT is obtained; and inputting actual malicious code data into the trained model, and outputting a category prediction result of the malicious code, thereby realizing malicious code classification detection based on the improved MobileViT. According to the method, high detection precision can be kept while the computing resource demand is reduced, and the method is suitable for the real-time malicious code detection demand in a resource-constrained environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of malicious code classification detection, and more specifically, relates to a malicious code classification detection method and system based on improved MobileViT. Background Art

[0002] With the rapid development of information technology, cyberspace has gradually become an important area of ​​national sovereignty, accompanied by the intensification of malicious code threats. Malicious code (or malware) is a type of code designed to destroy computer systems. Common types include viruses, Trojans, worms, and ransomware. This type of code spreads rapidly and mutates frequently, posing a serious threat to social and network security. In recent years, frequent network security incidents, such as system intrusions, information leaks, and ransomware attacks, are often manipulated by malicious code. Traditional malicious code detection methods mainly include static analysis and dynamic analysis. With the development of adversarial technology, many malicious codes have begun to use means such as shelling and code obfuscation, which affects the accuracy of static analysis. At the same time, dynamic analysis consumes a lot of resources because it needs to execute malicious code, making it difficult to meet the needs of large-scale detection. The rise of deep learning technology has gradually transitioned malicious code detection from traditional machine learning methods to automated detection methods based on deep learning. However, some current deep learning models (such as ResNet and DenseNet), although with high classification accuracy, have large computational complexity and slow detection speed in resource-constrained environments (such as embedded systems and mobile devices), and are difficult to meet the needs of real-time or large-scale deployment.

[0003] In view of the above situation, there is an urgent need for a malicious code classification and detection method designed specifically for resource-constrained scenarios.

[0004] Prior art document 1 (CN117574364B) discloses an Android malware detection method and system based on the PSEAM-MobileNet neural network. However, its shortcomings are strong dependence on feature extraction, large consumption of computing resources, insufficient robustness against adversarial attacks, and limited feature expression capabilities, and it is unable to cope with malicious code in large-scale scenarios. Summary of the invention

[0005] In order to solve the deficiencies in the prior art, the present invention provides a malicious code classification detection method and system based on improved MobileViT. The method realizes efficient and accurate detection of malicious code by innovatively integrating lightweight MobileViT network structure, deep convolution, channel shuffling and Transformer mechanism. At the same time, the present invention can maintain high detection accuracy while reducing the demand for computing resources, which is particularly suitable for the real-time malicious code detection needs in resource-constrained environments.

[0006] The present invention adopts the following technical solution.

[0007] A first aspect of the present invention provides a malicious code classification detection method based on an improved MobileViT, comprising:

[0008] Collect binary files containing malicious code samples and convert them into images;

[0009] Preprocessing the image to obtain preprocessed malicious code image data;

[0010] The MobileNetV2 model is introduced and combined with the ECA model for feature extraction. Then, deep convolution and channel shuffling are combined with the attention mechanism for deep feature extraction. A malicious code classification and detection model based on the improved MobileViT is constructed.

[0011] Input the preprocessed malicious code image data into the malicious code classification detection model training based on the improved MobileViT to obtain the trained malicious code classification detection model based on the improved MobileViT;

[0012] The actual malicious code data is input into the trained model, and the category prediction result of the malicious code is output to realize the malicious code classification detection based on the improved MobileViT.

[0013] Preferably, the preprocessing of the image specifically includes:

[0014] The image is standardized, the image size is adjusted to 224×224, the image pixel value is normalized, and one or more of the image enhancement methods of random cropping, center cropping and random flipping are used to enhance the image to obtain preprocessed malicious code image data.

[0015] Preferably, the MobileNetV2 model is introduced and combined with the ECA model for feature extraction, and then deep convolution and channel shuffling are combined with the attention mechanism for deep feature extraction to construct a malicious code classification and detection model based on the improved MobileViT. The malicious code classification and detection model based on the improved MobileViT specifically includes: a 3×3 convolution layer with a step size of 2, an MV2-ECA model constructed by combining the MobileNetV2 model with the ECA model, a MobileViT-DW model constructed by combining deep convolution and channel shuffling with the attention mechanism, a global average pooling layer and a classifier.

[0016] Preferably, based on the improved MobileViT malicious code classification detection model, the specific operation process includes:

[0017] The malicious code image data is input into the 3×3 convolutional layer with a step size of 2 in the malicious code classification and detection model based on the improved MobileViT for preliminary feature extraction and downsampling, and then input into the MobileNetV2 model for local feature extraction to obtain the first stage feature map;

[0018] The MobileNetV2 model is downsampled, and the MV2-ECA model is constructed based on the MobileNetV2 model and the ECA model. The first-stage feature map is sequentially input into the downsampled MobileNetV2 model and the MV2-ECA model for feature extraction to obtain the second-stage feature map.

[0019] The MV2-ECA model is downsampled, and the MobileViT-DW model is constructed by combining deep convolution and channel shuffling with the attention mechanism. The second-stage feature map is sequentially input into the downsampled MV2-ECA model and the MobileViT-DW model for feature extraction to obtain the third-stage feature map.

[0020] Repeat the process of inputting the third-stage feature map into the downsampled MV2-ECA model and the MobileViT-DW model for feature extraction to obtain the fourth-stage feature map.

[0021] Repeat the fourth stage feature map and input it into the downsampled MV2-ECA model and MobileViT-DW model feature extraction to obtain the fifth stage feature map;

[0022] The fifth-stage feature map is input into the classification layer for feature classification, and the classification result of the malicious code classification detection model based on the improved MobileViT is obtained.

[0023] Preferably, the MV2-ECA model is constructed based on the MobileNetV2 model and combined with the ECA model. The specific feature extraction process of the MV2-ECA model includes:

[0024] The initial feature map is used to assign attention weights to each channel in the initial feature through the ECA model, and the weighted feature map is output;

[0025] Input the weight feature map into 1×1 convolution for dimensionality reduction;

[0026] The feature map after dimensionality reduction is input into the 3×3 deep convolution layer to perform convolution processing on each channel independently to obtain a deep convolution feature map;

[0027] Then the deep convolution feature map is reduced in dimension twice through 1×1 convolution;

[0028] The feature map after the secondary dimensionality reduction is added to the initial feature map to achieve residual connection, and the feature map after feature extraction of the MV2-ECA model is obtained.

[0029] Preferably, deep convolution and channel shuffling are combined with the attention mechanism to construct the MobileViT-DW model. The specific feature extraction process of the MobileViT-DW model includes:

[0030] The input initial feature map is fed into a 3×3 deep convolutional layer for deep feature extraction;

[0031] The feature map extracted by deep convolution is reduced in dimension through 1×1 convolution, and the reduced-dimensional features are randomly rearranged through channel shuffling;

[0032] The rearranged feature map is split into non-overlapping block features and flattened through the unfold operation;

[0033] The flattened block features are input into the Transformer model, global associations between block features are established, and the associated block features are output;

[0034] The associated block features are then recombined into a two-dimensional feature map through a fold operation, and then the features are compressed through a 1×1 convolutional layer;

[0035] The compressed feature map is concatenated with the input initial feature map, and then the feature fusion is performed through a 3×3 convolutional layer to obtain the feature map after feature extraction of the MobileViT-DW model.

[0036] Preferably, the flattened block features are input into the Transformer model, a global association between the block features is established, and the associated block features are output. The specific process includes:

[0037] The flattened block features are input into the normalization layer for normalization;

[0038] The standardized block features capture the dependencies and global information between the standardized block features through a multi-head attention mechanism to obtain dependent block features;

[0039] The dependent block features are input into the multi-layer perceptron for nonlinear transformation and then input into the discard layer to output the associated block features.

[0040] Preferably, the fifth stage feature map is input into the classification layer for feature classification to obtain the classification result of the malicious code classification detection model based on the improved MobileViT. The specific feature classification process of the classification layer includes:

[0041] The fifth-stage feature map is compressed into a two-dimensional feature map through a 1×1 convolution layer in the classification layer, and then compressed into a single vector through a global average pooling layer. The probability distribution of each feature category is output through the Softmax function in the classification layer, and the classification result of the malicious code classification detection model based on the improved MobileViT is obtained.

[0042] Preferably, the preprocessed malicious code image data is input into the malicious code classification detection model training based on the improved MobileViT to obtain the trained malicious code classification detection model based on the improved MobileViT, which specifically includes:

[0043] Divide the preprocessed malicious code image data into a training set, a validation set, and a test set according to a set ratio;

[0044] The activation function of the malicious code classification detection model based on the improved MobileViT is set to the SiLU function, and the loss function is set to the cross entropy loss;

[0045] Then the training set and the validation set are input into the AdamW optimizer to train the model;

[0046] Continuously training the model, stopping the training when the number of training times of the model reaches a set number of training times for continuous training, and obtaining a trained model;

[0047] The trained model is tested using the test set to obtain a trained malicious code classification and detection model based on the improved MobileViT.

[0048] The second aspect of the present invention proposes a malicious code classification detection system based on improved MobileViT, and runs the malicious code classification detection method based on improved MobileViT, including:

[0049] Malicious code collection and conversion module: used to collect binary files containing malicious code samples and convert them into images;

[0050] Preprocessing module: used to preprocess the image to obtain preprocessed malicious code image data;

[0051] Model building module: used to build a malicious code classification detection model based on the improved MobileViT. The model introduces the MobileNetV2 model and combines it with the ECA model for feature extraction, and then combines deep convolution and channel shuffling for deep feature extraction;

[0052] Model training module: used to input the pre-processed malicious code image data into the malicious code classification detection model training based on the improved MobileViT, and obtain the trained malicious code classification detection model based on the improved MobileViT;

[0053] Malicious code classification module: used to input actual malicious code data into the trained model, output the category prediction result of the malicious code, and realize malicious code classification detection based on the improved MobileViT.

[0054] The third aspect of the present invention proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when loaded into the processor, implements the malicious code classification and detection method based on the improved MobileViT.

[0055] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the malicious code classification and detection method based on the improved MobileViT is implemented.

[0056] Compared with the prior art, the beneficial effects of the present invention include at least:

[0057] (1) Introducing the MV2-ECA model to enhance feature expression: In the initial feature extraction stage, the MV2-ECA model combines the ECA mechanism on the basis of MobileNetV2. Through cross-channel information interaction, the model’s ability to express important features is strengthened, and the model’s attention to important features is enhanced. This enables the model to maintain a high feature extraction effect while reducing computational complexity, and enables the model to focus on important features while remaining lightweight. This enhances the model’s ability to express features and robustness at different feature levels, thereby improving detection accuracy.

[0058] (2) The present invention introduces 3×3 deep convolution, channel shuffling and attention mechanism to form an improved MobileViT model, called MobileViT-DW model: 3×3 deep convolution is used to replace standard convolution to reduce the number of parameters and calculations of malicious code images, capture the deep information of malicious code images, introduce channel shuffling operation, and realize the information flow between different channels by rearranging channels, which helps to break the limitation of local receptive field, enhance the global understanding ability of the model, and improve the computational efficiency of the model. It is suitable for real-time malicious code detection tasks in resource-constrained environments, such as mobile devices and edge computing devices. Combined with Transformer mechanism, the model can focus on important parts of the input and give higher weights to certain areas, so that the model can focus on key features, reduce the interference of irrelevant details, and improve the quality of feature representation. The combination of 3×3 deep convolution, channel shuffling and attention mechanism improves the expression ability, generalization ability and computational efficiency of the model, making feature extraction more accurate.

[0059] (3) Multi-stage feature extraction and fusion: This method divides the entire network into five stages, and gradually extracts and fuses the features of malicious code images through multiple MV2, MV2-ECA and MobileViT-DW models. This staged structural design can ensure the effective expression of the model at different feature levels, improving the accuracy and generalization ability of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram of the malicious code classification detection method process based on the improved MobileViT provided in accordance with an embodiment of the present invention;

[0061] Figure 2 is a schematic diagram of a network structure of an improved MobileViT-DW model provided in accordance with an embodiment of the present invention;

[0062] Figure 3 is a schematic diagram of a network structure of an MV2-ECA model provided according to an embodiment of the present invention;

[0063] Figure 4 It is a schematic diagram of a process of inputting a processed image into a malicious code classification detection network based on an improved MobileViT model and finally outputting a category prediction result of the malicious code according to an embodiment of the present invention;

[0064] Figure 5 is a schematic diagram of a phased structure of an improved MobileViT network provided in accordance with an embodiment of the present invention;

[0065] Figure 6It is a schematic diagram of a malicious code classification detection system based on improved MobileViT provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.

[0067] like Figure 1 As shown, embodiment 1 of the present invention provides a malicious code classification detection method based on improved MobileViT, comprising the following steps:

[0068] Step 1: Collect binary files containing malicious code samples and convert them into images.

[0069] In a preferred but non-limiting embodiment of the present invention, malicious code sample data is obtained, and the binary file of the malicious code is converted into an image format (grayscale image) to meet the input requirements of the subsequent MobileViT model.

[0070] Step 2: preprocess the image to obtain preprocessed malicious code image data.

[0071] In a preferred but non-limiting embodiment of the present invention, data preprocessing includes the following operations:

[0072] Step 2.1: Standardization processing: adjust the image size to 224×224 to meet the network input requirements, and normalize the image pixel values ​​to improve network processing efficiency.

[0073] Step 2.2: Image enhancement: Use one or more data enhancement techniques such as random cropping, center cropping, and random flipping to enhance the image, increase the diversity of the data, and thus improve the robustness of the model. Among them, random cropping, center cropping, and random flipping are as follows:

[0074] Random cropping involves randomly selecting some areas from the image to increase data diversity and network robustness;

[0075] Random flipping involves randomly flipping the image horizontally with a given probability of 50% to increase the diversity of the training data;

[0076] Center cropping involves cutting out a portion of the image at the center to ensure that the main features are preserved.

[0077] Step 3: Introduce the MobileNetV2 model and combine it with the ECA model for feature extraction. Then combine deep convolution and channel shuffling with the attention mechanism for deep feature extraction, and build a malicious code classification and detection model based on the improved MobileViT.

[0078] In a preferred but non-restrictive embodiment of the present invention, in the MobileViT model, the key model is improved. Specifically, the traditional MobileViT is improved to MobileViT-DW, and the standard convolution is replaced by deep convolution (DWConv 3×3) to reduce the amount of calculation and parameters. In addition, the channel shuffle operation is introduced to promote the flow of information between channels and enhance the expressiveness of features, thereby improving the detection accuracy and generalization performance of the model.

[0079] Construct a malicious code classification and detection network model based on the improved MobileViT model, including:

[0080] Conv 3×3 / 2 is a 3×3 convolutional layer with a stride of 2, which is used for preliminary feature extraction and downsampling.

[0081] The significant difference from the existing MobileNetV2 model is that, for the feature extraction process of the MobileNetV2 model, the present invention combines the ECA model to construct the MV2-ECA model based on the MobileNetV2 model, where MV2 is the abbreviation of MobileNetV2. Figure 3 The figure shows the network structure of the MV2-ECA model. MV2-ECA combines the characteristics of the MobileNetV2 model and the ECA model. This model effectively extracts rich feature information from malicious code images while keeping the network lightweight.

[0082] Significantly different from the existing MobileViT model, the present invention combines 3×3 deep convolution (DWConv 3×3) and channel shuffle (Channel Shuffle) for the feature extraction process of the MobileViT model to capture deeper local and global semantic information. The improved MobileViT model is referred to as the MobileViT-DW model here. Here, 3×3 deep convolution is used to efficiently extract local features and significantly reduce the number of parameters and computational requirements; the channel shuffle operation rearranges the convolution channels to achieve information flow between different channels, further improving the diversity of feature expression. At the same time, the integrated Transformer mechanism captures global features through the self-attention mechanism, enhancing the model's overall understanding of malicious code images. In addition, the residual connection combines the processed features with the input features, effectively reducing information loss and improving the gradient flow in the deep network, improving the robustness of the model.

[0083] Conv 1×1 indicates that multiple 1×1 convolutional layers are used to further adjust the number of channels to ensure compact feature expression.

[0084] Global average pooling layer and classifier (Global Pool & Classifier), the global average pooling layer compresses the feature dimension into a single feature vector, and then the final classification is achieved through the linear classification layer.

[0085] like Figure 4 and Figure 5 As shown in the figure, the malicious code classification detection model based on the improved MobileViT includes the following specific operation processes:

[0086] First, the malicious code image data is subjected to preliminary feature extraction and downsampling through a 3×3 convolutional layer (Conv 3×3↓2) with a stride of 2 to obtain an initial feature map. The goal of this stage is to extract the basic features of the image and reduce the data dimension to reduce the amount of subsequent calculations. Subsequently, the feature map is input into the MobileNetV2 model to further extract local features and obtain the first-stage feature map. The MobileNetV2 model is referred to as the MV2 model.

[0087] Next, the MobileNetV2 model is downsampled, and the first-stage feature map is input into the downsampled MobileNetV2 model and the MV2-ECA model in turn for feature extraction to obtain the second-stage feature map. The ECA mechanism highlights the key information in the image by assigning attention weights between channels, allowing the model to focus on important features with low computational effort. These models generate more compact and discriminative feature maps by strengthening feature expression and further downsampling. The multiple applications of the MV2-ECA model ensure the diversity and depth of feature expression while maintaining a lightweight design.

[0088] like Figure 3 As shown in the figure, the feature extraction of the MV2-ECA model includes the following steps:

[0089] The initial feature map is used to assign attention weights to each channel in the initial feature through the ECA model, and the weighted feature map is output;

[0090] Input the weight feature map into 1×1 convolution for dimensionality reduction;

[0091] The feature map after dimensionality reduction is input into the 3×3 deep convolution layer to perform convolution processing on each channel independently to obtain a deep convolution feature map;

[0092] Then the deep convolution feature map is reduced in dimension twice through 1×1 convolution;

[0093] The feature map after the secondary dimensionality reduction is added to the initial feature map to achieve residual connection, and the feature map after feature extraction of the MV2-ECA model is obtained.

[0094] and MobileViT-DW model to obtain the third stage feature map;

[0095] Subsequently, the MV2-ECA model is downsampled, and the second-stage feature map is sequentially input into the downsampled MV2-ECA model and the MobileViT-DW model for feature extraction to obtain the third-stage feature map.

[0096] After the first round of processing by the MobileViT-DW model, the feature map is passed through another set of MV2-ECA models and MobileViT-DW models. These models work together to further integrate and deepen the feature expression, effectively capturing and fusing local and global features at each stage, thus ensuring that the final feature map has high discriminative power.

[0097] like Figure 2 As shown in the figure, the feature extraction of MobileViT-DW model includes the following steps:

[0098] The input initial feature map is input into a 3×3 deep convolutional layer with a stride of 2 for feature extraction;

[0099] The feature map extracted by deep convolution is reduced in dimension through 1×1 convolution, and the reduced-dimensional features are randomly rearranged through channel shuffling;

[0100] The rearranged feature map is split into non-overlapping block features and flattened through the unfold operation;

[0101] The flattened block features are input into the Transformer model, global associations between block features are established, and the associated block features are output;

[0102] The flattened block features are input into the Transformer model, global associations between block features are established, and associated block features are output. The specific process includes:

[0103] The flattened block features are input into the normalization layer for normalization;

[0104] The standardized block features capture the dependencies and global information between the standardized block features through a multi-head attention mechanism to obtain dependent block features;

[0105] The dependent block features are input into the multi-layer perceptron for nonlinear transformation and then input into the discard layer to output the associated block features.

[0106] The associated block features are then recombined into a two-dimensional feature map through a fold operation, and then the features are compressed through a 1×1 convolutional layer;

[0107] The compressed feature map is concatenated with the input initial feature map, and then the feature fusion is performed through a 3×3 convolutional layer to obtain the feature map after feature extraction of the MobileViT-DW model.

[0108] Repeat the process of inputting the third-stage feature map into the downsampled MV2-ECA model and the MobileViT-DW model for feature extraction to obtain the fourth-stage feature map.

[0109] Repeat the fourth stage feature map and input it into the downsampled MV2-ECA model and MobileViT-DW model feature extraction to obtain the fifth stage feature map;

[0110] Input the classification layer for feature classification and obtain the classification results of the malicious code classification detection model based on the improved MobileViT.

[0111] Finally, the fifth-stage feature map after all models are processed is input to the classification layer for final feature classification, and the classification results of the malicious code classification detection model based on the improved MobileViT are obtained. The classification layer first uses a 1×1 convolution layer to adjust the number of feature channels of the fifth-stage feature map, compresses the fifth-stage feature map to a two-dimensional feature map, and ensures the compact expression of the features; then the two-dimensional feature map is compressed into a single vector through the global average pooling layer, which retains the global feature information while reducing the data dimension. Subsequently, the feature vector is sent to the linear classification layer for processing, and the probability distribution of each feature category is output through the Softmax function, and the classification results of the malicious code classification detection model based on the improved MobileViT are obtained, thereby achieving accurate classification and detection of malicious code types.

[0112] Step 4: Input the training of the malicious code classification detection model based on the improved MobileViT to obtain the trained malicious code classification detection model based on the improved MobileViT.

[0113] Step 4: input the preprocessed malicious code image data into the training of the malicious code classification detection model based on the improved MobileViT model to obtain a trained malicious code classification detection model based on the improved MobileViT.

[0114] Further preferably, the divided malicious code data set is input into the improved MobileViT model for training. The present invention adopts the AdamW optimizer to optimize the training efficiency while reducing the risk of overfitting. The loss function adopts the cross entropy loss, which is suitable for processing probability prediction in classification tasks. The activation function selects the SiLU function for nonlinear feature conversion. The training strategy of the present invention is to continue training until the model performance is no longer improved, at which time the weight file of the network is saved for subsequent use.

[0115] The preprocessed malicious code image data is input into the malicious code classification detection model training based on the improved MobileViT to obtain the trained malicious code classification detection model based on the improved MobileViT, which specifically includes:

[0116] Divide the preprocessed malicious code image data into a training set, a validation set, and a test set according to a set ratio;

[0117] Image data division: In the present invention, image data division will be performed based on the category of malicious code (such as Adialer.C, Agent.FYI, Allaple.A, etc.) to ensure that samples of each category can be used independently and evenly for training, verification, and testing. The specific division method is as follows:

[0118] First, all image samples are preliminarily divided according to the malicious code family (category). Samples of each malicious code category will be processed separately to ensure that each category can be fully represented in subsequent training and testing. For each category, the image data will be divided and processed independently to avoid mixing of data between different categories. Training set, validation set and test set division: Within each category, the data will be divided into training set, validation set and test set according to a certain proportion. The training set accounts for 80%, the validation set accounts for 10%, and the test set accounts for 10%. In this way, it is ensured that the image data of each category can be reasonably used for model training, parameter adjustment and final test evaluation.

[0119] The final number of image samples (quantity) of the family (category) to which each image belongs is shown in Table 1:

[0120] Table 1 Number of samples in each category

[0121]

[0122]

[0123] The activation function of the malicious code classification detection model based on the improved MobileViT is set to the SiLU function, and the loss function is set to the cross entropy loss;

[0124] Then the training set and the validation set are input into the AdamW optimizer to train the model;

[0125] Continuously training the model, stopping the training when the number of training times of the model reaches a set number of training times for continuous training, and obtaining a trained model;

[0126] The trained model is tested using the test set to obtain a trained malicious code classification and detection model based on the improved MobileViT.

[0127] The MobileViT model is tested and validated using a test set consisting of original malicious code images. The test set can effectively evaluate the detection capability of the network on real data. In addition, this embodiment also selects a number of currently popular malicious code detection networks for comparison, verifying the superiority of the method of the present invention.

[0128] Precision, recall, and the harmonic mean of precision and recall are expressed as follows:

[0129]

[0130] Among them, TP, TN, FP, and FN respectively represent the number of correctly predicted positive samples, the number of correctly predicted negative samples, the number of incorrectly predicted positive samples, and the number of incorrectly predicted negative samples. This embodiment selects the network in the following table to compare using Precision, Recall, F1 (F2 score is the harmonic mean of precision and recall) and model parameter quantity to verify the superiority of the method of the present invention. The method of the present invention surpasses the same type of network in all indicators, among which Precision is as high as 96.24, Recall is as high as 97.02, F1 is as high as 96.63, and the model parameter quantity is 2.95.

[0131] Table 2 Performance comparison of different models

[0132]

[0133]

[0134] Step 5: Input the actual malicious code data into the trained model, output the category prediction result of the malicious code, and implement malicious code classification detection based on the improved MobileViT.

[0135] Compared with the prior art, the beneficial effects of the present invention include at least:

[0136] (1) Introducing the MV2-ECA model to enhance feature expression: In the initial feature extraction stage, the MV2-ECA model combines the ECA mechanism on the basis of MobileNetV2. Through cross-channel information interaction, the model’s ability to express important features is strengthened, and the model’s attention to important features is enhanced. This enables the model to maintain a high feature extraction effect while reducing computational complexity, and enables the model to focus on important features while remaining lightweight. This enhances the model’s ability to express features and robustness at different feature levels, thereby improving detection accuracy.

[0137] (2) The present invention introduces 3×3 deep convolution, channel shuffling and attention mechanism to form an improved MobileViT model, called MobileViT-DW model: 3×3 deep convolution is used to replace standard convolution to reduce the number of parameters and calculations of malicious code images, capture the deep information of malicious code images, introduce channel shuffling operation, and realize the information flow between different channels by rearranging channels, which helps to break the limitation of local receptive field, enhance the global understanding ability of the model, and improve the computational efficiency of the model. It is suitable for real-time malicious code detection tasks in resource-constrained environments, such as mobile devices and edge computing devices. Combined with Transformer mechanism, the model can focus on important parts of the input and give higher weights to certain areas, so that the model can focus on key features, reduce the interference of irrelevant details, and improve the quality of feature representation. The combination of 3×3 deep convolution, channel shuffling and attention mechanism improves the expression ability, generalization ability and computational efficiency of the model, making feature extraction more accurate.

[0138] (3) Multi-stage feature extraction and fusion: This method divides the entire network into five stages, and gradually extracts and fuses the features of malicious code images through multiple MV2, MV2-ECA and MobileViT-DW models. This staged structural design can ensure the effective expression of the model at different feature levels, improving the accuracy and generalization ability of detection.

[0139] like Figure 6 As shown, embodiment 2 of the present invention provides a malicious code classification detection system based on improved MobileViT, and runs a malicious code classification detection method based on improved MobileViT described in embodiment 1, including:

[0140] Malicious code collection and conversion module: used to collect binary files containing malicious code samples and convert them into images;

[0141] Preprocessing module: used to preprocess the image to obtain preprocessed malicious code image data;

[0142] Model building module: used to introduce the MobileNetV2 model and combine it with the ECA model for feature extraction, and then combine deep convolution and channel shuffling with the attention mechanism for deep feature extraction, and build a malicious code classification detection model based on the improved MobileViT;

[0143] Model training module: used to input the pre-processed malicious code image data into the malicious code classification detection model training based on the improved MobileViT, and obtain the trained malicious code classification detection model based on the improved MobileViT;

[0144] Malicious code classification module: used to input actual malicious code data into the trained model, output the category prediction result of the malicious code, and realize malicious code classification detection based on the improved MobileViT.

[0145] Embodiment 3 of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, a malicious code classification and detection method based on improved MobileViT as described in Embodiment 1 is implemented.

[0146] Embodiment 4 of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the malicious code classification and detection method based on the improved MobileViT according to embodiment 1 is implemented.

[0147] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A malicious code classification detection method based on improved MobileViT, characterized in that: include: Collect binary files containing malicious code samples and convert them into images; Preprocessing the image to obtain preprocessed malicious code image data; The MobileNetV2 model is introduced and combined with the ECA model for feature extraction. Then, deep convolution and channel shuffling are combined with the attention mechanism for deep feature extraction. A malicious code classification and detection model based on the improved MobileViT is constructed. Input the preprocessed malicious code image data into the malicious code classification detection model training based on the improved MobileViT to obtain the trained malicious code classification detection model based on the improved MobileViT; The actual malicious code data is input into the trained model, and the category prediction result of the malicious code is output to realize the malicious code classification detection based on the improved MobileViT.

2. The malicious code classification detection method based on improved MobileViT according to claim 1, characterized in that: The preprocessing of the image specifically includes: The image is standardized, the image size is adjusted to 224×224, the image pixel value is normalized, and one or more of the image enhancement methods of random cropping, center cropping and random flipping are used to enhance the image to obtain preprocessed malicious code image data.

3. The malicious code classification detection method based on improved MobileViT according to claim 1 is characterized in that: The MobileNetV2 model is introduced and combined with the ECA model for feature extraction. Then deep convolution and channel shuffling are combined with the attention mechanism for deep feature extraction. A malicious code classification and detection model based on the improved MobileViT is constructed. The malicious code classification and detection model based on the improved MobileViT specifically includes: a 3×3 convolution layer with a step size of 2, an MV2-ECA model constructed by combining the MobileNetV2 model with the ECA model, a MobileViT-DW model constructed by combining deep convolution and channel shuffling with the attention mechanism, a global average pooling layer and a classifier.

4. The malicious code classification detection method based on improved MobileViT according to claim 3 is characterized in that: Based on the improved MobileViT malicious code classification detection model, the specific operation process includes: The malicious code image data is input into the 3×3 convolutional layer with a step size of 2 in the malicious code classification and detection model based on the improved MobileViT for preliminary feature extraction and downsampling, and then input into the MobileNetV2 model for local feature extraction to obtain the first stage feature map; The MobileNetV2 model is downsampled, and the MV2-ECA model is constructed based on the MobileNetV2 model and the ECA model. The first-stage feature map is sequentially input into the downsampled MobileNetV2 model and the MV2-ECA model for feature extraction to obtain the second-stage feature map. The MV2-ECA model is downsampled, and the MobileViT-DW model is constructed by combining deep convolution and channel shuffling with the attention mechanism. The second-stage feature map is sequentially input into the downsampled MV2-ECA model and the MobileViT-DW model for feature extraction to obtain the third-stage feature map. Repeat the process of inputting the third-stage feature map into the downsampled MV2-ECA model and the MobileViT-DW model for feature extraction to obtain the fourth-stage feature map. Repeat the fourth stage feature map and input it into the downsampled MV2-ECA model and MobileViT-DW model feature extraction to obtain the fifth stage feature map; The fifth-stage feature map is input into the classification layer for feature classification, and the classification result of the malicious code classification detection model based on the improved MobileViT is obtained.

5. According to the malicious code classification detection method based on improved MobileViT according to claim 4, its characteristics are: exist The MV2-ECA model is constructed based on the MobileNetV2 model and the ECA model. The specific feature extraction process of the MV2-ECA model includes: The initial feature map is used to assign attention weights to each channel in the initial feature through the ECA model, and the weighted feature map is output; Input the weight feature map into 1×1 convolution for dimensionality reduction; The feature map after dimensionality reduction is input into the 3×3 deep convolution layer to perform convolution processing on each channel independently to obtain a deep convolution feature map; Then the deep convolution feature map is reduced in dimension twice through 1×1 convolution; The feature map after the secondary dimensionality reduction is added to the initial feature map to achieve residual connection, and the feature map after feature extraction of the MV2-ECA model is obtained.

6. The malicious code classification detection method based on improved MobileViT according to claim 4 is characterized in that: The MobileViT-DW model is constructed by combining deep convolution and channel shuffling with the attention mechanism. The specific feature extraction process of the MobileViT-DW model includes: The input initial feature map is fed into a 3×3 deep convolutional layer for deep feature extraction; The feature map extracted by deep convolution is reduced in dimension through 1×1 convolution, and the reduced-dimensional features are randomly rearranged through channel shuffling; The rearranged feature map is split into non-overlapping block features and flattened through the unfold operation; The flattened block features are input into the Transformer model, global associations between block features are established, and the associated block features are output; The associated block features are then recombined into a two-dimensional feature map through a fold operation, and then the features are compressed through a 1×1 convolutional layer; The compressed feature map is concatenated with the input initial feature map, and then the feature fusion is performed through a 3×3 convolutional layer to obtain the feature map after feature extraction of the MobileViT-DW model.

7. The malicious code classification detection method based on improved MobileViT according to claim 6 is characterized in that: The flattened block features are input into the Transformer model, global associations between block features are established, and associated block features are output. The specific process includes: The flattened block features are input into the normalization layer for normalization; The standardized block features capture the dependencies and global information between the standardized block features through a multi-head attention mechanism to obtain dependent block features; The dependent block features are input into the multi-layer perceptron for nonlinear transformation and then input into the discard layer to output the associated block features.

8. The malicious code classification detection method based on improved MobileViT according to claim 4 is characterized in that: The fifth stage feature map is input into the classification layer for feature classification, and the classification result of the malicious code classification detection model based on the improved MobileViT is obtained. The specific feature classification process of the classification layer includes: The fifth-stage feature map is compressed into a two-dimensional feature map through a 1×1 convolution layer in the classification layer, and then compressed into a single vector through a global average pooling layer. The probability distribution of each feature category is output through the Softmax function in the classification layer, and the classification result of the malicious code classification detection model based on the improved MobileViT is obtained.

9. The malicious code classification detection method based on improved MobileViT according to claim 1, characterized in that: The preprocessed malicious code image data is input into the malicious code classification detection model training based on the improved MobileViT to obtain the trained malicious code classification detection model based on the improved MobileViT, which specifically includes: Divide the preprocessed malicious code image data into a training set, a validation set, and a test set according to a set ratio; The activation function of the malicious code classification detection model based on the improved MobileViT is set to the SiLU function, and the loss function is set to the cross entropy loss; Then the training set and the validation set are input into the AdamW optimizer to train the model; Continuously training the model, stopping the training when the number of training times of the model reaches a set number of training times for continuous training, and obtaining a trained model; The trained model is tested using the test set to obtain a trained malicious code classification and detection model based on the improved MobileViT.

10. A malicious code classification detection system based on improved MobileViT, running a malicious code classification detection method based on improved MobileViT according to any one of claims 1 to 9, characterized in that: include: Malicious code collection and conversion module: used to collect binary files containing malicious code samples and convert them into images; Preprocessing module: used to preprocess the image to obtain preprocessed malicious code image data; Model building module: used to introduce the MobileNetV2 model and combine it with the ECA model for feature extraction, and then combine deep convolution and channel shuffling with the attention mechanism for deep feature extraction, and build a malicious code classification detection model based on the improved MobileViT; Model training module: used to input the pre-processed malicious code image data into the malicious code classification detection model training based on the improved MobileViT, and obtain the trained malicious code classification detection model based on the improved MobileViT; Malicious code classification module: used to input actual malicious code data into the trained model, output the category prediction result of the malicious code, and realize malicious code classification detection based on the improved MobileViT.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into the processor, the malicious code classification detection method based on improved MobileViT as described in any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a malicious code classification detection method based on improved MobileViT as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • An Android malware detection method and system based on PSEAM-MobileNet neural network

    CN117574364B

  • Malicious code family detection method and device, electronic equipment and storage medium

    CN116992446A

  • Malicious code classification method and system and electronic equipment

    CN117407875A

  • Malicious code classification method and device, storage medium and electronic equipment

    CN117521067A