A lightweight malware visualization classification method based on multi-scale features
By converting binary files to grayscale images and using multi-scale convolutional fusion attention with SimMobileNetV2, the method addresses inefficiencies in malicious software classification, achieving high accuracy and robustness across varying datasets.
Patent Information
- Application Number
- CN202510506336.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing methods for malicious software classification face challenges due to code obfuscation and encryption, leading to inefficiencies in feature extraction and model generalization, especially in deep learning-based approaches, which are sensitive to data format and can lose information during binary file conversion to images.
A method that converts binary files to grayscale images using Markov image representation and employs a multi-scale convolutional fusion attention module (MDFA) with channel and spatial attention mechanisms to enhance feature extraction and model robustness, utilizing SimMobileNetV2 for classification.
The method significantly improves classification accuracy on malicious software datasets, achieving rates of 99.78%, 98.71%, and 97.62% on Malimg, BIG2015, and enterprise datasets, demonstrating enhanced robustness and adaptability.
Smart Images

Figure CN120032141B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of malware detection, and particularly to a lightweight malware visualization classification method based on multi-scale features. Background Art
[0002] Malware is a program that, without the user's permission, is designed to damage computer systems, steal data, or perform other harmful behaviors, including viruses, Trojan horses, worms, and ransomware, etc., and has become one of the major threats in the field of network security. Malware classification is a key task for analyzing its behavior patterns, propagation paths, and potential impacts, which helps to reduce the workload of security analysts and provides support for the detection and defense of new malware and its variants. However, with the continuous evolution of malware, technologies such as code obfuscation, encryption, and polymorphic transformation are becoming increasingly complex, and efficient and accurate classification remains a major challenge in current research.
[0003] Traditional classification methods mainly rely on dynamic analysis and static analysis techniques. Dynamic analysis monitors API calls, system behaviors, and network interactions by executing malware in a virtual environment to capture malicious behaviors during runtime. This method can bypass the interference of code obfuscation and encryption technologies, but malware may use environment awareness technologies to evade classification, and executing malicious code in a sandbox environment may bring computational overhead and security risks, such as system damage or data leakage. Static analysis extracts code features of executable files (such as binary structure, opcode sequence, string features, imported / exported functions, and control flow graphs) to reduce the risk of code leakage and quickly process a large number of samples. However, static analysis relies on manual feature extraction and is easily affected by technologies such as code obfuscation, encryption, and polymorphic transformation, which limits its adaptability.
[0004] In recent years, machine learning technologies have shown potential in malware classification. Traditional methods rely on manually constructed features (such as opcode sequences, API calls, permission requests, etc.) and use classifiers such as support vector machines (SVMs), random forests (RFs), and naive Bayes (NBs) for classification. Although these methods can improve the classification accuracy, due to the dependence of feature design on expert experience, they are difficult to adapt to the polymorphic transformation and evasion technologies of malware, and the generalization ability of the models is limited.
[0005] To solve this problem, malware classification methods based on deep learning have gradually emerged and have demonstrated excellent pattern learning capabilities in large-scale data processing. In particular, image-based deep learning methods have attracted attention due to their automated feature extraction and strong generalization. This method converts the binary data of malware into image form and uses deep neural networks for classification through visual pattern recognition. Unlike traditional methods, image processing can show the global characteristics of malware in a higher dimension, so that malware of the same category presents similar texture features in the visualized image, while different categories show obvious differences. Specifically, visualization methods usually use grayscale images or RGB images to map binary files to pixel values, and use deep neural networks for feature learning and classification. This method reduces the reliance on manual feature engineering, and can mine the potential patterns of malware and enhance classification robustness.
[0006] However, visualization-based malware classification methods still have limitations. First, image conversion may lead to information loss. When binary files are converted to images, they need to be normalized, which may lose key features and affect the detection accuracy of the model. Second, deep learning models are sensitive to the input data format. Malware files of different sizes may present different resolutions and structures after being converted to images, which affects the generalization ability of the model and leads to misclassification. In addition, data preprocessing may introduce redundant or irrelevant features. For example, non-critical features such as padding bytes and alignment structures may be learned by the model, affecting the classification effect and causing the model to fail when facing variant malware. Summary of the invention
[0007] Purpose of the invention: The purpose of the present invention is to provide a lightweight malware visualization classification method based on multi-scale features, which converts malware binary files into grayscale images through the Markov image representation method; uses a multi-scale dilated convolution fusion attention module (MDFA) to capture multi-level features of images at multiple scales; and integrates image features of different dimensions by combining the channel attention mechanism and the spatial attention mechanism, so that the model has stronger robustness and adaptability when dealing with malware variants to solve the problems existing in the background technology.
[0008] Technical solution: The lightweight malware visualization classification method based on multi-scale features described in the present invention comprises the following steps:
[0009] (1) Convert the malware binary file into a grayscale image through Markov transition matrix modeling, which includes: constructing a discrete Markov chain and counting the state transition frequency matrix of adjacent bytes in the binary byte stream; normalizing and truncating the transition frequency matrix to generate a 256×256 dimensional Markov grayscale image;
[0010] (2)Extract image features through the multi-scale dilated convolution fusion attention module MDFA, including: using parallel dilated convolution layers to capture local texture, regional structure, and global semantic features at different dilation rates; combining channel attention mechanism and spatial attention mechanism to dynamically weight key channels and spatial regions;
[0011] (3)Use the SimMobileNetV2 network for classification. The network includes: multiple cascaded SimBottleneck modules, each module embedding a SimAM attention module in the inverted residual structure, generating 3D attention weights through an energy function, strengthening key features and suppressing redundant information.
[0012] Further, in step (1), the method for generating the Markov transition matrix is: for the binary byte stream
[0013] ; according to statistical state m to n transition frequency C(m,n); where is an indicator function, taking 1 when the condition is met, otherwise taking 0; Traverse the byte stream through a sliding window to construct a transition frequency matrix and truncate the frequency values exceeding 255 to 255.
[0014] Further, in step (2), the operations of the MDFA module include: using parallel dilated convolution layers with dilation rates of 1, 3, and 5 respectively to extract short-range, medium-range, and long-range features; through a global-local feature fusion strategy, concatenating the 1×1 convolution features of the local feature enhancement path and the pooling features of the global semantic path, and performing adaptive weighting.
[0015] Further, the calculation of the channel attention mechanism includes: performing channel average pooling on the input feature map to generate a channel descriptor; generating channel weights through a dimensionality reduction convolution and an activation function, and multiplying with the original feature map channel by channel.
[0016] Further, the calculation of the spatial attention mechanism includes: performing 1×1 convolution on the input feature map to compress the channel dimension, generating a spatial weight map; multiplying with the original feature map position by position after normalizing the weights through the Sigmoid function.
[0017] Further, in step (3), the SimAM attention in the SimBottleneck module is implemented through the following steps: defining an energy function to calculate the importance of neurons, obtaining the minimum energy ; The formula is as follows:
[0018] ;
[0019] where and are the linear transformation parameters for the target neuron and other neurons in the same channel ; is the number of neurons in the channel; and represent the neuron linear transformation parameters; and are the set binary label values used to distinguish the target neuron t from other neurons ;
[0020] Add a regularization term, and the final energy function is as follows:
[0021] ;
[0022] The final minimum energy can be calculated by the following formula:
[0023] ;
[0024] where and represent the mean and variance of all neurons in the channel except t, respectively;
[0025] Refine the feature map according to ; where includes all in the channel and spatial dimensions.
[0026] Furthermore, in step (3), the structure of the SimMobileNetV2 network includes: the underlying MDFA module extracts multi-scale features; the middle layer stacks 7 SimBottleneck modules to gradually abstract the regional structure features; the upper layer outputs the classification result through global average pooling and a fully connected layer.
[0027] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory, and the steps of any of the above methods are implemented when the processor executes the program.
[0028] A computer-readable storage medium according to the present invention stores a computer program, and the steps of any of the above methods are implemented when the program is executed by a processor.
[0029] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: By converting malware binary files into image representations using Markov technology, the present invention proposes a malware visualization classification method combining multi-scale dilated convolution and attention mechanism. By converting malware into Markov images, a multi-scale dilated convolution fusion attention module (MDFA) is used to extract image features, and a SimMobileNetV2 model based on simple attention (SimAM) is constructed for classification. Experimental results show that the proposed method achieves classification accuracies of 99.78%, 98.71%, and 97.62% on the Malimg, BIG2015, and enterprise datasets respectively, significantly improving the classification performance and having strong practical application value. Brief Description of the Drawings
[0030] Figure 1 is the overall architecture diagram of the present invention;
[0031] Figure 2 is the malware visualization process diagram of the present invention;
[0032] Figure 3 is the multi-scale dilated convolution fusion attention module (MDFA) of the present invention;
[0033] Figure 4 is the SimBottleneck structure of the present invention. Detailed Embodiment
[0034] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0035] As Figure 1 shown, the embodiment of the present invention provides a lightweight malware visualization classification method based on multi-scale features, including the following steps:
[0036] Figure 1 is the overall architecture diagram of the present invention, which is divided into three parts: malware visualization, MDFA module, and SimMobileNetV2 network model; Markov transition matrix modeling: First, convert the binary file into a standardized transition frequency matrix to establish the statistical feature space of the byte sequence; Subsequently, perform dilated convolution operations with different dilation rates in parallel through the multi-scale dilated convolution fusion attention module (MDFA), and realize cross-level feature fusion by combining channel-spatial dual-path attention; Finally, the lightweight SimMobileNetV2 classification network makes a decision. This network integrates depthwise separable convolution and parameter-free attention mechanism (SimAM), maintaining the classification accuracy while reducing the computational complexity, forming a complete solution for feature representation, multi-scale analysis, and efficient classification.
[0037] As Figure 2As shown in the figure, the core theoretical basis of the malware visualization process is as follows: by quantifying the state transition law of the byte stream, while avoiding the defects of traditional image mapping, the core code features of the malware family are captured. By modeling the malware binary byte stream as a discrete Markov chain:
[0038] ;
[0039] where L is the length of the byte stream, represents the value of the i-th byte.
[0040] The transition frequency from state m to state n satisfies the first-order Markov row hypothesis, and the current byte state only depends on the previous state:
[0041] ;
[0042] where, is an indicator function, taking 1 when the condition is satisfied, otherwise taking 0; .
[0043] The Markov image representation is achieved through a three-stage processing flow:
[0044] S1 State space construction: Establish a two-dimensional transition frequency space M and establish a mapping relationship from the transition frequency to the matrix coordinates:
[0045] ;
[0046] S2 Transition statistics: Use the sliding window method to traverse the byte stream (window size w = 2, step size s = 1), and count the transition frequency matrix:
[0047] ;
[0048] S3 Truncate the high-frequency transitions exceeding 256 times to maintain the balance of the feature distribution:
[0049] By constructing the Markov transition matrix , the high-fidelity encoding of the malware statistical characteristics into a two-dimensional feature image is achieved. As Figure 2 shown, this conversion process strictly adheres to the first-order Markov hypothesis of the byte stream, completely retaining the statistical law of the original byte sequence, and fundamentally avoiding the structural distortion problem caused by dimension conversion in traditional imaging methods. For the problem of high-frequency noise interference, a truncation processing strategy is introduced to limit the amplitude of the transition frequency, and its mathematical expression is the formula:
[0050] ;
[0051] Enhancing image representation ability through triple - technology optimization: First, based on the spatial mapping function, the local continuity features of byte values are retained to ensure the relative position relationship of adjacent bytes in the image space Secondly, a sliding window traversal algorithm with linear complexity O ( L ) is adopted, which is suitable for the efficient generation of large - scale malware images; Finally, an adaptive truncation function is introduced to suppress the influence of extremely few high - frequency noises on the feature space while retaining discriminative low - frequency transfer patterns.
[0052] As Figure 3 shown, the multi - scale dilated convolution fusion attention (MDFA) module is presented. Through the collaborative work of parallel convolutional layers with different dilation rates, it realizes hierarchical capture of short - range (local texture), mid - range (regional structure), and long - range (global semantics) features. This architecture not only inherits the multi - scale representation ability of the spatial pyramid but also significantly improves the classification performance of malware through the differential combination strategy of receptive fields. Aiming at the information redundancy problem in the multi - scale feature fusion process, a global - local feature fusion strategy is designed in this paper. In the local feature enhancement path, 1x1 convolution is used for feature recombination, mapping single - channel images to a high - dimensional feature space through channel - dimension transformation, enhancing the expression ability of features while retaining spatial resolution; A cascaded structure of batch normalization layer (BatchNorm2d) and ReLU activation function is introduced to achieve feature distribution regularization and enhanced non - linear expression ability. The global semantic extraction path fuses global average pooling and multi - scale features, generating a compact feature representation with a global view while retaining key semantic information by spatially compressing each channel's feature map. The dual - path features are collaboratively optimized through channel concatenation and adaptive weighting to construct a feature expression system with hierarchical perception ability.
[0053] To solve the heterogeneity problem of channel and spatial information in malware images, the present invention proposes a multi - scale dilated convolution fusion attention module (MDFA). After completing multi - scale feature extraction, this module parallels the channel attention and spatial attention sub - modules:
[0054] The channel attention mechanism dynamically adjusts the importance weights by evaluating the feature contribution degrees of each channel, enabling the network to focus on channels containing key information. The specific process is as follows:
[0055] Channel descriptor generation: Average pooling is performed on the spatial dimension of each channel of the input feature map to compress spatial information and obtain channel - level statistical features:
[0056] ;
[0057] Feature dimensionality reduction and activation: The number of channels is reduced from C to 1 through convolution, and then the reduced features are activated by the ReLU function to learn the non-linear relationship between channels:
[0058] ;
[0059] Weight application: The activated features are restored to the original spatial dimension through bilinear interpolation ( H, W ), a channel weight matrix is generated and multiplied with the input feature map channel by channel to enhance the response intensity of important channels:
[0060] ;
[0061] The inter-attention mechanism focuses on identifying and strengthening the salient regions in the image, enhancing the influence of key spatial positions in the feature extraction process. The specific processing flow includes:
[0062] Spatial weight generation: Use 1x1 convolution on the input feature map to compress the number of channels C to 1, generating a spatial weight map:
[0063] ;
[0064] Weight normalization: Apply the Sigmoid function to the spatial weight map to constrain the weight values to the range [0, 1], achieving adaptive calibration of the spatial attention coefficients:
[0065] ;
[0066] Feature enhancement: Multiply the normalized spatial weight map with the input feature map position by position to highlight key regions such as section boundaries and information-dense areas in malware images:
[0067] ;
[0068] As Figure 3 shown, the MDFA module enhances features by concatenating global-local features in the channel dimension and jointly using channel and spatial attention mechanisms. This design combines local features with global features to ensure that the network simultaneously focuses on local byte arrangements and global structural rules. Finally, feature dimensionality reduction is achieved through 1x1 convolution, integrating information from multiple branches while reducing the computational amount, generating a more discriminative feature representation.
[0069] As Figure 4As shown, the inverted residual structure (SimBottleneck) integrating the SimAM module. To solve the problem of insufficient feature interaction after the expansion-convolution-compression process of the Bottleneck, which is the core component of MobileNetV2, the present invention constructs an improved SimBottleneck module based on SimAM (Simple Attention Module). In the expansion-convolution-compression process of the basic Bottleneck, the improved SimBottleneck structure introduces the SimAM module after compression and dimensionality reduction to calculate the 3D attention weights to strengthen the attention to key features, while expanding the depth and breadth of feature fusion and enhancing the ability to capture complex features. Finally, the input feature map and the output feature map are added through skip connection to complete feature fusion.
[0070] The SimBottleneck layer controls the change of the feature map size through different strides (1 or 2). The stride is used for downsampling to reduce the size of the feature map. The downsampling operation is only performed in the first inverted residual structure of each SimBottleneck layer, and the shortcut connection is not used at this time. Specifically, when the stride = 2, the number of input and output channels is inconsistent, and the inverted residual structure will no longer use the shortcut connection.
[0071] Without adding extra parameters, this module generates 3D attention weights for each neuron to fuse channels through the defined energy function to strengthen the channel and spatial feature responses in key regions, while suppressing redundant information and enhancing the model's ability to jointly model multi-scale heterogeneous features.
[0072] SimBottleneck defines the following energy function for each neuron:
[0073] ;
[0074] where and are the linear transformation parameters for the target neuron and other neurons in the same channel; is the number of neurons in the channel; and represent the neuron linear transformation parameters;
[0075] Minimizing the energy function is equivalent to finding the linear separability between the target neuron and all other neurons in the same channel. To simplify the problem, binary labels (i.e., 1 and -1) are introduced, and a regularization term is added. The final energy function is as follows:
[0076] ;
[0077] Among them and are the linear transformation parameters for the target neuron and other neurons in the same channel ; is the number of neurons in the channel; and represent the neuron linear transformation parameters; and are the set binary label values used to distinguish the target neuron t from other neurons ;
[0078] Relative to and there is a closed - form solution and it can be quickly obtained:
[0079] ;
[0080] ;
[0081] Among them, and respectively represent the mean and variance of all neurons except t in this channel. The final minimum energy can be calculated by the following formula:
[0082] ;
[0083] Energy The lower it is, the greater the difference between neuron t and its surrounding neurons, and the higher its importance in visual processing. Therefore, the importance of each neuron can be quantified by . After deriving the energy function and obtaining the importance of neurons, a scaling operator is used for feature refinement. The entire refinement stage of the module is as follows:
[0084] ;
[0085] Among them, contains all in the channel and spatial dimensions. The Sigmod function restricts overly large values to ensure that the relative importance of each neuron is not affected.
[0086] Such as Figure 1As shown in the figure, the SimMobileNetV2 network adopts a hierarchical processing architecture. Bottom layer feature extraction: The MDFA module is used to extract multi-scale features of grayscale images; Middle layer feature abstraction: By stacking 7 SimBottleneck layers integrated with the SimAM module, the regional structure features are gradually abstracted; High layer semantic aggregation: The classification results are output through global average pooling and fully connected layers.
Claims
1. A lightweight malware visualization classification method based on multi-scale features, characterized in that, It includes the following steps: (1) Convert the malware binary file into a grayscale image through Markov transition matrix modeling, specifically including: constructing a discrete Markov chain, and statistically analyzing the state transition frequency matrix of adjacent bytes in the binary byte stream; performing normalization and truncation processing on the transition frequency matrix to generate a 256×256-dimensional Markov grayscale image; (2) Extract image features through a multi-scale dilated convolution fusion attention module MDFA, including: adopting parallel dilated convolutional layers to capture local texture, regional structure, and global semantic features at different dilation rates; combining channel attention mechanism and spatial attention mechanism to dynamically weight key channels and spatial regions; The operations of the MDFA module include: adopting parallel dilated convolutional layers with dilation rates of 1, 3, and 5 respectively to extract short-range, mid-range, and long-range features; through a global-local feature fusion strategy, splicing the 1×1 convolutional features of the local feature enhancement path and the pooling features of the global semantic path, and performing adaptive weighting; The calculation of the channel attention mechanism includes: performing channel average pooling on the input feature map to generate a channel descriptor; generating channel weights through a dimensionality reduction convolution and an activation function, and multiplying them with the original feature map channel by channel; The calculation of the spatial attention mechanism includes: performing 1×1 convolution on the input feature map to compress the channel dimension and generate a spatial weight map; multiplying it with the original feature map position by position after normalizing the weights through the Sigmoid function; (3) Use the SimMobileNetV2 network for classification, and the network includes: multiple cascaded SimBottleneck modules, each module embeds a SimAM attention module in the inverted residual structure, generates 3D attention weights through an energy function, strengthens key features, and suppresses redundant information.
2. The lightweight malware visualization classification method based on multi-scale features according to claim 1, characterized in that, In step (1), the method for generating the Markov transition matrix is: for the binary byte stream where L is the length of the byte stream, represents the value of the i-th byte; According to count the transition frequency C(m, n) from state m to state n; where is an indicator function that takes 1 when the condition is satisfied and 0 otherwise; Traverse the byte stream through a sliding window to construct a transition frequency matrix and truncate the frequency values exceeding 255 to 255.
3. A lightweight malware visualization classification method based on multi-scale features according to claim 1, characterized in that In step (3), the SimAM attention in the SimBottleneck module is implemented through the following steps: defining an energy function to calculate the importance of neurons and obtaining the minimum energy ; The formula is as follows: ; Among them and are the linear transformation parameters for the target neuron and other neurons in the same channel ; is the number of neurons in the channel; and represent the neuron linear transformation parameters; and are the set binary label values used to distinguish the target neuron t from other neurons ; Add a regularization term, and the final energy function is as follows: ; The final minimum energy can be calculated by the following formula: ; Among them, and respectively represent the mean and variance of all neurons in the channel except t; According to refine the feature map; wherein includes all in the channel and spatial dimensions .
4. A lightweight malware visualization classification method based on multi-scale features according to claim 1, characterized in that, In step (3), the structure of the SimMobileNetV2 network includes: the underlying MDFA module extracts multi-scale features; the middle layer stacks 7 SimBottleneck modules to gradually abstract regional structure features; the upper layer outputs classification results through global average pooling and a fully connected layer.
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-4.
6. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Intelligent terminal malicious software dynamic detection method based on system call
CN109753801A
Malicious software detection and family classification method based on MAAM and CliqueNet
CN113836530A