Lightweight malicious software visual classification method based on multi-scale features
By converting malware binary files into Markov images and using multi-scale hollow convolution convolution to extract features, combining channel and spatial attention mechanisms, the problem of malware classification in the existing technology is difficult to adapt to polymorphic deformation and evasion technologies, and malware classification with high accuracy and strong generalization capabilities is achieved.
Patent Information
- Application Number
- CN202510506336.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing malware classification methods are difficult to maintain efficient and accurate classification performance when facing the polymorphic deformation and evasion technology of malware. Traditional methods rely on manual feature extraction and deep learning models to be sensitive to input data formats, affecting the generalization ability of the model.
采用基于多尺度特征的轻量化恶意软件可视化分类方法,通过马尔可夫图像表征将恶意软件二进制文件转换为灰度图像,并利用多尺度空洞卷积融合注意力模块(MDFA)提取图像特征,结合通道注意力机制和空间注意力机制,整合不同维度的图像特征,最终通过SimMobileNetV2网络进行分类。
It significantly improves the accuracy and robustness of malware classification, can better adapt to the polymorphic deformation and evasion technology of malware, and improves the generalization ability and practical application value of the model.
Smart Images

Figure CN120032141A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of malware detection, and in particular to a lightweight malware visualization classification method based on multi-scale features. Background Art
[0002] Malware is a program that is designed to damage computer systems, steal data, or perform other harmful behaviors without user permission. It includes viruses, Trojan horses, worms, and ransomware, and has become one of the major threats in the field of network security. Malware classification is a key task to analyze its behavior patterns, propagation paths, and potential impacts. It helps to reduce the workload of security analysts and provides support for the detection and defense of new malware and its variants. However, as malware continues to evolve, techniques such as code obfuscation, encryption, and polymorphism are becoming increasingly complex, and efficient and accurate classification remains a major challenge in current research.
[0003] Traditional classification methods mainly rely on dynamic analysis and static analysis techniques. Dynamic analysis captures malicious behaviors at runtime by executing malware in a virtual environment, monitoring its API calls, system behaviors, and network interactions. This method can bypass the interference of code obfuscation and encryption technologies, but malware may use environmental perception technology to evade classification, and dynamic analysis may bring computational overhead and security risks such as system damage or data leakage when executing malicious code in a sandbox environment. Static analysis reduces the risk of code leakage and quickly processes a large number of samples by extracting code features of executable files (such as binary structure, opcode sequence, string features, import / export functions, and control flow graphs). However, static analysis relies on manual feature extraction and is easily affected by technologies such as code obfuscation, encryption, and polymorphism, which limits its adaptability.
[0004] In recent years, machine learning technology has shown potential in malware classification. Traditional methods rely on manually constructed features (such as opcode sequences, API calls, permission requests, etc.) and use classifiers such as support vector machines (SVM), random forests (RF), and naive Bayes (NB) for classification. Although these methods can improve classification accuracy, they are difficult to adapt to the polymorphic deformation and evasion techniques of malware because feature design relies on expert experience, and the model generalization ability is limited.
[0005] To solve this problem, malware classification methods based on deep learning have gradually emerged and have demonstrated excellent pattern learning capabilities in large-scale data processing. In particular, image-based deep learning methods have attracted attention due to their automated feature extraction and strong generalization. This method converts the binary data of malware into image form and uses deep neural networks for classification through visual pattern recognition. Unlike traditional methods, image processing can show the global characteristics of malware in a higher dimension, so that malware of the same category presents similar texture features in the visualized image, while different categories show obvious differences. Specifically, visualization methods usually use grayscale images or RGB images to map binary files to pixel values, and use deep neural networks for feature learning and classification. This method reduces the reliance on manual feature engineering, and can mine the potential patterns of malware and enhance classification robustness.
[0006] However, visualization-based malware classification methods still have limitations. First, image conversion may lead to information loss. When binary files are converted to images, they need to be normalized, which may lose key features and affect the detection accuracy of the model. Second, deep learning models are sensitive to the input data format. Malware files of different sizes may present different resolutions and structures after being converted to images, which affects the generalization ability of the model and leads to misclassification. In addition, data preprocessing may introduce redundant or irrelevant features. For example, non-critical features such as padding bytes and alignment structures may be learned by the model, affecting the classification effect and causing the model to fail when facing variant malware. Summary of the invention
[0007] Purpose of the invention: The purpose of the present invention is to provide a lightweight malware visualization classification method based on multi-scale features, which converts malware binary files into grayscale images through the Markov image representation method; uses a multi-scale dilated convolution fusion attention module (MDFA) to capture multi-level features of images at multiple scales; and integrates image features of different dimensions by combining the channel attention mechanism and the spatial attention mechanism, so that the model has stronger robustness and adaptability when dealing with malware variants to solve the problems existing in the background technology.
[0008] Technical solution: The lightweight malware visualization classification method based on multi-scale features described in the present invention comprises the following steps: (1) Convert the malware binary file into a grayscale image through Markov transition matrix modeling, which includes: constructing a discrete Markov chain and counting the state transition frequency matrix of adjacent bytes in the binary byte stream; normalizing and truncating the transition frequency matrix to generate a 256×256 dimensional Markov grayscale image; (2) Extract image features through the multi-scale dilated convolution fusion attention module MDFA, including: using parallel dilated convolution layers to capture local texture, regional structure and global semantic features with different dilation rates; combining channel attention mechanism and spatial attention mechanism to dynamically weight key channels and spatial regions; (3) Classification is performed using the SimMobileNetV2 network, which includes: multiple serially connected SimBottleneck modules, each of which embeds a SimAM attention module in an inverted residual structure, generates 3D attention weights through an energy function, strengthens key features and suppresses redundant information.
[0009] Furthermore, in step (1), the Markov transition matrix is generated by: ;according to Count the transition frequency C(m,n) from state m to state n; where, is an indicator function, which takes 1 when the condition is met and 0 otherwise; Traverse the byte stream through the sliding window and build the transfer frequency matrix Frequency values exceeding 255 are truncated to 255.
[0010] Furthermore, in step (2), the operation of the MDFA module includes: using parallel atrous convolutional layers with atrous rates of 1, 3, and 5 to extract short-range, medium-range, and long-range features, respectively; and concatenating the 1×1 convolutional features of the local feature enhancement path with the pooled features of the global semantic path through a global-local feature fusion strategy, and performing adaptive weighting.
[0011] Furthermore, the calculation of the channel attention mechanism includes: performing channel average pooling on the input feature map to generate a channel descriptor; generating channel weights through dimensionality reduction convolution and activation function, and multiplying them channel by channel with the original feature map.
[0012] Furthermore, the calculation of the spatial attention mechanism includes: performing a 1×1 convolution on the input feature map to compress the channel dimension and generate a spatial weight map; normalizing the weights through a Sigmoid function and then multiplying them position by position with the original feature map.
[0013] Furthermore, in step (3), the SimAM attention in the SimBottleneck module is implemented by the following steps: defining an energy function to calculate the importance of neurons and obtaining the minimum energy ; The formula is as follows: ; in and The target neuron and other neurons in the same channel The linear transformation parameters of is the number of neurons in the channel; and represents the neuron linear transformation parameters; and is the set binary label value, used to distinguish the target neuron t from other neurons ; Adding the regularization term, the final energy function is as follows: ; The final minimum energy can be calculated by the following formula: ; in, and Respectively represent the mean and variance of all neurons in the channel except t; according to The feature map is refined; wherein, Contains all channels and spatial dimensions .
[0014] Furthermore, in step (3), the structure of the SimMobileNetV2 network includes: the bottom-level MDFA module extracts multi-scale features; the middle-level stacks 7 SimBottleneck modules to gradually abstract regional structural features; and the high-level outputs the classification results through global average pooling and a fully connected layer.
[0015] An electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory, and the processor implements the steps of any one of the methods when executing the program.
[0016] A computer-readable storage medium according to the present invention stores a computer program, and when the program is executed by a processor, the steps of any one of the methods are implemented.
[0017] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: The present invention converts malware binary files into image representations using Markov technology, and proposes a malware visualization classification method that combines multi-scale dilated convolution with attention mechanism. By converting malware into Markov images, multi-scale dilated convolution fusion attention module (MDFA) is used to extract image features, and a SimMobileNetV2 model based on simple attention (SimAM) is constructed for classification. Experimental results show that the classification accuracy of the proposed method on Malimg, BIG2015 and enterprise datasets reached 99.78%, 98.71% and 97.62% respectively, which significantly improved the classification performance and has strong practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is the overall architecture diagram of the present invention; Figure 2 A diagram of the malware visualization process of the present invention; Figure 3 The multi-scale dilated convolution fusion attention module (MDFA) of the present invention; Figure 4 It is the SimBottleneck structure of the present invention. DETAILED DESCRIPTION
[0019] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0020] like Figure 1 As shown, an embodiment of the present invention provides a lightweight malware visualization classification method based on multi-scale features, comprising the following steps: Figure 1 This is the overall architecture diagram of the present invention, which is divided into three parts: malware visualization, MDFA module, and SimMobileNetV2 network model; Markov transfer matrix modeling: first, the binary file is converted into a standardized transfer frequency matrix to establish a statistical feature space of the byte sequence; then, the multi-scale dilated convolution fusion attention module (MDFA) is used to perform dilated convolution operations with different expansion rates in parallel, and the channel-space dual-path attention is combined to achieve cross-level feature fusion; finally, the decision is made by the lightweight SimMobileNetV2 classification network, which integrates deep separable convolution and parameter-free attention mechanism (SimAM), which reduces the computational complexity while maintaining classification accuracy, forming a complete solution for feature representation, multi-scale analysis, and efficient classification.
[0021] like Figure 2 As shown in the figure, the core theoretical basis of the malware visualization process is to capture the core code features of the malware family while avoiding the defects of traditional image mapping by quantifying the state transition rules of the byte stream. By modeling the malware binary byte stream as a discrete Markov chain: ; Where L is the length of the byte stream, Represents the i-th byte value.
[0022] The transition frequency from state m to state n satisfies the first-order Markov row assumption, and the current byte state only depends on the previous state: ; in, is an indicator function, which takes 1 when the condition is met and 0 otherwise; .
[0023] Markov image characterization is achieved through a three-stage processing pipeline: S1 state space construction: establish a two-dimensional transition frequency space M, and establish the mapping relationship from transition frequency to matrix coordinates: ; S2 transfer statistics: Use sliding window method to traverse the byte stream (window size w = 2, step length s = 1), statistics Transition frequency matrix: ; S3 truncates high-frequency transfers exceeding 256 times to maintain a balanced feature distribution: By constructing the Markov transition matrix , achieving high-fidelity encoding of malware statistical characteristics into two-dimensional feature images. Figure 2 As shown in the figure, the conversion process strictly adheres to the first-order Markov hypothesis of the byte stream, completely retains the statistical laws of the original byte sequence, and fundamentally avoids the structural distortion problem caused by dimensional conversion in traditional visualization methods. In order to solve the problem of high-frequency noise interference, a truncation processing strategy is introduced to limit the transfer frequency, and its mathematical expression is the formula: ; The image representation capability is improved through three technical optimizations: First, the local continuity characteristics of byte values are preserved based on the spatial mapping function to ensure that adjacent bytes are consistent in the image space. Secondly, the linear complexity is used O ( L )'s sliding window traversal algorithm is suitable for the efficient generation of large-scale malware images; finally, the introduction of an adaptive truncation function can suppress the impact of very few high-frequency noises on the feature space while retaining the discriminative low-frequency transfer patterns.
[0024] like Figure 3As shown in the figure, the multi-scale dilated convolution fusion attention (MDFA) module is demonstrated. Through the collaborative work of parallel convolution layers with differentiated dilation rates, the hierarchical capture of short-range (local texture), medium-range (regional structure) and long-range (global semantics) features is achieved. This architecture not only inherits the multi-scale representation capability of the spatial pyramid, but also significantly improves the classification performance of malware through the differentiated combination strategy of the receptive field. In view of the information redundancy problem in the multi-scale feature fusion process, this paper designs a global-local feature fusion strategy. In the local feature enhancement path, 1x1 convolution is used for feature reorganization, and the single-channel image is mapped to the high-dimensional feature space through channel dimension transformation, which improves the feature expression ability while retaining the spatial resolution; the cascade structure of the batch normalization layer (BatchNorm2d) and the ReLU activation function is introduced to realize feature distribution regularization and nonlinear expression ability enhancement. The global semantic extraction path integrates global average pooling and multi-scale features, and generates a compact feature representation with a global perspective while retaining key semantic information by spatially compressing the feature maps of each channel. The dual-path features are collaboratively optimized through channel splicing and adaptive weighting to construct a feature expression system with hierarchical perception capabilities.
[0025] In order to solve the heterogeneity problem of channel and spatial information in malware images, this paper proposes a multi-scale dilated convolution fusion attention module (MDFA). After completing multi-scale feature extraction, this module connects the channel attention (ChannelAttention) and spatial attention (SpatialAttention) submodules in parallel: The channel attention mechanism dynamically adjusts the importance weight by evaluating the feature contribution of each channel, so that the network focuses on the channels containing key information. The specific process is: Channel descriptor generation: Perform average pooling on the spatial dimension of each channel of the input feature map to compress the spatial information and obtain channel-level statistical features: ; Feature dimensionality reduction and activation: The number of channels is reduced from C to 1 through convolution, and then the reduced features are activated by the ReLU function to learn the nonlinear relationship between channels: ; Weight application: restore the activated features to the original spatial dimensions through bilinear interpolation ( H,W ), generate a channel weight matrix and multiply it with the input feature map channel by channel to enhance the response strength of important channels: ; The inter-attention mechanism focuses on identifying and strengthening the salient areas in the image, increasing the influence of key spatial locations in the feature extraction process. The specific processing flow includes: Spatial weight generation: Use 1x1 convolution on the input feature map to compress the number of channels C to 1 and generate a spatial weight map: ; Weight normalization: Apply the Sigmoid function to the spatial weight map to constrain the weight value to the [0, 1] interval to achieve adaptive calibration of the spatial attention coefficient: ; Feature enhancement: The normalized spatial weight map is multiplied with the input feature map position by position to highlight key areas such as section boundaries and information-dense areas in the malware image: ; like Figure 3 As shown in the figure, the MDFA module combines global-local features in the channel dimension and combines channel and spatial attention mechanisms to enhance features. This design combines local features with global features to ensure that the network pays attention to both local byte arrangement and global structural rules. Finally, feature dimensionality reduction is achieved through 1x1 convolution, which integrates multi-branch information while reducing the amount of computation and generates more discriminative feature representations.
[0026] like Figure 4 As shown, the inverted residual structure (SimBottleneck) of the integrated SimAM module. In order to solve the problem of insufficient feature interaction after the expansion-convolution-compression process of the inverted residual module (Bottleneck), which is the core component of MobileNetV2, the present invention constructs an improved SimBottleneck module based on SimAM (Simple Attention Module). In the expansion-convolution-compression process of the basic Bottleneck, the improved SimBottleneck structure introduces the SimAM module after compression and dimensionality reduction to calculate the three-dimensional attention weights to strengthen the focus on key features, while expanding the depth and breadth of feature fusion, and improving the ability to capture complex features. Finally, the input feature map and the output feature map are added through jump connections to complete the feature fusion.
[0027] The SimBottleneck layer controls the change in feature map size by differentiating the stride (1 or 2). The stride is used for downsampling, thereby reducing the size of the feature map. The downsampling operation is only performed in the first inverted residual structure of each SimBottleneck layer, and the shortcut connection is not used at this time. Specifically, when the number of input and output channels is inconsistent when the stride=2, the inverted residual structure will no longer use the shortcut connection.
[0028] This module generates 3D attention weights of fused channels for each neuron through a defined energy function without adding additional parameters to strengthen the channel and spatial feature responses in key areas, while suppressing redundant information and improving the model's ability to jointly model multi-scale heterogeneous features.
[0029] SimBottleneck defines the following energy function for each neuron: ; in and The target neuron and other neurons in the same channel The linear transformation parameters of is the number of neurons in the channel; and represents the neuron linear transformation parameters; Minimizing the energy function is equivalent to finding the linear separability between the target neuron and all other neurons in the same channel. To simplify the problem, binary labels (i.e., 1 and -1) are introduced and regularization terms are added. The final energy function is as follows: ; in and The target neuron and other neurons in the same channel The linear transformation parameters of is the number of neurons in the channel; and represents the neuron linear transformation parameters; and is the set binary label value, used to distinguish the target neuron t from other neurons ; Relative to and There is a closed-form solution that can be found quickly: ; ; in, and Respectively represent the mean and variance of all neurons in the channel except t. The final minimum energy can be calculated by the following formula: ; energy The lower the value, the greater the difference between the neuron t and the surrounding neurons, and the higher its importance in visual processing. Therefore, the importance of each neuron can be calculated by to quantify. After deriving the energy function and obtaining the importance of neurons, feature refinement is performed using a scaling operator. The entire refinement phase of the module is as follows: ; in, Contains all channels and spatial dimensions The Sigmod function limits excessively large values to ensure that the relative importance of each neuron is not affected.
[0030] like Figure 1 As shown in the figure, the SimMobileNetV2 network adopts a hierarchical processing architecture. Bottom-level feature extraction: use the MDFA module to extract multi-scale features of grayscale images; middle-level feature abstraction: by stacking 7 SimBottleneck layers that integrate the SimAM module, the regional structural features are gradually abstracted; high-level semantic aggregation: output classification results through global average pooling and fully connected layers.
Claims
1. A lightweight malware visualization classification method based on multi-scale features, characterized in that: The following steps are involved: (1) Convert the malware binary file into a grayscale image through Markov transition matrix modeling, which includes: constructing a discrete Markov chain and counting the state transition frequency matrix of adjacent bytes in the binary byte stream; normalizing and truncating the transition frequency matrix to generate a 256×256 dimensional Markov grayscale image; (2) Extract image features through the multi-scale dilated convolution fusion attention module MDFA, including: using parallel dilated convolution layers to capture local texture, regional structure and global semantic features with different dilation rates; combining channel attention mechanism and spatial attention mechanism to dynamically weight key channels and spatial regions; (3) Classification is performed using the SimMobileNetV2 network, which includes: multiple serially connected SimBottleneck modules, each of which embeds a SimAM attention module in an inverted residual structure, generates 3D attention weights through an energy function, strengthens key features and suppresses redundant information.
2. According to the lightweight malware visualization classification method based on multi-scale features according to claim 1, it is characterized in that: In step (1), the Markov transition matrix is generated by: ; Where L is the length of the byte stream, Indicates the value of the i-th byte; according to Count the transition frequency C(m,n) from state m to state n; where, is an indicator function, which takes 1 when the condition is met and 0 otherwise; Traverse the byte stream through the sliding window and build the transfer frequency matrix Frequency values exceeding 255 are truncated to 255.
3. According to the lightweight malware visualization classification method based on multi-scale features according to claim 1, it is characterized in that: In step (2), the operation of the MDFA module includes: using parallel atrous convolutional layers with atrous rates of 1, 3, and 5 to extract short-range, medium-range, and long-range features, respectively; and concatenating the 1×1 convolutional features of the local feature enhancement path with the pooled features of the global semantic path through a global-local feature fusion strategy, and performing adaptive weighting.
4. According to the lightweight malware visualization classification method based on multi-scale features according to claim 3, it is characterized in that: The calculation of the channel attention mechanism includes: performing channel average pooling on the input feature map to generate a channel descriptor; generating channel weights through dimensionality reduction convolution and activation function, and multiplying them channel by channel with the original feature map.
5. According to the lightweight malware visualization classification method based on multi-scale features according to claim 3, it is characterized in that: The calculation of the spatial attention mechanism includes: performing a 1×1 convolution on the input feature map to compress the channel dimension and generate a spatial weight map; normalizing the weights through the Sigmoid function and then multiplying them position by position with the original feature map.
6. The lightweight malware visualization classification method based on multi-scale features according to claim 1 is characterized in that: In step (3), the SimAM attention in the SimBottleneck module is implemented by the following steps: define the energy function to calculate the importance of neurons and obtain the minimum energy ; The formula is as follows: ; in and The target neuron and other neurons in the same channel The linear transformation parameters of is the number of neurons in the channel; and represents the neuron linear transformation parameters; and is the set binary label value, used to distinguish the target neuron t from other neurons ; Adding the regularization term, the final energy function is as follows: ; The final minimum energy can be calculated by the following formula: ; in, and Respectively represent the mean and variance of all neurons in the channel except t; according to The feature map is refined; wherein, Contains all channels and spatial dimensions .
7. The lightweight malware visualization classification method based on multi-scale features according to claim 1 is characterized in that: In step (3), the structure of the SimMobileNetV2 network includes: the bottom-level MDFA module extracts multi-scale features; the middle-level stacks 7 SimBottleneck modules to gradually abstract regional structural features; the high-level outputs the classification results through global average pooling and fully connected layers.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Intelligent terminal malicious software dynamic detection method based on system call
CN109753801A
Malicious software detection and family classification method based on MAAM and CliqueNet
CN113836530A
Winter fish classification and identification method based on MG-ShuffleNet network structure
CN118942123A
Cited By
Lightweight malicious software classification method based on multi-feature fusion
CN120448932A
A lightweight malware classification method based on multi-feature fusion
CN120448932B
PRNU anonymity method based on multi-scale and hierarchical feature fusion
CN120599058A