Zero-confusion sample dependence anti-confusion Android malicious software detection method
By extracting static features from Android applications and constructing anti-obfuscation feature images, and combining deep learning models for feature fusion and classification, the problem of insufficient robustness of the existing Android malware detection methods under obfuscation technology is solved, and efficient malware detection is achieved.
Patent Information
- Application Number
- CN202510682876.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-29
AI Technical Summary
The existing Android malware detection methods are not robust in the face of obfuscation technology, their detection performance is easily affected, and they need to rely on obfuscation samples for training, which lacks adaptability and flexibility.
By decompiling Android applications, static features such as Dex bytecode, sensitive API and permission fields are extracted, and feature changes are analyzed using different obfuscation techniques, anti-obfuscation feature images are constructed, and feature fusion and classification are used for deep learning models to achieve detection without relying on obfuscated samples.
In a variety of obfuscation scenarios, the accuracy and recall of detection are significantly improved, the robustness and adaptability of the model are enhanced, and it can effectively deal with unknown or complex obfuscation technologies.
Smart Images

Figure K9EZNMSLHUN9MLYGJCQKCHTKXWK4BCF4RV6WXBU8 
Figure RDUUUDUFZR6JTGO3G87T3MDXWASLBRYGXAISSBDJ 
Figure SAZUAMTC41H7GJREHINQBBCKXCYGQBEF5PFVUQCA
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer security technology, and particularly to an anti-obfuscation Android malware detection method that is independent of obfuscated samples. Background Art
[0002] With the wide application of the Android operating system, malware has become a major threat to the security of mobile devices. Existing malware detection methods are mainly divided into two categories: static analysis and dynamic analysis. Static analysis can identify potential malicious behaviors without executing the program by scanning the source code or binary files of the application; while dynamic analysis discovers malicious operations by monitoring the behaviors of the program during runtime. However, with the wide application of obfuscation techniques, these traditional detection methods are facing unprecedented challenges. Obfuscation techniques, such as code obfuscation, API renaming, and control flow obfuscation, greatly obscure the true intentions of malware by modifying the code structure and behaviors of the application, thus significantly reducing the effectiveness of traditional detection methods.
[0003] Although malware detection methods based on machine learning have improved the detection accuracy to a certain extent, due to the fact that most methods do not perform specialized anti-obfuscation processing, they usually exhibit poor robustness and generalization ability when facing obfuscation techniques, resulting in the detection performance being easily affected. In addition, existing detection methods usually require the use of obfuscated data for model training, which makes them lack sufficient adaptability and flexibility when facing rapidly changing obfuscation techniques.
[0004] Therefore, the current technology urgently needs an innovative solution that can cope with various malware attacks and, without relying on obfuscated samples, improve the robustness and generalization ability of the detection model through more efficient feature extraction and fusion strategies. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention proposes an anti-obfuscation Android malware detection method that is independent of obfuscated samples to solve the technical problems of insufficient robustness and easily affected detection performance in malware detection in the face of complex obfuscation techniques in the existing technology.
[0006] The technical solution adopted by the present invention is as follows: In a first aspect, an anti-obfuscation Android malware detection method that is independent of obfuscated samples is provided, including the following steps: Decompile multiple Android applications to extract static features; the static features include Dex bytecode, sensitive APIs, and permission fields; Obfuscate the static features using different obfuscation techniques, then analyze the feature changes brought about by obfuscation, and reconstruct the static features according to the analysis results; Convert the reconstructed static features into an image representation to obtain an anti-obfuscation feature image; Based on a deep learning model, perform feature fusion and classification on the anti-obfuscation feature image to complete the detection of Android malware.
[0007] Further, when extracting static features, use a static analysis tool to decompile each APK file, restore the internal bytecode to structured information, and obtain static features.
[0008] Further, when extracting static features, the sensitive APIs are identified through a custom screening method, and the sensitive APIs whose names start with a predefined string prefix are selected.
[0009] Further, use different obfuscation techniques to obfuscate the APK files, and then analyze the feature changes brought about by obfuscation, including: For Dex bytecode, first obtain the obfuscated samples and extract the same static features before and after obfuscation; for text features, use byte-level difference to compare the text feature changes before and after obfuscation, and evaluate the obfuscation effect based on the byte-level difference metric; for image features, use structural similarity to compare the image features before and after obfuscation and calculate the image similarity.
[0010] Further, use different obfuscation techniques to obfuscate the APK files, and then analyze the feature changes brought about by obfuscation, including: For sensitive APIs, analyze the change in the number of APIs and the name similarity after obfuscation, and use text similarity to evaluate the degree of damage to the API call path structure caused by obfuscation.
[0011] Further, use different obfuscation techniques to obfuscate the APK files, and then analyze the feature changes brought about by obfuscation, including: For permission fields, quantify the interference intensity of obfuscation on the permission expression form by comparing the changes in the number and name of permission fields before and after obfuscation.
[0012] Further, reconstruct the static features according to the analysis results, including: Convert the Dex file into a Markov matrix, map the transition probability of adjacent bytes in the byte sequence to a two-dimensional matrix structure, where each element represents the conversion probability between adjacent byte values; For sensitive APIs, directly use the initially extracted API set; For permission fields, adopt an improved preprocessing method for obfuscation robustness, only retain the last part of the permission name and remove the redundant prefix part.
[0013] Further, convert the reconstructed static features into an image representation, including: Map the Markov matrix obtained by converting the Dex bytecode into a grayscale image, where each element in the matrix corresponds to a grayscale value, used to represent the transition probability between adjacent bytes; Convert the sensitive API and permission fields into semantic vectors, and perform normalization processing on the semantic vectors, so that the normalized values fall within a unified pixel grayscale range.
[0014] Further, perform feature fusion and classification on the anti-confusion feature images based on a deep learning model to complete the detection of Android malware, including: Use a shared convolutional layer to extract local patterns and global structure information from different types of anti-confusion feature images respectively, and fuse them after unifying the encoding of the local patterns and global structure information; Adopt a spatial pyramid pooling module for multi-scale pooling, and splice the feature vectors after multi-scale pooling to form a unified feature representation; Input the unified feature representation into a fully connected neural network for classification to identify Android malware.
[0015] In a second aspect, a computer program product is provided, including a computer program / instructions, which when executed by a processor, implement the steps of the anti-confusion Android malware detection method described in the first aspect that is independent of obfuscated samples.
[0016] As can be seen from the above technical solutions, the beneficial technical effects of the present invention are as follows: By analyzing the impact of obfuscation on different static features, the present invention designs and constructs multiple static features, and uses a deep learning model for feature fusion and classification; without relying on obfuscated samples to participate in training, the present invention can effectively cope with various obfuscation techniques and maintain strong robustness and accuracy. The present invention has broad application prospects in the field of Android malware detection and can provide effective protection against various malware attacks and obfuscation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally denoted by similar reference numerals. In the drawings, the elements or parts do not necessarily draw according to the actual scale.
[0018] Figure 1 It is a flowchart of the Android malware detection method according to the embodiment of the present invention; Figure 2 Flow chart for analyzing the impact of obfuscation processing on different static features in an embodiment of the present invention; Figure 3 Schematic diagram of the feature fusion and classification process in an embodiment of the present invention. Detailed implementation manners
[0019] Hereinafter, embodiments of the technical solution of the present invention will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and thus are only examples and cannot be used to limit the protection scope of the present invention.
[0020] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meanings understood by those skilled in the art to which the present invention belongs.
[0021] Embodiment This embodiment provides an anti-obfuscation Android malware detection method that does not rely on obfuscated samples, and the method includes the following steps: Step 1: Collect an Android software dataset, and decompile multiple Android applications in the dataset to extract static features, including Dex files, sensitive APIs, and permissions; In this step, a dataset covering various types of Android applications is first constructed, and the Android applications of the type are processed in combination with decompilation tools to extract static features for subsequent analysis and modeling. Specifically as follows: To ensure the diversity and representativeness of the samples, multiple public datasets are selected, including AndroZoo, CICAndMal2017, and Drebin+GooglePlay. These datasets cover Android application samples of different versions, sources, and functional categories. Among them, the AndroZoo dataset contains 9,308 malicious samples and 17,139 benign samples, CICAndMal2017 contains 1,700 malicious samples and 426 benign samples, and the Drebin+GooglePlay dataset contains 6,147 malicious samples and 5,469 benign samples. By integrating the above resources, an Android application dataset with a reasonable distribution and wide coverage is constructed.
[0022] Subsequently, a static analysis tool (such as Androguard) is used to decompile each APK (Android application) file, restore its internal bytecode to structured information, and extract static features. The static features mainly include Dex bytecode, sensitive APIs, and permission fields.
[0023] Dex bytecode is the core execution logic of Android applications, carrying the main behavior path of the program.
[0024] Sensitive APIs involve key operations such as system services, permission control, and network communication, and are important bases for identifying potential malicious behaviors; in some embodiments, sensitive APIs are identified through a custom screening method, which is based on in-depth analysis of common system or third-party APIs, and selects those sensitive APIs whose API names start with a predefined string prefix; different from the simple API selection of traditional methods, this method accurately screens APIs closely related to functions such as system service interaction, permission management, and network operations, ensuring that the selected APIs have extremely high recognition value for malware detection.
[0025] The permission field reflects the system permissions requested by the application during operation and can reveal the types of sensitive operations it may perform. These static features will serve as the basic data for subsequent analysis of the impact of obfuscation and construction of robust features.
[0026] Step 2: Use different obfuscation techniques to obfuscate the APK file, then analyze the feature changes brought by obfuscation, and reconstruct the static features according to the analysis results In this step, first analyze the interference of various obfuscation techniques on the APK, and on this basis, customize the design and construction of static features to enhance their stability and identifiability in the obfuscated environment. Specifically as follows: Use an obfuscation tool (such as Obfuscapk) to apply various types of obfuscation to each APK sample, and use a static analysis tool (such as Androguard) to extract the static features before and after obfuscation, and then compare and analyze the feature changes brought by obfuscation; the analysis content covers three key features: Dex bytecode, sensitive APIs, and permission fields.
[0027] During analysis, for Dex bytecode, for text features, use byte-level difference comparison to analyze the text feature changes before and after obfuscation, and evaluate the obfuscation effect based on byte-level difference metrics. For image features, evaluate the impact of obfuscation through byte-level similarity calculation and file size change, and analyze the structural perturbation of its image representation form in combination with image similarity metrics such as structural similarity index (SSIM) and mean squared error (MSE). For example: first obtain the obfuscated sample, extract the same static features before and after obfuscation; then use structural similarity (SSIM) to compare the image features before and after obfuscation, and calculate the image similarity using the following formula: where and are the averages of the images and respectively, and is the variance of the image, is the covariance between images, and is a small constant introduced for stability.
[0028] During analysis, for sensitive APIs, analyze the change in the number of obfuscated APIs and name similarity, and use text similarity to evaluate the degree of damage to the API call path structure caused by obfuscation.
[0029] During analysis, for permission fields, quantify the interference intensity of obfuscation on the permission expression form by comparing the quantity and name changes of permission fields before and after obfuscation.
[0030] The above evaluation process comprehensively uses multi-dimensional similarity metrics such as byte-level, structure-level, and image-level (e.g., PSNR) for quantitative analysis, providing data support and basis for subsequent feature construction.
[0031] In the static feature reconstruction stage, for Dex bytecode, obfuscation generally has a significant impact on Dex files, especially changing the file header and data segment; since the data segment contains key logical information, directly deleting these parts may cause a large amount of feature loss; to effectively alleviate this problem, convert the Dex file into a Markov matrix, mapping the transition probability of adjacent bytes in the byte sequence to a two-dimensional matrix structure, where each element represents the conversion probability between adjacent byte values. This method can not only retain the basic structural features of the Dex file but also minimize the interference caused by obfuscation, thus ensuring that the core logical information of the program can be effectively extracted even in an obfuscated environment.
[0032] In the static feature reconstruction stage, for sensitive APIs, analysis shows that they exhibit strong adaptability when facing different obfuscation methods. Especially under the interference of encryption and reflection obfuscation, they can still maintain high stability; therefore, directly use the initially extracted API set without additional processing.
[0033] In the static feature reconstruction stage, for permission fields, adopt an improved preprocessing method for obfuscation robustness, that is, only retain the last part of the permission name and remove the redundant prefix part to eliminate the influence of the obfuscation prefix. For example, the original permission name is: "com.kunpeng.babyting.permission.MEDIA_PLAYER" may be obfuscated to "p4d236d9a.p4f4a8fc4.pfcbeb739.permission.MEDIA_PLAYER". By retaining the stable field "MEDIA_PLAYER", the anti-obfuscation ability of permission features can be effectively improved, and its core semantics can be retained, providing key information related to malicious behavior detection and ensuring higher detection accuracy.
[0034] Step 3: Convert the reconstructed static features into an image representation to obtain the anti-obfuscation feature image In this step, the preprocessed static features are uniformly converted into images to meet the input format requirements of the deep learning model. Specifically as follows: Map the Markov matrix obtained by converting Dex bytecode into a grayscale image, where each element in the matrix corresponds to a grayscale value, representing the transition probability between adjacent bytes; this image form not only retains the structural information of the Dex file but also effectively reduces the interference of obfuscation operations on the expression of image features.
[0035] For sensitive APIs and permission fields, use text classification and word vector representation models (such as the FastText model) to convert sensitive APIs or permission fields into semantic vectors and normalize the semantic vectors so that their values fall within a unified pixel grayscale range, thereby constructing an input representation that can be processed by the image model. Through the above conversion, different types of static features are unified into a two-dimensional image format to obtain the anti-obfuscation feature image, ensuring efficient feature fusion and consistent modeling in subsequent deep models. The anti-obfuscation feature image includes Dex feature images, sensitive API feature images, and permission feature images.
[0036] Step 4: Based on the deep learning model, perform feature fusion and classification on the anti-obfuscation feature image to complete the detection of Android malware First, construct a shared convolutional layer to extract the local patterns and global structural information of different types of anti-obfuscation feature images respectively, and fuse them after unified encoding of the local patterns and global structural information. Then, introduce a spatial pyramid pooling (SPP) module to perform multi-scale pooling operations on feature maps of different sizes to capture multi-level semantic information in the features and avoid information loss caused by inconsistent sizes. After pooling, each feature map is converted into a feature vector with a fixed dimension and spliced and integrated to form a unified feature representation.
[0037] Subsequently, the feature representation is input into a fully-connected neural network for classification. This fully-connected neural network consists of several fully-connected layers, gradually compressing the feature space and enhancing the feature discriminability, and finally outputting the prediction results of malicious or benign software. In the training stage, the model uses cross-entropy as the loss function and the Adam optimizer to iteratively update the parameters to achieve a stable and efficient optimization convergence process and improve the classification accuracy.
[0038] In some embodiments, for the anti-obfuscation Android malware detection method that relies on zero obfuscated sample dependence described above, the following iterative optimization and effectiveness verification are performed: First, the malware samples obtained by this method are input into the Android malware classification model for prediction; according to the output results of the classification model, the parameters in the feature selection and preprocessing process are adjusted to further optimize the performance of the classification model. Specifically, the optimization process involves refining the input anti-obfuscation feature images to improve the robustness of the model when facing obfuscated samples.
[0039] To verify the effectiveness of the proposed Android malware detection method of the present invention, experiments were conducted on the AndroZoo2017 dataset. This dataset contains 9,308 malicious samples and 17,139 benign samples. During the experiment, indicators such as accuracy, precision, recall, and F1-score were used to evaluate the detection performance of the model.
[0040] The experimental results show that in the test of the original dataset (without obfuscation), the accuracy of the method of the present invention is 94.44%, the precision is 93.22%, the recall is 95.30%, and the F1-score is 94.25%. In the test of renamed obfuscation, the accuracy of the method of the present invention is 89.45% and the F1-score is 85.91%, showing strong robustness. In the encrypted obfuscation test, the recall is 92.16% and the F1-score is 90.97%, indicating that the method has good adaptability when dealing with encrypted malware.
[0041] Through the above verification tests, it can be seen that the Android malware detection method provided in this embodiment can effectively improve the detection ability of Android malware, especially showing strong anti-interference in obfuscated samples; through experimental results verification, the proposed detection method is superior to existing methods in multiple indicators such as accuracy, recall, and F1-score, indicating that the method has strong potential in practical applications.
[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered within the scope of the claims and the description of the present invention.
Claims
1. A method for detecting anti-confusion Android malware that is independent of zero-confusion samples, characterized in that, It includes the following steps: Decompile multiple Android applications and extract static features; The static features include Dex bytecode, sensitive APIs, and permission fields; Use different obfuscation techniques to obfuscate the static features, then analyze the feature changes brought by obfuscation, and reconstruct the static features according to the analysis results; Convert the reconstructed static features into an image representation to obtain an obfuscation-resistant feature image; Based on a deep learning model, perform feature fusion and classification on the obfuscation-resistant feature image to complete the detection of Android malware.
2. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, characterized in that When extracting static features, use a static analysis tool to decompile each APK file, restore the internal bytecode to structured information, and obtain static features.
3. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, characterized in that When extracting static features, the sensitive APIs are identified through a custom screening method, and sensitive APIs with API names starting with a predefined string prefix are selected.
4. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, characterized in that, Use different obfuscation techniques to obfuscate the APK files, and then analyze the feature changes brought by obfuscation, including: For Dex bytecode, first obtain the obfuscated samples and extract the same static features before and after obfuscation; for text features, use byte-level difference to compare the text feature changes before and after obfuscation, and evaluate the obfuscation effect based on the byte-level difference metric; for image features, use structural similarity to compare the image features before and after obfuscation and calculate the image similarity.
5. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, characterized in that Use different obfuscation techniques to obfuscate the APK files, and then analyze the feature changes brought by obfuscation, including: For sensitive APIs, analyze the change in the number of APIs and the name similarity after obfuscation, and use text similarity to evaluate the degree of damage to the API call path structure by obfuscation.
6. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, wherein Use different obfuscation techniques to obfuscate the APK files, and then analyze the feature changes brought by obfuscation, including: For permission fields, quantify the interference intensity of obfuscation on the permission expression form by comparing the changes in the number and name of permission fields before and after obfuscation.
7. The anti-confusion Android malware detection method that depends on zero-confusion samples according to claim 1, characterized in that, Reconstruct the static features according to the analysis results, including: Convert the Dex file into a Markov matrix, map the transition probability between adjacent bytes in the byte sequence to a two-dimensional matrix structure, where each element represents the transition probability between adjacent byte values; For sensitive APIs, directly use the initially extracted API set; For permission fields, adopt an improved preprocessing method for obfuscation robustness, only retain the last part of the permission name, and remove the redundant prefix part.
8. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, wherein Convert the reconstructed static features into an image representation, including: Map the Markov matrix obtained by converting Dex bytecode into a grayscale image, where each element in the matrix corresponds to a grayscale value, used to represent the transition probability between adjacent bytes; Convert sensitive APIs and permission fields into semantic vectors, and perform normalization processing on the semantic vectors to make the normalized values fall within a unified pixel grayscale range.
9. The zero-confusion sample-dependent anti-confusion Android malware detection method according to claim 1, characterized in that, Based on a deep learning model, perform feature fusion and classification on the obfuscation-resistant feature image to complete the detection of Android malware, including: Use a shared convolutional layer to separately extract local patterns and global structure information from different types of anti-aliasing feature images, and fuse them after unified encoding of the local patterns and global structure information; Adopt a spatial pyramid pooling module for multi-scale pooling, and splice the feature vectors after multi-scale pooling to form a unified feature representation; Input the unified feature representation into a fully connected neural network for classification to identify Android malware.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-9 are implemented.