LLM-based few-sample multi-label Android malicious software detection method

Through the core set strategy and chain thinking module of the LeoDroid framework, a large language model is used to perform multi-label detection of Android malware, solving the detection problems in noise and data scarcity environments, and achieving high precision and rapid adaptability.

CN120337218APending Publication Date: 2025-07-18TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510479388.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively carry out multi-label detection of Android malware in the environment of scarce noise data and data, especially when malware rapidly evolves and labels are inconsistent, traditional methods are difficult to maintain high detection accuracy.

Method used

Using the LeoDroid framework based on large language models, representative samples are selected through core set strategies, and combined with label description and chain thinking module design structured prompts, LLM is guided to perform multi-label classification tasks to optimize the performance of the model in the case of few samples.

Benefits of technology

It significantly improves the robustness and accuracy of malware detection, which is better than traditional methods, and can maintain high detection accuracy in noise and data scarcity environments, quickly adapt to malware changes, and reduce model update costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337218A_ABST
    Figure CN120337218A_ABST
Patent Text Reader

Abstract

The invention provides a few-sample multi-label Android malicious software detection method based on LLM, and solves the major challenge of keeping stable malicious software detection performance under the condition of data noise and label inconsistency. Two main innovations are introduced into the provided LeoDdroid framework to deal with the challenges. Firstly, a complex core set strategy is realized, representative samples are carefully selected, and the influence of noise is reduced to the maximum extent. Secondly, the advanced reasoning ability of a large language model is utilized through a customized prompt project. In addition, a novel Multi-Sample-ACC measure is introduced, and the measure provides more meaningful evaluation for the multi-label classification performance in the malicious software detection context. The method is characterized in that a consistent MS-ACC (Maximum Sequence-Adaptive Cracking Code) score, which is realized by the LoDandroid on an anonymmouscept data set, a Drebin data set and a VirusShare data set, is higher than 0.93. Due to the powerful framework, the framework is superior to a traditional machine learning method by more than 300% in an anonymmousert data set. These results verify the framework's ability to maintain high detection accuracy with varying degrees of noise and data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Android malware detection, and proposes a few-shot multi-label detection framework based on large language models to improve the robustness of malware detection in noisy data and data-scarce environments. Background Art

[0002] Data noise is a fundamental challenge in Android malware detection, which significantly degrades the performance of machine learning models. Traditional methods using third-party services introduce inconsistencies due to malware evolution and changing time stamps, while deep learning methods struggle to handle noisy data as they rely on large clean datasets and tend to memorize rather than generalize from noisy labels. The quality of training data plays a fundamental role in determining the effectiveness of machine learning models, and data noise is a particularly important challenge that severely degrades model performance. This challenge becomes particularly evident in Android malware detection systems, where accurately labeling malware samples is essential but difficult to achieve. Despite the widespread reliance on third-party services such as VirusTotal for malware labeling based on a voting mechanism, these methods often introduce a large amount of noise into the dataset. This noise stems from inconsistencies in the labeling method and the evolving characteristics of malware, which leads to a decline in the performance of multiple security tasks such as malware detection, behavior analysis, and home contribution. Due to the rapid evolution of malware variants and their increasing complexity, even the most advanced noise reduction techniques face continuous challenges in maintaining label quality. The practical implications of noise in malware detection are reflected in several aspects that motivate research. A good example emerges in time analysis, where a model trained on fully denoised 2022 data shows reduced effectiveness when applied to 2023 malware samples due to the noise in the new dataset. Expert-driven manual behavior analysis can address this issue, but its resource-intensive nature makes it unsuitable for large-scale applications. The problem is further complicated by the temporal inconsistency of VirusTotal's labeling system, where the same malware sample receives different classification identifiers in different years. Due to these different classifications, it becomes increasingly difficult to maintain consistent detection criteria. Although reducing the sample size helps avoid noise, it creates insufficient training data for traditional machine learning and deep learning models. Specifically, Android malware detection faces several interrelated challenges that require innovative solutions. The main challenge lies in sample selection, where traditional methods often struggle to identify representative samples for model training. Although active learning suggests selecting ambiguous samples from each category, these samples typically contain a high level of noise that can harm model performance. The second major challenge comes from the reduction in the amount of training samples, as traditional machine learning and neural network models exhibit a significant decline in performance when trained on limited data. This limitation is particularly evident in multi-label classification scenarios, where multiple malicious behaviors must be detected simultaneously against the backdrop of the increasing diversity of malware types and the emergence of zero-day vulnerabilities. Due to the complexity of malware behavior, multi-label classification has proven to be more challenging than traditional single-label detection such as family classification or binary detection. Summary of the Invention

[0003] The present invention adopts a two-stage process, which integrates a core set strategy for selecting representative samples and an elaborately designed prompting engineering method. The prompt design combines label descriptions, core set examples, and chain-of-thought reasoning to guide large language models in multi-label classification tasks. Through this integration, LeoDroid effectively manages the balance between the sample size and noise tolerance to maintain a high detection accuracy. The LeoDroid framework was evaluated on three real-world datasets: anonymouscert, Drebin, and VirusShare. The experimental results show that the MS-ACC is higher than 0.93 on all datasets, with excellent performance, exceeding traditional machine learning methods by more than three times on anonymousCERT.

[0004] The specific steps are as follows:

[0005] S1: Extract features from Android applications, including behavioral features such as API calls, permission usage, and intent patterns;

[0006] S2: Adopt clustering techniques based on the core set strategy to select the most representative samples from each category to reduce the impact of noisy samples on the training process;

[0007] S3: Combine label descriptions, core set sample examples, and chain-of-thought reasoning to design structured prompts to guide the LLM in multi-label classification tasks;

[0008] S4: Utilize the powerful reasoning ability of the LLM to perform multi-label classification in the few-shot setting and improve the model's performance in data-scarce environments;

[0009] S5: Verify the performance of the model on different datasets through multiple rounds of experiments and update the model as needed to adapt to new malware features.

[0010] Furthermore, step S1 includes: API calls: Record the API functions called by the application during runtime and their call frequencies; Permission usage: Record the types of permissions requested by the application and their usage frequencies. Intent patterns: Record the intents sent and received by the application and their patterns. Other behavioral features: Include network requests, file operations, and system calls.

[0011] Furthermore, the clustering algorithm adopted in step S2 is a hierarchical clustering method based on the KNN similarity matrix, and the optimal number of clusters is adaptively determined by maximizing the weighted combination of the silhouette coefficient and the Calinski-Harabasz index; its step S2 includes:

[0012] S21: Clustering algorithm: Adopt a hierarchical clustering method based on the KNN similarity matrix, divide the dataset into multiple clusters, and select the center point of each cluster as the core sample. Similarity matrix calculation: Calculate the similarity between samples to form a similarity matrix. Hierarchical clustering: Recursively merge the most similar clusters until the predetermined number of clusters is reached. Cluster center selection: Select the sample with the smallest average distance from other samples within the cluster as the core sample;

[0013] S22: Adaptive number of clusters selection: Adaptively determine the optimal number of clusters by maximizing the weighted combination of the silhouette coefficient and the Calinski-Harabasz index to ensure the robustness of the clustering result. Silhouette coefficient: Measure the intra-cluster similarity and inter-cluster separation. Calinski-Harabasz index: Measure the intra-cluster compactness and inter-cluster separation. Weighted combination: Select the optimal number of clusters by weighted combining the silhouette coefficient and the Calinski-Harabasz index.

[0014] Furthermore, the prompts designed in step S3 include: Label description: Clearly give the definition of each malware behavior label; Core set sample examples: Integrate the feature and label information of the core samples into the prompt; Chain of thought module: Guide the LLM to perform a structured reasoning process by decomposing the relationship between features and labels. Its step S3 includes:

[0015] S31: Label description: Clearly give the definition of each malware behavior label in the prompt to provide a semantic basis for the classification task.

[0016] Label definition: Describe in detail the meaning and features of each malware behavior label.

[0017] S32: Core set sample examples: Integrate the feature and label information of the core samples into the prompt as a template for the LLM to learn.

[0018] Feature information: Includes API calls, permission usage, and intent patterns.

[0019] Label information: The malware behavior label corresponding to each core sample.

[0020] S33: Chain of thought module: Guide the LLM to perform a structured reasoning process by decomposing the relationship between features and labels, reducing errors caused by overly simplified feature interpretations.

[0021] Reasoning steps: Decompose the relationship between features and labels into multiple reasoning steps to gradually guide the LLM for classification.

[0022] Furthermore, the LLM used in step S4 is the Qwen series model. By means of iteration, the LLM is used to classify the core samples to optimize the model performance. Step S4 includes:

[0023] S41: LLM Selection: Select a large language model suitable for handling few-shot learning tasks.

[0024] Model Selection: Select a suitable LLM model according to the task requirements.

[0025] S42: Prompt Input: Input the designed prompt into the LLM to enable multi-label classification in the few-shot case.

[0026] Prompt Format: Integrate the label description, core set sample examples, and the chain of thought module into the prompt.

[0027] S43: Model Training and Inference: By means of iteration, the LLM is used to classify the core samples to optimize the model performance.

[0028] Iterative Training: Optimize the classification performance of the model through multiple rounds of iterative training.

[0029] Furthermore, the performance evaluation metrics adopted in step S5 include multi-sample accuracy, Hamming loss, zero-one loss, and F1 score. Step S5 includes:

[0030] S51: Performance Evaluation: Evaluate the model performance using metrics such as multi-sample accuracy, Hamming loss, zero-one loss, and F1 score.

[0031] Multi-sample Accuracy: Measure the overall accuracy of the model in multi-label classification tasks.

[0032] Hamming Loss: Measure the average error rate of the model in multi-label classification tasks.

[0033] Zero-one Loss: Measure the sample-level error rate of the model in multi-label classification tasks.

[0034] F1 Score: Measure the comprehensive performance of the model in multi-label classification tasks.

[0035] S52: Model Update: According to the experimental results, make necessary updates to the model to adapt to new malware features and data distributions.

[0036] Incremental Learning: Conduct incremental learning on the model by regularly collecting new malware samples.

[0037] Real-time Update: Update the model in real time according to new malware features to maintain the robustness of the model.

[0038] Further, the method further includes performing a noise robustness test on the model by injecting a certain proportion of noise into the training data to verify the performance of the model in a noisy environment.

[0039] Further, the method further includes performing a data scarcity test on the model by reducing the number of training samples to verify the performance of the model in the few-shot case.

[0040] Further, the method further includes performing real-time updates on the model by regularly collecting new malware samples and performing incremental learning on the model to adapt to new malware characteristics.

[0041] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are as follows:

[0042] Significantly improve detection robustness: Through the core set strategy and the chain of thought module, the present invention can maintain high detection accuracy in noisy data and data scarcity environments. Compared with traditional methods, the performance of the present invention on the noisy dataset is particularly significant, especially in the face of label inconsistency problems caused by the rapid evolution of malware, and it can still maintain stable detection effects.

[0043] Optimize performance: Utilizing the powerful reasoning ability of the LLM, the present invention performs excellently in few-shot multi-label classification tasks and is significantly superior to traditional machine learning methods. In the experiment, the LeoDroid framework proposed by the present invention achieved a multi-sample accuracy of over 0.93 on all datasets, while the accuracy of traditional methods under the same conditions was usually lower than 0.6.

[0044] Enhance adaptability: Through adaptive cluster number selection and few-shot learning, the present invention can quickly adapt to new malware families and feature changes, reducing the model update cost. Compared with traditional deep learning methods that rely on a large amount of labeled data, the present invention can still work effectively in data scarcity situations, reducing the dependence on large-scale clean datasets.

[0045] Reduce the model update cost: The present invention performs incremental learning on the model by regularly collecting new malware samples to adapt to new malware characteristics. Compared with traditional methods that require retraining the entire model, the real-time update mechanism of the present invention can significantly reduce the time and computational costs of model updates.

[0046] Improve detection efficiency: Through structured prompt design, the present invention guides the LLM to perform efficient multi-label classification tasks. Compared with traditional methods, when dealing with complex malware behavior patterns, the present invention can complete the classification task faster, improving the detection efficiency.

[0047] Enhancing Model Interpretability: Through the chain of thought module, the present invention can decompose the relationship between features and labels into multiple reasoning steps, gradually guiding the LLM for classification. This not only improves the classification performance of the model but also enhances the interpretability of the model, making the detection results easier to understand and verify. In summary, the present invention shows significant advantages in improving detection robustness, optimizing performance, enhancing adaptability, reducing model update costs, improving detection efficiency, and enhancing model interpretability, and can effectively address the challenges existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 : Schematic diagram of the framework of LeoDroid.

[0049] Figure 2 : Schematic diagrams of MS-ACC and F1-Score performance of the model on different datasets under different parameters.

[0050] Figure 3 : Schematic diagrams of Hamming Loss and Zero-One Loss performance of the model on different datasets under different parameters.

[0051] Figure 4 : Ablation performance comparison chart of anonymous scert.

[0052] Figure 5 : Ablation performance comparison chart on Derbin.

[0053] Figure 6 : Ablation performance comparison chart on VirusShare. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To verify that the model proposed by the present invention effectively balances the trade-off between sample size and noise tolerance while maintaining high detection accuracy and can enhance the robustness of malware detection in noisy and data-scarce environments, the specific steps are as follows:

[0055] (1) Dataset

[0056] This study utilized three different Android malware datasets. The main dataset came from anonymmouscert, which contains expert-verified security reports providing reliable ground truth labels. The DREBIN dataset was also used, which covered 5,560 Android applications during the period from August 2010 to October 2012, covering 179 malware families. The third dataset came from VirusShare, collecting malware samples between 2019 and 2022. Through detailed analysis of these datasets, six different types of malicious behaviors were identified. The feature extraction process followed the method established in previous work b to capture three fundamental aspects of Android applications: API calls, permission usage, and intent patterns. This foundation was enriched by integrating additional insights from recent technical research and documentation, resulting in a comprehensive set of 531 behavioral features. To ensure robust evaluation, 90 noise-free samples were designated from each dataset as the test set. Several measures were designed and implemented to maintain objectivity. A small number of training samples were anonymized to prevent potential biases, and a balanced label distribution was maintained across all datasets. Each LLM-based method received the same prompt structure for fair comparison. The experimental setup included specific considerations for noise handling. Although anonymmouscert samples remained noise-free due to their expert verification, 10% noise was deliberately injected into a small number of samples from DREBIN and VirusShare. This design choice enabled the evaluation of model performance under both ideal and realistic conditions, as noise typically occurs in real-world malware detection scenarios.

[0057] (2) Evaluation Metrics

[0058] (201) Multi-Sample Accuracy. In multi-label classification tasks, accurately predicting all labels for a given sample is inherently challenging, especially when the labeled dataset is limited. Achieving perfectly correct label predictions is often impractical. Previous work has emphasized that traditional classification methods struggle to maintain accuracy across multiple labels simultaneously. To address this issue, the Multi-Sample Accuracy metric is proposed, which introduces a flexible evaluation criterion that allows for a degree of tolerance, making it more suitable for multi-label classification scenarios. For a sample xi, its classification is considered correct if the following two conditions are met: At least one correct label is selected: At least one of the true labels of sample xi is correctly identified, i.e., Yi,j = 1 and Zi,j = 1, where Yi,j represents the ground truth label and Zi,j represents the predicted label. Not all labels are selected: Sample xi is not assigned all possible labels, ensuring that at least one label is not selected (i.e., the model does not indiscriminately choose all labels). This criterion ensures that MS-ACC provides a more realistic evaluation of model performance in multi-label classification, allowing for a degree of tolerance while preventing the trivial solution of all labels being selected indiscriminately.

[0059]

[0060] (202) Hamming Loss. The Hamming Loss metric measures label-level accuracy by calculating the ratio of misclassified labels to the total number of labels across all samples. A reduction in Hamming Loss reflects a closer agreement between the predictions and the ground truth, indicating improved model precision in label assignment.

[0061] (203) Zero-One Loss. Zero-One Loss provides a sample-level evaluation perspective by measuring the proportion of samples in the dataset that have any misclassified labels. Models with lower 0-1 Loss exhibit better overall classification accuracy as they successfully identify the correct label combinations for a larger portion of the samples.

[0062] (204) F1-Score. The F1-Score, as a balanced performance metric, combines precision and recall through their harmonic mean. A high F1-Score indicates that the model performs well both in terms of precise label prediction and comprehensive coverage of true positive cases. Due to this dual focus, the F1-Score provides valuable insights into the practical effectiveness of the model in real-world applications.

[0063] (3) Baseline Methods

[0064] (301) LLM Model

[0065] Qwen2.5: Qwen2.5 is the latest version of the Qwen2 series and has made substantial improvements in various aspects compared to Qwen2. There are several reasons for choosing Qwen2.5. The parameters of this model range from 500 million to 7.2 billion, which can systematically evaluate the impact of scale effects on malware detection performance. Its advanced Transformer architecture, enhanced by RoPE and SwiGLU technologies, provides a solid foundation for processing sequential code features and understanding complex malware patterns. The LeoDroid framework takes the Qwen2.5-7B architecture as its core component. To comprehensively evaluate the effectiveness of the method, Qwen2.5 variants with different parameter scales were compared, ranging from 0.5B to 7B billion parameters. This experimental design can evaluate how model scale affects performance in the malware detection task while verifying the ar architecture selection LongAlign: LongAlign is used as a benchmark model for long-context alignment tasks because it has specialized capabilities in processing extended context sequences. The architecture of this model is based on extensive training on the LongAlign-10k dataset, which contains 10,000 instruction samples ranging from 8K to 64K tokens. This foundation demonstrates the potential advantages of processing complex malware feature sequences.

[0066] (302) Machine learning methods

[0067] Based on the successful implementation of previous multi-label Android malware classification research, MEKA v1.9.2 was used as the basic machine learning framework. This open-source system extends WEKA and provides specialized support for multi-label classification tasks. In this study, CDN was used as the basic classifier and integrated with LMT and REPtree as multi-label classifiers to classify Android malware.

[0068] (4) Experimental environment

[0069] The experiment used a high-performance Ubuntu server, an Intel(R) Xeon(R) Platinum 8336C processor, and an NVIDIA GeForce RTX 4090 GPU to handle the LLM computing requirements. The implementation framework combines Python with PyTorch v2.0.1, Transformers v4.47.0, and CUDA v11.8.89. For the machine learning-based multi-label classification comparison, MEKA v1.9.2

[50] was used, which is a specialized extension of WEKA

[49] and provides comprehensive support for multi-label classification algorithms and evaluation metrics.

[0070] (5) Comparison of multi-label malware classification in anonymousCERT

[0071] Experiments have demonstrated the excellent performance of Qwen7B in all metrics. The MS-ACC of this model is 0.978, which is significantly better than traditional methods, being 95.6% and 266.3% higher than CDN-REPTree and CDN-LMT respectively. The MS-ACC value of Qwen3B is 0.811, but it does not match the performance of Qwen7B. Although Qwen0.5B achieved a relatively low MS-ACC of 0.422, it still outperformed some traditional methods. Traditional learning methods face significant challenges in few-shot learning environments. The MS-ACC of CDN-REPTree reached 0.500, and the performance of CDN-LMT and LongAlign was weaker. Although it was designed for long-context tasks, LongAlign performed the worst, with the lowest MS-ACC of 0.200 and the highest zero-loss of 0.956. Qwen7B also maintained its advantages in other metrics. It achieved the lowest Hamming loss of 0.156 and zero-loss of 0.456, indicating better label prediction accuracy. The F1-Score of this model also led with 0.820, demonstrating a balanced performance of accuracy and recall in effective malware detection. These results show that the qwen7 series of models, especially Qwen7B, perform excellently in few-shot multi-label classification tasks. Due to their advanced reasoning capabilities, they effectively capture the complex relationships between labels in data-scarce scenarios. The LeoDroid model, built on Qwen7B, has a significant advantage over traditional methods in Android malware classification. The framework schematic diagram of LeoDroid is as Figure 1 shown.

[0072] Table 1: Anonymous CERT dataset

[0073]

[0074] (6) Robustness evaluation of the noisy dataset

[0075] To further study the robustness of the proposed method, additional experiments were conducted on two noisy datasets (i.e., Drebin and VirusShare) with a noise level of 10%. As mentioned before, few-shot learning was adopted to ensure that the extracted training samples were minimally affected by noise. Through multiple rounds of experiments, it was confirmed that the core set strategy effectively eliminated the noise in few-shot samples, thereby enhancing the performance of the model.

[0076] (601) Performance on the Drebin dataset. Experiments on the Drebin dataset further verified the superior performance of Qwen7B in multi-label classification tasks. As shown in Table 2, Qwen7B achieved an MS-ACC of 0.978, demonstrating its special ability to capture complex inter-label dependencies in several-shot settings. Qwen3B followed closely, with an MS-ACC of 0.878 slightly lower than Qwen7B but still significantly better than traditional machine learning methods. This highlights the potential of large language models in few-shot learning scenarios. In contrast, Qwen0.5B achieved an MS-ACC of 0.456, which, although not high, still exceeded traditional methods such as CDN-LMT and LongAlign, indicating its practicality in data-scarce environments. In terms of Hamming loss, Qwen7B maintained the leading position with a score of 0.139, significantly lower than other models, indicating its ability to minimize label misclassification errors. Qwen3B achieved a Hamming loss of 0.250, demonstrating its robustness in handling noisy data. However, the Hamming loss of Qwen0.5B was 0.422, indicating limitations in its parameter configuration, which may hinder its ability to capture fine-grained label relationships. For zero loss, Qwen7B again performed excellently with a score of 0.400, showing its ability to effectively avoid misclassification. The zero loss of Qwen3B was 0.578, which, although higher than Qwen7B, was still better than traditional methods. The zero loss of Qwen0.5B was 0.933, further emphasizing its limitations in few-shot learning. Traditional methods such as CDN-LMT and CDN-REPTree showed higher zero loss values of 0.844 and 0.622 respectively, indicating their inability to adapt to noisy, data-scarce environments. In the F1-Score metric, Qwen7B scored 0.835, indicating its ability to effectively balance precision and recall. Qwen3B ranked second with an F1-Score of 0.694, and Qwen0.5B scored 0.297, indicating its limitations in few-shot classification tasks. The F1-Scores of traditional methods such as CDN-LMT and CDN-REPTree were 0.362 and 0.247 respectively, significantly lower than Qwen7B and Qwen3B. The F1-Score of LongAlign was 0.173, further emphasizing its inadequacy in few-shot learning scenarios.

[0077] (602) Performance of the VirusShare dataset. Experiments conducted on the VirusShare dataset, which is sparser and noisier than Drebin, provide further insights into the robustness of the methods. As shown in Table 2, the MS-ACC of Qwen7B is 0.933, significantly outperforming other models. Qwen3B follows closely, with an MS-ACC of 0.922, and Qwen0.5B scores 0.267. Traditional methods such as LongAlign and CDN-LMT perform poorly, with MS-ACC scores of 0.189 and 0.133 respectively. In terms of Hamming loss, CDN-LMT unexpectedly obtains the lowest score of 0.163, indicating its ability to minimize label misclassification errors. Followed by Qwen3B and Qwen7B, with Hamming loss scores of 0.222 and 0.311 respectively. Qwen0.5B and LongAlign exhibit higher Hamming loss values of 0.414, highlighting their limitations in handling noisy data. For Zero-One Loss, Qwen3B has the lowest score of 0.478, indicating its robustness in avoiding misclassification. Followed by Qwen7B, with a score of 0.733, and Qwen0.5B and LongAlign score 0.956 and 0.933 respectively. CDN-REPTree obtains a moderate Zero-One Loss of 0.689. In the F1-Score metric, Qwen3B leads other models with a score of 0.741, and Qwen7B follows closely with a score of 0.644. Compared with traditional methods such as LongAlign and CDN-LMT, Qwen0.5B has significantly lower F1-Scores, which are 0.181, 0.139, and 0.188 respectively.

[0078] (603) Performance on noisy datasets. The performance gap between the Drebin and VirusShare noisy datasets can be attributed to their inherent differences in data quality and sparsity. Although noisy, Drebin contains more structured and representative samples, enabling models like Qwen7B to effectively capture label dependencies. In contrast, VirusShare is sparser and noisier, posing greater challenges for "few-shot learning". Despite these challenges, Qwen7B and Qwen3B demonstrated remarkable robustness, significantly outperforming traditional methods. Experiments on the Drebin and VirusShare datasets showed that LeoDroid could maintain robust performance on different datasets even in the presence of noise and data sparsity. Qwen7B and Qwen3B consistently outperformed traditional methods, demonstrating their ability to adapt to different data conditions. This robustness was particularly evident in their superior performance on the noisy and sparse VirusShare dataset, where traditional methods struggled to achieve competitive results. In summary, the results from the Drebin and VirusShare datasets highlight the robustness of the qwen series of models, especially Qwen7B, in few-shot multi-label classification tasks. Their ability to handle noisy and sparse data, combined with their excellent performance on multiple evaluation metrics, underscores their potential for real-world applications in Android malware detection. On the other hand, traditional methods exhibited significant limitations, further emphasizing the advantages of LLM-based methods in challenging data environments.

[0079] Table 2: Datasets with 10% noise

[0080]

[0081] (7) Impact of model parameter scale on performance

[0082] A key factor influencing the performance of large language models (LLMs) is their parameter scale. As Figure 2 Figure 3 shown by previous studies, under similar conditions, models with larger parameter scales tend to exhibit stronger reasoning capabilities. To investigate this, a comparative analysis of Qwen models with different parameter scales was conducted on the anonymous-mouscert, Drebin, and VirusShare datasets for several multi-label classification tasks of Android malware.

[0083] (701) Performance on the anonymous scert dataset. The experimental results are as Figure 2 Figure 3As shown, with the increase in the scale of the pa - parameter, the performance has been significantly improved. On the anonymous scert dataset, the MS - ACC metric increased from approximately 0.4 for the 0.5B model to 0.8 for the 3B model, and finally exceeded 0.9 for the 7B model. This indicates that there is an obvious positive correlation between the parameter scale and the inference ability of the model in the Android malware multi - label classification task. Similarly, as Figure 2 Figure 3 shown, the F1 - Score metric is positively correlated with the parameter scale. In contrast, the Zero - One Loss and Hamming Loss metrics decrease with the increase in the parameter scale, which indicates that larger models are more capable of reducing prediction errors and classification biases. However, Figure 2 and Figure 3 the slope analysis in shows that the performance improvement is not linear. Although the inference ability of the model increases significantly as the parameter scale increases from small to medium size, after exceeding a certain scale, the improvement rate gradually slows down. This indicates that the return on performance improvement is diminishing. Considering the computational limitations, it is not possible to explore the full range of this trend in larger models. Nevertheless, the research results show that although increasing the parameter scale enhances the inference ability of the model on the anonymous scert dataset, the marginal benefit decreases. Therefore, it is crucial to balance the model scale and computational cost when selecting the optimal scale.

[0084] (702)Performance on the Drebin dataset. A similar trend was also observed on the Drebin dataset, as Figure 4 and Figure 5 shown. The MS - ACC and F1 - Score metrics are positively correlated with the parameter scale, while the Zero - One Loss and Hamming Loss metrics are negatively correlated. This further confirms that increasing the parameter scale can enhance the inference ability of the model in the Android malware multi - label classification task. However, the diminishing marginal effect is also obvious on the Drebin dataset. As Figure 2 and Figure 3 shown, after the parameter scale reaches 3B, the performance improvement rate significantly slows down, which is consistent with the observations on the anonymous scert dataset, strengthening the fact that the benefits of increasing the parameter scale are affected by diminishing returns.

[0085] (703)Performance on the VirusShare dataset. The results on the VirusShare dataset, as Figure 2 shown, Figure 3, presents a more nuanced picture. While the MS-ACC metric is positively correlated with the parameter scale, other metrics (Zero-OneLoss, Hamming Loss, and F1-Score) decline from 3B to 7B. This deviation from the trend observed in the anonymmouscert and Drebin datasets is unexpected and worthy of further investigation. To explain this phenomenon, an in-depth analysis of the VirusShare dataset was conducted. It was found that, compared to the anonymmouscert and Drebin datasets, the data quality of VirusShare is significantly poorer. Specifically, the samples in VirusShare are sparser, with most samples containing fewer than 2-4 features. Additionally, the data distribution is highly imbalanced. These factors may hinder the model's ability to effectively utilize its increased parameter scale, resulting in suboptimal performance. In contrast, the anonymmouscert and Drebin datasets have stronger robustness and evenly distributed data, enabling the model to fully utilize its enhanced inference capabilities. The poor data quality of VirusShare exacerbates the diminishing marginal effect of increasing the parameter scale and may even lead to overfitting, where the model makes unreasonable inferences due to sparse training data - a phenomenon commonly referred to as "hallucination".

[0086] (704) Summary: The experimental results show the impact of model scale on the performance of few-shot multi-label malware classification. The systematic analysis indicates that a larger parameter scale generally leads to an improvement in model capabilities. The parameter size is positively correlated with the MS-ACC and F1-Score metrics and negatively correlated with the Zero-OneLoss and Hamming Loss metrics, providing strong evidence for this relationship. Despite this overall trend, the findings reveal important nuances in the scale behavior. Although larger models tend to perform better, beyond certain parameter thresholds, the improvement shows diminishing returns. The benefits of scaling also depend heavily on dataset characteristics, as demonstrated by the unexpected performance patterns observed in the VirusShare dataset. These insights have important practical implications for model deployment. Due to the comprehensive evaluation, it can be concluded that the optimal model selection requires careful consideration of dataset attributes and computational resource constraints. The choice of parameter scale must balance the potential performance gains and practical limitations to achieve efficient and effective malware detection.

[0087] (8) Ablation Study

[0088] To validate the effectiveness of each module in the proposed LeoDroid model, a comprehensive ablation study was conducted. This study aimed to evaluate the contribution of the COT module and the choice of few-shot and zero-shot learning methods. The experiments were conducted on the anonymmouscert, Drebin, and VirusShare datasets, and the results are asFigure 4 Figure 5 Figure 6 as shown

[0089] (801) Performance of the anonymmouscert dataset. As shown in the figure Figure 4 as shown, the impact of the COT module was analyzed, and the performance of few-shot and zero-shot learning methods on the anonymmouscert dataset was compared. The results show that, on multiple metrics, the few-shot learning method significantly outperforms the zero-shot learning. Specifically, the improvement in MS-ACC indicates an increase in accuracy in multi-label classification tasks, while the increase in F1-Score highlights the ability of few-shot learning to achieve a better balance between accuracy and recall. Additionally, the reduction in Zero-One Loss and Hamming Loss further emphasizes the superiority of few-shot learning in minimizing classification errors. These findings reveal the crucial role of few-shot learning in improving model performance, especially when compared to the contribution of the COT module. Furthermore, from a computational perspective, the overhead difference between few-shot and zero-shot methods is negligible, making few-shot learning the preferred choice.

[0090] (802) Performance of Drebin and VirusShare datasets. In the experiments on Drebin and VirusShare datasets, zero-shot and few-shot learning methods were not directly compared. This is because these datasets contain noise in their training environments, and few-shot learning was adopted to select more reliable samples for training, thus mitigating the impact of noise on model performance. In such scenarios, zero-shot learning is incompatible with noise control strategies and is therefore not suitable for fair comparison. However, as shown in Table 2, the few-shot learning method continues to demonstrate its advantages on these datasets, and the combination with the COT module further improves the model's performance. Ablation studies also emphasize the importance of the COT module in the LeoDroid model. As Figure 4 Figure 5 Figure 6As shown, the COT module significantly improves the model's performance in few-shot and zero-shot settings. Removing the COT module results in a significant decline in the evaluation metrics of all components, highlighting its crucial role in multi-label classification tasks. For example, on the anonymmouscert dataset, removing the COT module caused the MS-ACC to drop from 0.978 to 0.850 and the F1-Score to drop from 0.820 to 0.680. Similarly, on the Drebin and VirusShare datasets, the absence of the COT module led to a significant performance decline, further validating its contribution to the entire model.

[0091] (803)Summary: The ablation study provides strong empirical evidence for the effectiveness of each component in the LeoDroid architecture. The Chain-of-Thought module significantly enhances the model's analytical capabilities for malware detection tasks. Due to its structured reasoning method, the model demonstrates an enhanced ability to identify complex malware patterns and their relationships. The few-shot learning method is proven to be superior to the zero-shot method by achieving higher classification accuracy and lower error rates in the evaluation metrics. The synergistic integration of these components creates a powerful framework for multi-label malware classification. Experimental results on different datasets show that removing any one component leads to a measurable performance decline. Even though each component shows individual merits, their combined implementation yields the strongest results. These findings validate the architectural decisions and lay the foundation for effective few-shot multi-label classification in malware detection scenarios.

Claims

1. A few-shot multi-label Android malware detection method based on LLM, characterized in that It includes the following steps: S1: Extract features from Android applications, including behavioral features such as API calls, permission usage, and intent patterns; S2: Adopt a clustering technique based on the core set strategy to select the most representative samples from each category to reduce the impact of noisy samples on the training process; S3: Combine label descriptions, core set sample examples, and chain-of-thought reasoning to design structured prompts to guide the LLM in performing multi-label classification tasks; S4: Utilize the powerful reasoning ability of the LLM to perform multi-label classification in the few-shot setting and improve the performance of the model in data-scarce environments; S5: Verify the performance of the model on different datasets through multiple rounds of experiments and update the model as needed to adapt to new malware features.

2. The few-shot multi-label Android malware detection method based on LLM according to claim 1, characterized in that Step S1 includes: API calls: Record the API functions called by the application during runtime and their call frequencies; Permission usage: Record the types of permissions requested by the application and their usage frequencies; Intent patterns: Record the intents sent and received by the application and their patterns, Other behavioral features: Include network requests, file operations, system calls.

3. The few-shot multi-label Android malware detection method based on LLM according to claim 1, wherein The clustering algorithm adopted in step S2 is a hierarchical clustering method based on the KNN similarity matrix, and the optimal number of clusters is adaptively determined by maximizing the weighted combination of the silhouette coefficient and the Calinski-Harabasz index. Step S2 includes: S21: Clustering algorithm: Adopt a hierarchical clustering method based on the KNN similarity matrix to divide the dataset into multiple clusters and select the center point of each cluster as the core sample. Similarity matrix calculation: Calculate the similarity between samples to form a similarity matrix. Hierarchical clustering: Recursively merge the most similar clusters until the predetermined number of clusters is reached. Cluster center selection: Select the sample with the smallest average distance to other samples within the cluster as the core sample; S22: Adaptive number of clusters selection: Adaptively determine the optimal number of clusters by maximizing the weighted combination of the silhouette coefficient and the Calinski-Harabasz index to ensure the robustness of the clustering result. Silhouette coefficient: Measure the intra-cluster similarity and inter-cluster separation. Calinski-Harabasz index: Measure the intra-cluster compactness and inter-cluster separation. Weighted combination: Select the optimal number of clusters by weighted combining the silhouette coefficient and the Calinski-Harabasz index.

4. The few-shot multi-label Android malware detection method based on LLM according to claim 1, characterized in that, The prompts designed in step S3 include: Label description: Clearly give the definition of each malware behavior label; Core set sample examples: Integrate the features and label information of the core samples into the prompts; Chain-of-thought module: Guide the LLM in performing a structured reasoning process by decomposing the relationship between features and labels. Step S3 includes: S31: Label description: Clearly give the definition of each malware behavior label in the prompt to provide a semantic basis for the classification task; Label definition: Describe in detail the meaning and features of each malware behavior label; S32: Core set sample examples: Integrate the features and label information of the core samples into the prompts as a template for the LLM to learn; Feature information: Includes API calls, permission usage, intent patterns; Label information: Malware behavior labels corresponding to each core sample; S33: Chain-of-Thought Module: By decomposing the relationship between features and labels, guide the LLM to perform a structured reasoning process and reduce errors caused by overly simplified feature explanations; Inference steps: Decompose the relationship between features and labels into multiple inference steps and gradually guide the LLM for classification.

5. The few-shot multi-label Android malware detection method based on LLM according to claim 1, wherein, The LLM used in step S4 is the Qwen series of models. Through an iterative approach, use the LLM to classify core samples and optimize the model performance; Its step S4 includes: S41: LLM Selection: Select a large language model suitable for handling few-shot learning tasks; Model selection: Select a suitable LLM model according to task requirements; S42: Prompt Input: Input the designed prompt into the LLM to enable multi-label classification in the few-shot case; Prompt format: Integrate label descriptions, core set sample examples, and the Chain-of-Thought Module into the prompt; S43: Model Training and Inference: Through an iterative approach, use the LLM to classify core samples and optimize the model performance; Iterative training: Optimize the classification performance of the model through multiple rounds of iterative training.

6. The few-shot multi-label Android malware detection method based on LLM according to claim 1, characterized in that, The performance evaluation metrics used in step S5 include multi-sample accuracy, Hamming loss, zero-one loss, and F1 score; Its step S5 includes: S51: Performance Evaluation: Evaluate the model performance using metrics such as multi-sample accuracy, Hamming loss, zero-one loss, and F1 score; Multi-sample accuracy: Measure the overall accuracy of the model in multi-label classification tasks; Hamming loss: Measure the average error rate of the model in multi-label classification tasks; Zero-one loss: Measure the sample-level error rate of the model in multi-label classification tasks; F1 score: Measure the comprehensive performance of the model in multi-label classification tasks; S52: Model Update: According to the experimental results, make necessary updates to the model to adapt to new malware features and data distributions; Incremental learning: Regularly collect new malware samples to perform incremental learning on the model; Real-time update: Update the model in real time according to new malware features to maintain the robustness of the model.

7. The few-shot multi-label Android malware detection method based on LLM according to claim 1, wherein The method also includes performing a noise robustness test on the model. By injecting a certain proportion of noise into the training data, verify the performance of the model in a noisy environment.

8. The few-shot multi-label Android malware detection method based on LLM according to claim 1, wherein The method also includes performing a data scarcity test on the model. By reducing the number of training samples, verify the performance of the model in the few-shot case.

9. The few-shot multi-label Android malware detection method based on LLM according to claim 1, characterized in that, The method also includes performing a real-time update on the model. By regularly collecting new malware samples, perform incremental learning on the model to adapt to new malware features.

Citation Information

Cited By

  • Malware family classification method

    CN120579008A

  • Malware family classification method

    CN120579008B