Malware Detection Using Static and Dynamic Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Static malware models require extensive labeled training data and continuous updates to remain effective, which is time-consuming and resource-intensive, and they struggle to detect new types of malware developed by malicious actors.
Innovation Solution
A dual-model approach where a static malware model is used for initial detection, and if inconclusive, a dynamic malware model is executed to determine the file's malware status, with the server retraining the static model using feedback from user devices to distribute updated models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If static malware models are used for malware detection, then detection can be performed before file execution, but the models require extensive labeled training data and continuous updates which are time-consuming and resource-intensive
Solution Approach 1:
The system enables user devices to automatically generate labeled training data by executing files and using the dynamic malware model to determine malware status. This self-service mechanism eliminates the need for manual labeling by security researchers, allowing the static malware model to be continuously retrained with fresh data from the wild without requiring external human intervention or resources.
Solution Approach 2:
The system implements a feedback loop where the dynamic malware model's determinations on executed files are sent back to the server. The server uses this feedback to retrain the static malware model, creating a continuous improvement cycle. This feedback mechanism ensures the static model remains effective against new malware types without requiring manual updates or extensive external training data collection.
2Adaptability or versatility
If static malware models are continuously updated with new training data, then detection effectiveness against new malware types is maintained, but additional resources are required to distribute updated models
Solution Approach 1:
User devices autonomously generate local training data by executing suspicious files and obtaining malware determinations from the dynamic model. This decentralized data generation eliminates the need for centralized collection and distribution of training data, allowing each device to adapt its static model independently using locally generated evidence from new malware encounters.
Solution Approach 2:
The system divides the malware detection adaptation process into independent segments at each user device. Instead of centrally managing and distributing updated models to all devices, each device independently retrains its own static malware model using local feedback from the dynamic model. This segmentation eliminates distribution overhead and allows parallel adaptation across the network.
3Measurement precision
If a dual-model approach is used where dynamic model is executed when static model is inconclusive, then detection accuracy is improved, but additional computational resources are required on user devices
Solution Approach 1:
The system employs a dynamic selection mechanism where the choice of which model to use (static or dynamic) depends on the confidence level of the static model's initial assessment. When the static model's probability falls within the inconclusive range, the system dynamically transitions to using the dynamic model for that specific file. This dynamic approach ensures the more computationally intensive dynamic model is only activated when necessary, optimizing the balance between accuracy and resource consumption.
Data Source
AI summary
In an embodiment, systems and methods for detecting malware are provided. A server trains a static malware model and a dynamic malware model to detect malware in files. The models are distributed to a plurality of user devices for use by antimalware software executing on the user devices. When a user device receives a file, the static malware model is used to determine whether the file contains malware. If the static malware model is unable to make the determination, when the file is later executed, the dynamic malware model is used to determine whether the file contains malware. The file along with the determination made by the dynamic malware model are then provided to the server. The server then retrains the static malware model using the received files and the received determinations. The server then distributes the updated static malware model to each of the devices.


