Machine Learning Malicious Program Identification System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying malicious programs rely heavily on manual analysis and character string signatures, leading to low efficiency and hysteresis, making it difficult to detect new viruses and allowing virus makers to evade detection.
Innovation Solution
A method and device for program identification using machine learning that analyzes unknown programs, extracts features, coarsely classifies them, and uses decision-making machines trained on large datasets to accurately identify malicious programs, improving efficiency and preventing undetected threats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual analysis and character string signature methods are used to identify malicious programs, then identification accuracy for known viruses is maintained, but identification efficiency is low and hysteresis occurs making it difficult to detect new viruses
Solution Approach 1:
The system performs preliminary action by pre-training the machine learning model offline with extensive virus samples and features before actual detection. The model is prepared in advance with learned patterns and characteristics, enabling rapid detection of new viruses without manual analysis delays. This resolves the hysteresis problem by having the identification capability ready beforehand rather than requiring real-time manual intervention.
Solution Approach 2:
The patent replaces the mechanical manual analysis system with an automated machine learning-based identification system. Instead of relying on human analysts to manually examine and classify viruses, the system uses trained models that automatically analyze program features and identify malicious software. This substitution dramatically improves identification efficiency and eliminates the time loss associated with manual processing while maintaining high accuracy for detecting both known and new viruses.
2Ease of manufacture
If simple feature or rule-based identification is used, then implementation is straightforward, but virus makers can easily evade detection
Solution Approach 1:
The system applies parameter changes by transforming the detection approach from simple rule-based parameters to complex multi-dimensional feature parameters. Instead of using basic characteristics that viruses can easily bypass, the model analyzes numerous parameters including code structure, behavior patterns, and statistical features. This parameter transformation maintains implementation feasibility through automated processing while significantly improving detection reliability by making it much harder for viruses to evade all parameter checks simultaneously.
Solution Approach 2:
The patent employs composite materials analogy by combining multiple different feature types and analysis methods into a unified detection model. Rather than relying on a single simple feature or rule, the system integrates diverse parameters such as code characteristics, behavioral patterns, and structural features into a composite identification approach. This composite method maintains relative implementation simplicity through systematic integration while greatly enhancing detection reliability by requiring viruses to evade multiple complementary detection mechanisms simultaneously.
3Measurement precision
If extensive manual analysis is performed to improve identification accuracy, then detection precision increases, but the need for experienced personnel increases and efficiency decreases
Solution Approach 1:
The system implements self-service by enabling the machine learning model to automatically perform identification tasks without requiring continuous human intervention or expert analysis. The trained model independently examines program features, applies learned patterns, and generates identification results autonomously. This self-service capability maintains high identification accuracy that would otherwise require experienced personnel while dramatically improving processing efficiency by eliminating the bottleneck of manual expert analysis for each detected program.
Data Source
AI summary
A method and device perform program identification based on machine learning. The method includes: analyzing an inputted unknown program, and extracting a feature of the unknown program; coarsely classifying the unknown program according to the extracted feature; judging by inputting the unknown program into a corresponding decision-making machine generated by training according to a result of the coarse classification; and outputting an identification result of the unknown program. The identification result is a malicious program or a non-malicious program. The method can save a lot of manpower and improve the identification efficiency for a malicious program by using the decision-making machine.


