Malware Detection Metric Estimation via Synthetic Distributions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current cybersecurity evaluation methods focus on correct detections and false alarms, but they do not provide a complete picture as they fail to quantify missed detections accurately and do not allow for a direct calculation of true or false detection rates.
Innovation Solution
The proposed solution involves using statistical properties of known malware distributions to improve estimates of malware detection metrics. This is achieved by generating synthetic sample distributions based on a base data set and observed data, which are then used to identify malware distributions producing detection statistics similar to live target data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional cybersecurity evaluation methods are used that focus on correct detections and false alarms, then the evaluation process is simple and straightforward, but the measurement precision of malware detection accuracy is insufficient and missed detections cannot be quantified
Solution Approach 1:
The patent creates synthetic copies of malware distributions by generating multiple synthetic data sets that replicate the statistical properties of known malware distributions. These synthetic copies are then used to evaluate cybersecurity systems, allowing quantification of missed detections and improved measurement precision without requiring direct access to all actual malware instances.
Solution Approach 2:
The patent performs preliminary actions by pre-generating synthetic data sets with known malware distributions before evaluating the cybersecurity system. This allows the system to have reference data ready for comparison, enabling more accurate measurement of detection precision and missed detections without adding complexity during the actual evaluation process.
2Measurement precision
If synthetic sample distributions are generated to improve malware detection metric estimates, then the measurement precision of detection accuracy improves, but the device complexity and computational resources required increase
Solution Approach 1:
Instead of creating complex systems to analyze every possible malware instance, the patent creates simplified synthetic copies that capture the essential statistical properties of malware distributions. These synthetic data sets serve as representative models that can be processed more efficiently while still providing accurate detection metric estimates.
Solution Approach 2:
The patent changes parameters by generating multiple synthetic data sets with varying statistical parameters that reflect different malware distribution scenarios. This allows the evaluation system to assess performance across a range of conditions without requiring a single complex system, thereby improving measurement precision while managing complexity through parameter variation.
3Reliability
If multiple synthetic data sets are generated and analyzed to quantify missed detections, then the reliability of cybersecurity system evaluation improves, but the loss of time and computational resources increases
Solution Approach 1:
The patent performs preliminary generation of multiple synthetic data sets with known properties before the actual evaluation. This allows the system to have pre-prepared reference data that can be quickly compared against cybersecurity system outputs, improving reliability through multiple reference points while reducing the time penalty during actual evaluation by having data ready in advance.
Solution Approach 2:
The patent creates multiple synthetic copies of malware distributions that can be processed in parallel or pre-processed. These copies serve as reference benchmarks that speed up the evaluation process by providing known answers for comparison, thereby improving reliability without proportionally increasing evaluation time.
Data Source
AI summary
Statistical properties of known malware distributions may be used to improve estimates of malware detection metrics such as a base rate of malicious events in a target environment or missed detections (also referred to as false negatives). In particular, numerous synthetic sample distributions may be generated based on the statistical properties of a base data set and/or additional observed data, and used to identify malware distributions that produce overall detection statistics corresponding to model output for live target data. The malware detection metrics for the live target data can then be characterized using the observed distributions of malware (and malware detections) for the synthetic sample distributions.


