Malware Detection Model Visualization via Unified Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing malware detection models struggle to effectively integrate and visualize the performance of multiple sources of classification, leading to inconsistent and biased evaluations of malicious artifacts.

Innovation Solution

An apparatus and method that generate labels for the maliciousness of artifacts and evaluate the performance of multiple sources of classification by determining aggregate measures of performance, allowing for a unified label to be assigned to each artifact based on multiple sources of classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple sources of classification are integrated to evaluate malware detection models, then the reliability and comprehensiveness of the evaluation improves, but the complexity of the system increases

Engineering Contradiction:
Improveevaluation reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the evaluation system into distinct modules: data collection from multiple classification sources, data processing and normalization, model evaluation components, and visualization interfaces. This segmentation allows each module to handle specific tasks independently, managing complexity while maintaining comprehensive multi-source evaluation capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary components including a standardized data format layer that mediates between diverse classification sources and the evaluation engine, and an aggregation layer that intermediates between multiple models and their combined assessment. These intermediaries simplify integration of multiple sources while preserving evaluation reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple sources of classification are integrated to reduce biases, then the measurement precision of malware detection evaluation improves, but the device complexity increases

Engineering Contradiction:
Improveevaluation precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms diverse classification outputs into standardized parameters through normalization processes. Different classification sources with varying output formats are converted to a common parameter space, enabling precise comparison and aggregation while managing system complexity through parameter standardization

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal evaluation framework that can process and integrate multiple types of classification sources (different malware detection models, classification algorithms, and evaluation metrics) through a single standardized interface, improving measurement precision across diverse sources without proportionally increasing complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If comprehensive data from multiple classification sources is collected and processed, then the information completeness improves, but the loss of time in processing increases

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent implements preliminary actions including pre-defining evaluation criteria, pre-normalizing data formats from common classification sources, and pre-establishing aggregation rules. This preparation work is done beforehand to enable faster processing when actual evaluation occurs, maintaining information completeness while reducing processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent allows for partial evaluation where not all classification sources need to be processed in every evaluation scenario. Users can select subsets of sources based on specific needs, enabling faster processing when full comprehensiveness is not required, while still maintaining the option for complete information gathering when needed

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250119451A1Methods and apparatus for visualization of machine learning malware detection models
Publication Date: 2025.04.10 SOPHOS LTD
  • US20250119451A1 patent drawing
  • US20250119451A1 patent drawing
  • US20250119451A1 patent drawing

AI summary

Embodiments disclosed include methods and apparatus for visualization of data and models (e.g., machine learning models) used to monitor and/or detect malware to ensure data integrity and/or to prevent or detect potential attacks. Embodiments disclosed include receiving information associated with artifacts scored by one or more sources of classification (e.g., models, databases, repositories). The method includes receiving inputs indicating threshold values or criteria associated with a classification of maliciousness of an artifact and for selecting sample artifacts. The method further includes classifying and selecting the artifacts, based on the criteria, to define a sample set, and based on the sample set, generating a ground truth indication of classification of maliciousness for each sample artifact in the sample set. The method further includes using the ground truth indications to evaluate and display, via an interface, a representation of a performance of sources of classification and/or quality of data.