Malware Data Clustering for Investigation Prioritization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for financial and security investigations require analysts to manually repeat searches, leading to time-consuming and resource-intensive processes, and struggle to prioritize investigations effectively due to insignificant differences between initial data entities.
Innovation Solution
A data analysis system that generates clusters of related data entities from initial 'seed' data entities, using customizable analysis strategies, and assigns scores to prioritize clusters, allowing analysts to start investigations with prioritized clusters rather than individual data entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If analysts manually repeat searches for each investigation, then they can identify related data entities, but the process becomes time-consuming and resource-intensive
Solution Approach 1:
The system pre-computes and stores relationship data between data entities in advance, creating a ready-to-use knowledge base of connections. When an investigation is initiated, analysts can immediately query pre-established relationships without performing manual searches, thus maintaining accurate identification of related entities while dramatically reducing investigation time
Solution Approach 2:
The system creates automated copies of search and analysis processes through algorithms that replicate manual investigation logic. These automated copies can process multiple data entities simultaneously, maintaining the thoroughness of manual analysis while operating at machine speed and handling large volumes of data without additional time cost
2Ease of operation
If analysts examine individual data entities separately, then they can focus on specific suspicious characteristics, but they miss contextual relationships with other entities
Solution Approach 1:
The system implements a hierarchical view where individual data entities are nested within clusters of related entities. Analysts can examine specific entity characteristics in detail while simultaneously viewing the broader contextual cluster, allowing them to maintain focus on suspicious features while understanding relationships to other entities through the nested structural context
Solution Approach 2:
The system introduces relationship data and clustering algorithms as intermediaries between individual data entities. These intermediaries preserve and present contextual relationships without obscuring the specific characteristics of individual entities, enabling analysts to see both the tree and the forest simultaneously through structured relationship mappings
3Productivity
If analysts prioritize investigations based on seed characteristics, then they can manage investigation workload, but insignificant differences between seeds make prioritization difficult
Solution Approach 1:
The system transforms the prioritization process from subjective analyst judgment to objective algorithmic scoring based on multiple parameters. It calculates cluster scores using defined criteria such as relationship density, entity types, and suspicious pattern frequencies, converting vague seed characteristics into quantifiable metrics that enable consistent and scalable prioritization of investigations
Solution Approach 2:
The system creates a universal prioritization framework that handles diverse data entity types and relationship patterns through a single scoring mechanism. This multi-functional approach can evaluate different kinds of investigations (financial fraud, cyber threats, etc.) using the same prioritization logic, simplifying the process while maintaining productivity across various investigation domains
Data Source
AI summary
In various embodiments, systems, methods, and techniques are disclosed for generating a collection of clusters of related data from a seed. Seeds may be generated based on seed generation strategies or rules. Clusters may be generated by, for example, retrieving a seed, adding the seed to a first cluster, retrieving a clustering strategy or rules, and adding related data and/or data entities to the cluster based on the clustering strategy. Various cluster scores may be generated based on attributes of data in a given cluster. Further, cluster metascores may be generated based on various cluster scores associated with a cluster. Clusters may be ranked based on cluster metascores. Various embodiments may enable an analyst to discover various insights related to data clusters, and may be applicable to various tasks including, for example, tax fraud detection, beaconing malware detection, malware user-agent detection, and/or activity trend detection, among various others.


