Threat Intelligence Aggregation Using Self-Supervised Scanner Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional threat detection systems inaccurately assume uniform quality across threat assessment sources and require substantial manually labeled data, leading to sub-optimal performance and high recurring costs.
Innovation Solution
A generative model is developed to learn scanner dependencies and temporal dynamics through self-supervised learning, using three pretext tasks to aggregate threat data without labeled data, and fine-tune the model with semi-supervised or unsupervised methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional aggregation systems assume all threat sources are of similar quality, then the system is simple to implement, but detection accuracy deteriorates
Solution Approach 1:
The patent applies local quality by assigning different quality weights to different threat sources based on their individual performance characteristics. Instead of treating all sources uniformly, the system evaluates each source's accuracy, completeness, and reliability metrics, then applies localized quality adjustments to their contributions in the aggregation process, thereby improving overall detection accuracy while maintaining manageable complexity
Solution Approach 2:
The system changes parameters by dynamically adjusting source quality weights based on observed performance metrics. The quality weights are not fixed but are updated over time based on how accurately each source identifies threats, allowing the system to adapt to changing source reliability without requiring complete system redesign
2Measurement precision
If supervised machine learning models are used for threat detection, then detection accuracy can be improved, but the requirement for manually labeled data increases substantially
Solution Approach 1:
The patent implements self-service by enabling the system to automatically generate quality weights and performance metrics for threat sources without requiring manual labeling. The system uses unsupervised learning techniques to evaluate source quality based on the inherent patterns in the threat data itself, eliminating the need for expensive manual annotation while maintaining detection accuracy
Solution Approach 2:
The system substitutes the mechanical process of manual data labeling with automated computational methods. Instead of human experts manually categorizing threats, the system uses algorithmic approaches to evaluate source quality and detect threats, replacing labor-intensive operations with scalable computational processes
3Measurement precision
If manually labeled data is collected frequently to update models, then detection accuracy is maintained, but recurring costs and time consumption increase
Solution Approach 1:
The patent applies preliminary action by pre-computing quality metrics and weights for threat sources using historical data. Instead of repeatedly collecting and processing labeled data for each model update, the system performs quality evaluation in advance and uses these pre-computed metrics to guide ongoing detection, reducing the frequency and cost of model updates while maintaining accuracy
Data Source
AI summary
Generating high-quality threat intelligence from aggregated threat reports is provided via developing a generative model that identifies relationships between a plurality of threat assessment scanners; pre-training a plurality of individual encoders based on a corresponding plurality of pretext tasks and the generative model; combining the individual encoders into a pre-trained encoder; fine-tuning the pre-trained encoder using threat data; and marking a candidate threat, as evaluated via the pre-trained encoder as fine-tuned, as one of benign or malicious.


