Host Deduplication for Accurate Vulnerability Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Vulnerability detection and management in IT systems is challenging due to dynamic changes in configurations, software, and exposure to evolving threats, with scanners providing inconsistent and incomplete data, leading to difficulties in accurate and efficient detection and management.
Innovation Solution
A system and method for container image deduplication that processes scanner data to identify and manage vulnerabilities by matching data bits with a container image dataset, using a tiered set of rules to ensure accurate tracking and display of vulnerability status over time, and a graphical user interface for management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple scanners are used to detect vulnerabilities in IT systems, then detection coverage is improved, but data inconsistency and re-identification errors increase
Solution Approach 1:
The patent segments the asset identification process into multiple hierarchical levels: primary identification using consistent unique identifiers, secondary identification using physical characteristics, and tertiary identification using operational patterns. This segmentation allows different scanners to contribute different types of data without causing re-identification errors, as each scanner can focus on specific identification aspects.
Solution Approach 2:
The patent introduces an intermediary normalization layer that receives data from multiple scanners, standardizes the data formats, and resolves inconsistencies before processing. This intermediary layer acts as a mediator that transforms heterogeneous scanner outputs into a unified format, enabling accurate matching and deduplication across different scanning tools.
2Productivity
If scanner data is processed without deduplication, then processing speed is maintained, but vulnerability management accuracy deteriorates due to duplicate assets
Solution Approach 1:
The patent performs preliminary deduplication processing during data ingestion, matching and consolidating duplicate asset records before they enter the main vulnerability management workflow. This preliminary action prevents duplicates from propagating through subsequent processing stages, ensuring accurate vulnerability status without requiring slow post-processing deduplication.
Solution Approach 2:
The patent changes the parameter of asset identification from relying on single scanner-specific identifiers to using multiple parameters including unique persistent identifiers, physical characteristics, and operational patterns. This parameter transformation enables accurate deduplication while maintaining processing efficiency through optimized matching algorithms.
3Reliability
If comprehensive asset data is collected from multiple sources, then vulnerability detection completeness is improved, but data complexity and processing difficulty increase
Solution Approach 1:
The patent segments comprehensive asset data into distinct categories: identification data, configuration data, vulnerability data, and contextual data. Each category is processed and stored separately with optimized schemas, reducing overall system complexity while maintaining complete vulnerability detection capabilities through integrated querying across segments.
Data Source
AI summary
Disclosed are methods, systems and non-transitory computer readable memory for container image or host deduplication in vulnerability management systems. For instance, a method may include: obtaining source data from at least one source, wherein the source data includes a plurality of assets and/or findings; extracting data bits for each asset or finding from the source data; determining a first asset or finding concerns a first container image or first host based on the data bits for the first asset or finding; in response to determining the first asset or finding concerns the first container image or first host, obtaining a container image dataset or a search structure; determining whether the data bits match any of the plurality of sets of values of the container image dataset or the search structure; and, based on a match result, generating or updating records for the first container image or the first host.


