Fingerprint Prevalence Database for Encrypted Traffic

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying processes associated with encrypted network traffic sessions are hindered by the lack of clear-text descriptions in the Transport Layer Security (TLS) protocol, requiring extensive ground truth data that is difficult to collect, especially in environments like IoT, mobile networks, and containers.

Innovation Solution

A technique using passively collected network data to automatically generate a fingerprint prevalence database without endpoint ground truth, clustering observations by fingerprint strings and source/destination context, and annotating clusters with informative names through a rule-based system, allowing for the identification of processes in encrypted traffic sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If TLS fingerprinting is used to identify processes in encrypted traffic, then process identification capability is improved, but the requirement for extensive ground truth data increases collection difficulty and system complexity

Engineering Contradiction:
Improveprocess identification accuracyVSAvoiddatabase construction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically generating the fingerprint prevalence database through passive network traffic observation and clustering algorithms, eliminating the need for manual ground truth collection from endpoints. The database constructs itself by analyzing patterns in encrypted traffic without requiring external intervention or cooperation from monitored systems.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary passive observation system that collects network traffic data without directly interacting with endpoint processes. This intermediary approach allows database construction without requiring ground truth from difficult-to-access endpoints like IoT devices, mobile networks, and containers, thereby reducing collection complexity while maintaining identification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If ground truth collection from endpoints is performed, then fingerprint database accuracy is improved, but collection feasibility deteriorates in IoT, mobile, and container environments

Engineering Contradiction:
Improvefingerprint database reliabilityVSAvoiddata collection feasibility
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

Instead of collecting ground truth from endpoints upward to build the database, the patent inverts the approach by observing network traffic downward to infer endpoint behavior patterns. This inversion allows database construction in environments where traditional endpoint collection is infeasible, such as IoT devices, mobile networks, and containers, while maintaining database reliability through pattern-based inference.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent replaces the mechanical system of direct endpoint data collection with a passive network observation system that infers process information from encrypted traffic patterns. This substitution eliminates the need for direct endpoint access and ground truth collection, making database construction feasible in previously inaccessible environments while maintaining reliability through clustering algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If fingerprint prevalence database is constructed with extensive traffic observations, then process identification reliability is improved, but time and resources required for database construction increase

Engineering Contradiction:
Improveprocess identification reliabilityVSAvoiddatabase construction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by constructing the fingerprint prevalence database in advance through passive observation of network traffic. Once constructed, the database can be reused for multiple process identification tasks without requiring repeated extensive observations, thereby reducing time loss for subsequent operations while maintaining high reliability through the comprehensive initial data collection.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11936690B2Automatically generating a fingerprint prevalence database without ground truth
Publication Date: 2024.03.19 CISCO TECHNOLOGY INC
  • US11936690B2 patent drawing
  • US11936690B2 patent drawing
  • US11936690B2 patent drawing

AI summary

Techniques and mechanisms for using passively collected network data to automatically generate a fingerprint prevalence database without the need for endpoint ground truth. The process first clusters all observations with the same fingerprint string and similar source and destination context. The process then annotates each cluster with descriptive information and uses a rule-based system to derive an informative name from that descriptive information, e.g., “winnt amp client” or “cross-platform browser”. Optionally, the learned database may be augmented by a user to clarify custom process labels. Additionally, the generated database may be used to report the inferred processes in the same way as databases generated with endpoint ground truth.