User Behavior Vector Generation for Malware Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional signature-based security algorithms fail to detect novel or polymorphic malware and advanced persistent network threats, and learning algorithms struggle with accurately representing user behavior due to limitations in vector representations and the need for extensive, costly labeling processes.

Innovation Solution

Generating a single behavioral vector representative of user behavior from network telemetry data, which transforms complex user traffic structures into a compact vector form without requiring time-intensive labeling, enabling improved classification and detection of infected users and groups with similar behavior.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If signature-based security algorithms are used to detect network threats, then known threats can be identified through byte sequence comparison, but new threats, polymorphic malware, and zero-day attacks cannot be detected

Engineering Contradiction:
Improvethreat detection accuracyVSAvoidcapability to detect new and polymorphic threats
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the detection approach by changing the parameter representation from fixed byte sequences to dynamic vector representations of user behavior. Instead of comparing static signatures, the system converts network traffic into vectors that capture behavioral patterns, enabling detection of threats that exhibit different byte sequences but maintain similar behavioral characteristics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical signature-matching system with a learning-based vector representation system. Rather than directly comparing byte sequences, the system uses learning algorithms to process and represent user behavior as vectors, substituting the rigid mechanical comparison process with a more flexible computational approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If traditional vector representations are used in learning algorithms, then computational processing is simplified, but complete user traffic structure and behavioral complexities cannot be represented

Engineering Contradiction:
Improvecomputational processing simplicityVSAvoiduser traffic structure and behavioral complexity
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies nesting by organizing user traffic data into hierarchical structures where multiple levels of traffic patterns are nested within vector representations. The vector representation captures nested patterns of user behavior across different time scales and traffic types, preserving structural information while maintaining computational tractability.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent transforms multi-dimensional user traffic structures into vector representations by adding temporal and behavioral dimensions. Instead of losing structural information during dimensionality reduction, the system incorporates temporal patterns and behavioral contexts as additional dimensions in the vector space, preserving complexity while enabling efficient processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If learning algorithms are trained with extensively labeled training data, then detection accuracy may improve, but training becomes prohibitively expensive and time intensive

Engineering Contradiction:
Improvedetection accuracyVSAvoidtraining time and cost
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent enables the system to partially label training data automatically through the vector representation process itself. The learning algorithm processes raw traffic data and generates vector representations that inherently contain labeled information about behavioral patterns, reducing the need for manual annotation while maintaining training effectiveness.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary transformation of raw traffic data into vector representations before the actual training process. This preliminary action organizes and structures the data in advance, making subsequent training more efficient and reducing the time required for both data preparation and model training.

Inventive Principle:
Principle #10Preliminary action

4Ease of manufacture

If samples are labeled without context in training data, then labeling process is simplified, but labels become improper or unreliable

Engineering Contradiction:
Improvelabeling process simplicityVSAvoidlabel accuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent incorporates contextual feedback into the labeling process by using the vector representation of user behavior to inform label assignment. Rather than labeling samples in isolation, the system uses the behavioral context captured in vector form to guide labeling decisions, ensuring that labels reflect the actual behavioral patterns while maintaining process efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11271954B2Generating a vector representative of user behavior in a network
Publication Date: 2022.03.08 CISCO TECHNOLOGY INC
  • US11271954B2 patent drawing
  • US11271954B2 patent drawing
  • US11271954B2 patent drawing

AI summary

Presented herein are techniques for classifying devices as being infected with malware based on learned indicators of compromise. A method includes receiving, at a security analysis device, a set of feature vectors extracted from one or more flows of traffic to domains for a given user in a network during a period of time. The security analysis device analyzes the feature vectors included in the set of feature vectors with a set of operators to generate a set of per-flow vectors for the given user. Based on the set of per-flow vectors for the user, the security analysis device generates a single behavioral vector representative of the given user. The security analysis device classifies a computing device associated with the given user based on the single behavioral vector and at least one of known information or other behavioral vectors for other users.