Anomaly Detection via Unsupervised Learning for Data Loss Prevention

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Rule-based systems for data loss prevention are inadequate as they cannot detect unknown activities, require time to create new rules, and become cumbersome with growing rule sets, limiting their effectiveness in identifying suspicious behavior.

Innovation Solution

A method that uses unsupervised machine learning to identify characteristic user behaviors from historical data, allowing for the detection of suspicious activities without pre-defined rules, by determining relative frequencies of user actions and comparing them to criteria for anomaly detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If rule-based systems are used for data loss prevention, then known suspicious activities can be detected, but unknown activities cannot be detected and the system becomes cumbersome with growing rule sets

Engineering Contradiction:
Improvedetection capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical rule-based system with an unsupervised machine learning system that automatically learns user behavior patterns from historical data. Instead of manually creating and maintaining rules, the system uses algorithms to identify characteristic behaviors and detect anomalies automatically, resolving the contradiction between detection capability and system complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-learning by automatically analyzing historical user behavior data to determine characteristic behaviors without human intervention. The unsupervised machine learning algorithm continuously adapts to new patterns, eliminating the need for manual rule creation and maintenance while improving detection capability over time

Inventive Principle:
Principle #25Self-service

2Measurement precision

If new rules are created to address new knowledge, then detection accuracy improves, but the process is not instantaneous and requires human intervention

Engineering Contradiction:
Improvedetection accuracyVSAvoidtime delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The unsupervised machine learning system performs self-learning by automatically analyzing historical data to identify new characteristic behaviors and update detection models without human intervention. This eliminates the time delay associated with manual rule creation while maintaining high detection accuracy through continuous automatic adaptation to new patterns

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If the number of rules grows over time, then coverage of suspicious activities increases, but maintenance difficulty increases

Engineering Contradiction:
ImprovecoverageVSAvoidmaintenance ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces the manual rule maintenance process with an automated unsupervised machine learning system that continuously learns from historical data. The system automatically identifies and adapts to new user behavior patterns, maintaining comprehensive coverage of suspicious activities while eliminating the need for manual rule maintenance as the system self-updates its detection models

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10192050B2Methods, systems, apparatus, and storage media for use in detecting anomalous behavior and/or in preventing data loss
Publication Date: 2019.01.29 BLUE RIDGE INNOVATIONS LLC
  • US10192050B2 patent drawing
  • US10192050B2 patent drawing
  • US10192050B2 patent drawing

AI summary

In one aspect, a method includes: receiving information defining a plurality of different actions that may be performed by users; receiving information indicating a relative frequency at which each of the different actions was performed by each of a plurality of users over each of one or more periods of time; determining a plurality of different characteristic behaviors based at least in part on the information indicating the relative frequency at which each of the different actions was performed by each of the plurality of users over each of one or more periods of time, wherein each one of the different characteristic behaviors defines a relative frequency of performance of each of the different actions; receiving information indicating a relative frequency at which each of the different actions was performed by a user over a period of time; and determining a representation of the relative frequency at which each of the different actions was performed by the user over the period of time as a weighted combination of the different characteristic behaviors each of which defines a relative frequency of performance of each of the different actions.