Data-Aware Anomaly Detection for Distributed Sensitive Data Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increased complexity and likelihood of data breaches across distributed computing environments, particularly due to the fragmentation of sensitive data and the ability for anonymous access, hinder quick and accurate identification of unauthorized data access.
Innovation Solution
A data-aware anomaly detection system that combines knowledge of sensitive data locations with activity logs from various sources, normalizes and enriches these logs, and performs anomaly detection at a data-type level to identify anomalous access patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is distributed across multiple sources in computing environments, then data availability and accessibility are improved, but the complexity of detecting and responding to data breaches increases
Solution Approach 1:
The patent combines activity logs from multiple distributed data sources into a unified analysis framework. By merging log data across different sources and enriching it with data classification information, the system creates a consolidated view that enables centralized anomaly detection despite the distributed nature of the underlying data infrastructure.
Solution Approach 2:
The patent introduces an intermediary resource management system that acts as a mediator between distributed data sources and anomaly detection capabilities. This intermediary collects, normalizes, and enriches log data from various sources, then feeds it to anomaly detection algorithms, thereby simplifying the detection process without requiring changes to the distributed data architecture.
2Speed
If traditional anomaly detection methods are used without data awareness, then detection speed is maintained, but detection precision and accuracy deteriorate
Solution Approach 1:
The patent performs preliminary actions by pre-classifying data and pre-enriching log data with data location and sensitivity information before anomaly detection occurs. This preliminary enrichment ensures that when anomaly detection runs, it has immediate access to contextual information about data types and locations, improving precision without adding computational overhead during the actual detection process.
Solution Approach 2:
The patent replaces traditional mechanical log analysis methods with data-aware anomaly detection that leverages enriched log data containing data classification information. By substituting basic pattern matching with intelligence-driven analysis that understands data context, the system achieves higher detection precision while maintaining operational speed.
3Measurement precision
If logs from multiple data sources are collected and analyzed, then detection accuracy is improved, but processing complexity and computational overhead increase
Solution Approach 1:
The patent changes key parameters by normalizing log data from different sources into a unified format and enriching it with standardized data classification attributes. By transforming heterogeneous log data into a consistent structure with standardized fields for data type, location, and sensitivity, the system simplifies processing complexity while maintaining the ability to analyze multiple data sources for improved detection accuracy.
4Reliability
If data classification and log enrichment are performed, then anomaly detection capability is improved, but processing time and computational resources increase
Solution Approach 1:
The patent performs data classification and log enrichment as preliminary actions before anomaly detection is needed. By pre-processing and enriching log data with data context information in advance, the system eliminates the need for time-consuming processing during actual anomaly detection events, thereby improving detection capability while minimizing additional processing time overhead.
Data Source
AI summary
Methods, systems, and devices for data and resource management are described. A location of a type of data that is present in different locations across multiple data sources of a computing environment may be identified. Logs from multiple data sources may be obtained, where the logs may capture activity information with respective data sources. The logs may be converted to a normalized format to obtain normalized logs. Based on determining the location of the type of data, the normalized logs may be associated with the type of data to obtain supplemented activity logs, which may be used to identify anomalous activity associated with the type of data being accessed in the computing environment.


