Source Side Data Classification via OS Call Interception
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data classification processes are laborious and inefficient, relying on IT administrators who lack knowledge of data importance and priority, leading to suboptimal classification and increased security vulnerabilities, especially with unstructured data like emails and documents.
Innovation Solution
Implementing source side classification methods that intercept operating system calls to classify files at the endpoint device, leveraging user knowledge and idle periods to improve efficiency and accuracy, with a system that includes a classification engine and user rating engine to prioritize and manage file classification based on user history and geolocation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If IT administrators perform remote data classification, then data classification can be centralized and managed, but the process becomes manually laborious and inefficient
Solution Approach 1:
The system enables endpoint computing devices to autonomously identify and classify their own files by intercepting operating system calls and analyzing file access patterns, eliminating the need for manual classification by IT administrators. The endpoint device self-manages the classification process by sending file lists to the server and receiving classification requests autonomously.
Solution Approach 2:
The system performs preliminary classification actions by intercepting operating system calls before files are fully accessed or processed. The endpoint computing device proactively identifies files being accessed or recently accessed and prepares classification requests in advance, rather than waiting for manual intervention after the fact.
2Measurement precision
If IT administrators remotely classify data, then centralized control is achieved, but IT administrators lack knowledge of data importance and priority
Solution Approach 1:
The system incorporates user feedback mechanisms where active users can review and confirm file classifications. The server computing device receives file lists from endpoint devices, processes classification requests, and returns confirmation, creating a feedback loop that ensures classifications align with actual data importance and business context.
Solution Approach 2:
The system divides the classification responsibility into segments: the endpoint computing device identifies files and sends lists to the server, the server processes classification requests using benchmarking data, and active users provide context-specific approvals. This segmented approach combines automated efficiency with human expertise on data importance.
3Productivity
If source side classification is implemented at endpoint devices, then classification speed improves, but computing resources at the endpoint are consumed
Solution Approach 1:
The system performs partial classification actions at the endpoint by intercepting only the operating system calls related to file access and sending file lists to the server for processing. The endpoint device does not perform complete classification locally but rather initiates and manages the classification process, balancing local resource usage with remote processing power.
Solution Approach 2:
The server computing device acts as an intermediary between the endpoint computing device and the classification process. The endpoint sends file lists to the server, which processes the classification requests and returns confirmations, distributing the computational burden and preventing excessive resource consumption at the endpoint device.
4Measurement precision
If active users perform data classification, then classification accuracy improves through user knowledge, but the process becomes more complex
Solution Approach 1:
The system creates a universal platform that serves multiple functions: the endpoint computing device both identifies files and manages the classification request process, the server both receives file lists and processes classifications, and active users both provide context and approve classifications. This multi-functional design simplifies the overall system architecture despite the involvement of multiple stakeholders.
Data Source
AI summary
Disclosed herein are methods, systems, and processes for source side classification of five and active data. Operating system calls associated with files being accessed or files recently accessed by an endpoint computing device are intercepted. A list including the files is generated and sent to a server computing device. A confirmation is received that a request to classify the files has been received from the server computing device.


