Backend Data Classifier for Storage Device DLP
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data loss prevention (DLP) systems in computer networks face challenges in accurately and efficiently identifying and classifying stored data, particularly due to indirect access methods and lack of direct interaction with storage devices.
Innovation Solution
Implementing a backend data classification system that directly interacts with storage devices using a processing platform with a file analyzer and assignment module, which compares file characteristics with a file history database to assign classifications and label files with metadata, enabling more precise data loss prevention operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If indirect access methods (NFS, CIFS) are used to scan files, then the system can access stored data, but the accuracy and efficiency of data classification is insufficient
Solution Approach 1:
The patent introduces a backend data classifier as an intermediary component that directly interacts with storage devices. This mediator performs file inspection and classification by accessing storage devices through specialized interfaces, bypassing the limitations of conventional indirect access methods like NFS or CIFS. The intermediary enables precise data classification while managing the complexity of direct storage access.
Solution Approach 2:
The patent replaces the mechanical/protocol-based indirect access system (NFS, CIFS) with a direct access mechanism. Instead of relying on file system protocols that traverse directory structures, the system uses direct storage device interfaces to access and inspect files, thereby improving classification accuracy and efficiency while reducing the complexity of protocol handling.
2Productivity
If direct interaction with storage devices is implemented, then data classification efficiency is improved, but system complexity increases
Solution Approach 1:
The patent segments the data classification function into a dedicated backend data classifier component that operates independently from conventional file systems. This segmentation allows direct storage device interaction for efficient classification while isolating the complexity management to a specialized module, preventing it from propagating throughout the entire system.
Solution Approach 2:
The backend data classifier implements self-service capabilities by autonomously performing file inspection, characteristic comparison with historical data, and classification assignment. This self-service mechanism eliminates the need for complex external coordination while maintaining direct storage access, thereby improving efficiency without proportionally increasing system complexity.
3Reliability
If file inspection and classification is performed continuously, then data loss prevention accuracy is improved, but production impacts increase
Solution Approach 1:
The patent implements periodic action through scheduled classification operations that run at predetermined intervals rather than continuously. The backend data classifier can be configured to perform file inspection and classification at specific times, allowing DLP accuracy to be maintained through regular updates while minimizing disruption to production operations during peak periods.
Solution Approach 2:
The system employs dynamic scheduling capabilities that allow the classification operations to be adjusted based on production conditions. The backend data classifier can modify its operation timing and frequency dynamically, running more extensively during low-impact periods and reducing activity during critical production windows, thereby balancing DLP accuracy with production continuity.
Data Source
AI summary
An apparatus in one embodiment comprises a processing platform that includes one or more processing devices each comprising a processor coupled to a memory. The processing platform is associated with at least one storage device. The processing platform comprises a backend data classifier configured for communication with a data loss prevention system. The backend data classifier comprises a file analyzer configured to compare characteristics relating to current states of respective files stored in the storage device with information stored in a file history database, and an assignment module configured to assign classifications to respective ones of the files stored in the storage device based at least in part on comparison results from the file analyzer. The data loss prevention system is configured to perform different data loss prevention operations on different ones of the files stored in the storage device based at least in part on their respective assigned classifications.

