Data Classification System for Disparate Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face challenges in identifying and classifying data stored across disparate sources to ensure compliance with regulations, policies, and procedures due to the lack of infrastructure support for data management systems designed for regulatory compliance.
Innovation Solution
A computer system maintains a storage state database to track data fields from various sources, allowing users to define classes and rules for classification, and provides graphical interfaces for interactive classification and visualization without accessing the actual data, enabling organizations to understand and visualize the types and locations of data stored.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in multiple disparate sources across different environments and infrastructures, then organizations can retain large volumes of data, but it becomes difficult to identify and classify the data for compliance monitoring
Solution Approach 1:
The patent introduces an intermediary classification system that acts as a mediator between disparate data sources and compliance requirements. This system includes a classification hierarchy with multiple levels (e.g., data type, sensitivity, regulation category) and a set of rules that automatically classify data fields without requiring direct inspection of the actual data content. The intermediary layer aggregates information from multiple sources and presents a unified view for compliance monitoring.
Solution Approach 2:
The patent segments the classification process into distinct components: a classification hierarchy structure, classification rules, and a classification engine. Each data field is segmented and evaluated against multiple classification criteria independently. The system divides the complex task of data classification into manageable segments that can be processed systematically across disparate data sources.
2Adaptability or versatility
If data sources are designed to solve problems unrelated to regulations, then they can fulfill their primary functions, but they lack infrastructure support for regulatory compliance and data classification
Solution Approach 1:
The patent creates a universal classification framework that can be applied across multiple types of data sources and infrastructures without requiring modification to the original data sources. The classification system is designed to work with relational databases, flat files, cloud storage, and other disparate formats through a unified interface. This multi-functional approach allows the same classification infrastructure to serve compliance, security, and data management needs across diverse environments.
3Measurement precision
If organizations manually identify and classify data across multiple sources, then they can ensure compliance accuracy, but it requires significant time and resources
Solution Approach 1:
The patent implements preliminary action by pre-defining classification hierarchies and rules before data classification is needed. The classification framework is established in advance with predefined categories, criteria, and mapping rules. When data needs to be classified, the system simply applies the pre-established rules to the data fields, eliminating the need for manual analysis and significantly reducing classification time while maintaining consistency and accuracy.
Data Source
AI summary
A computer system receives data identifying data fields used by sources of data and infrastructure information. The computer system maintains a storage state database to track information about the data fields. The computer system provides a graphical user interface through which a user can create, view, modify, and delete data defining classes and rules associated with classes. The computer system applies the rules to classify the data fields. The computer system provides a graphical user interface through which the classification of the data fields can be visualized in the context of infrastructure information for the sources of data.


