Data Classification System for Disparate Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Organizations face challenges in identifying and classifying data stored across disparate sources to ensure compliance with regulations, policies, and procedures due to the lack of infrastructure support for data management systems designed for regulatory compliance.

Innovation Solution

A computer system maintains a storage state database to track data fields from various sources, allowing users to define classes and rules for classification, and provides graphical interfaces for interactive classification and visualization without accessing the actual data, enabling organizations to understand and visualize the types and locations of data stored.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in multiple disparate sources across different environments and infrastructures, then organizations can retain large volumes of data, but it becomes difficult to identify and classify the data for compliance monitoring

Engineering Contradiction:
Improvevolume of dataVSAvoiddifficulty of identifying and classifying data
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces an intermediary classification system that acts as a mediator between disparate data sources and compliance requirements. This system includes a classification hierarchy with multiple levels (e.g., data type, sensitivity, regulation category) and a set of rules that automatically classify data fields without requiring direct inspection of the actual data content. The intermediary layer aggregates information from multiple sources and presents a unified view for compliance monitoring.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the classification process into distinct components: a classification hierarchy structure, classification rules, and a classification engine. Each data field is segmented and evaluated against multiple classification criteria independently. The system divides the complex task of data classification into manageable segments that can be processed systematically across disparate data sources.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If data sources are designed to solve problems unrelated to regulations, then they can fulfill their primary functions, but they lack infrastructure support for regulatory compliance and data classification

Engineering Contradiction:
Improvefunctionality of data sourcesVSAvoidcomplexity of compliance infrastructure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal classification framework that can be applied across multiple types of data sources and infrastructures without requiring modification to the original data sources. The classification system is designed to work with relational databases, flat files, cloud storage, and other disparate formats through a unified interface. This multi-functional approach allows the same classification infrastructure to serve compliance, security, and data management needs across diverse environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If organizations manually identify and classify data across multiple sources, then they can ensure compliance accuracy, but it requires significant time and resources

Engineering Contradiction:
Improveaccuracy of data classificationVSAvoidtime required for classification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-defining classification hierarchies and rules before data classification is needed. The classification framework is established in advance with predefined categories, criteria, and mapping rules. When data needs to be classified, the system simply applies the pre-established rules to the data fields, eliminating the need for manual analysis and significantly reducing classification time while maintaining consistency and accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11100141B2Monitoring organization-wide state and classification of data stored in disparate data sources of an organization
Publication Date: 2021.08.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11100141B2 patent drawing
  • US11100141B2 patent drawing
  • US11100141B2 patent drawing

AI summary

A computer system receives data identifying data fields used by sources of data and infrastructure information. The computer system maintains a storage state database to track information about the data fields. The computer system provides a graphical user interface through which a user can create, view, modify, and delete data defining classes and rules associated with classes. The computer system applies the rules to classify the data fields. The computer system provides a graphical user interface through which the classification of the data fields can be visualized in the context of infrastructure information for the sources of data.