Automated Data Classification via Vector Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for data classification in multi-layer service-oriented platforms are inefficient, prone to human errors, and struggle to keep pace with the rapid growth of new applications and services, leading to compliance issues and misclassification of data objects.
Innovation Solution
A data classification system that retrieves data objects from a repository, parses them into text-based elements, converts them into vector data objects, and maps these objects to a trained data classification vector set to determine accurate labels, enabling automated and real-time classification and access control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data classification methods are used, then human review and judgment can be applied, but human errors occur and efficiency is low
Solution Approach 1:
The patent replaces manual mechanical classification processes with an automated machine learning system. The system uses trained classifiers, vector space models, and automated parsing to substitute human reviewers, thereby eliminating human errors while maintaining high classification accuracy and improving efficiency through automated batch processing.
Solution Approach 2:
The system enables self-service classification by automatically parsing data objects, generating vector representations, and applying classification labels without human intervention. The automated pipeline includes self-training capabilities where the system can learn from labeled examples and improve its classification performance independently.
2Adaptability or versatility
If existing classification systems are used, then they can handle current data, but they struggle to keep pace with rapid growth of new applications and services
Solution Approach 1:
The classification system is designed to be dynamic and adaptable to changing data landscapes. The system can be retrained with new labeled data as new applications and services are introduced, allowing it to adapt its classification models and vector spaces to maintain reliability and compliance with evolving data governance requirements.
Solution Approach 2:
The system performs preliminary classification and labeling of data objects as they are created or imported into the data lake. By applying classification labels early in the data lifecycle, the system ensures compliance is maintained from the outset, and the preliminary structured data facilitates faster adaptation to new applications.
3Productivity
If automated classification is implemented, then efficiency improves, but complexity of the system increases
Solution Approach 1:
The complex automated classification system is segmented into distinct modular components: data parsing modules, vector generation modules, classification engine modules, and labeling modules. Each component handles a specific aspect of the classification process, making the overall system more manageable and easier to maintain while achieving high throughput through parallel processing of these segmented functions.
4Measurement precision
If comprehensive data parsing is performed, then classification accuracy improves, but processing time increases
Solution Approach 1:
The system applies partial parsing and vector generation to data objects, focusing computational resources on the most critical fields and data elements that contribute most to classification accuracy. By selectively processing only the necessary portions of data objects rather than exhaustive parsing, the system achieves high labeling accuracy while minimizing processing time and computational overhead.
Data Source
AI summary
Methods, apparatuses, or computer program products are disclosed providing for the dynamic data classification of data objects. Examples enable prediction of candidate data classification labels for data objects associated with one or more applications, services, or computing devices. Examples enable the assignment of one or more data classification labels to a data object for transmission to one or more computing devices. Examples enable the interactive and progressive application of machine learning techniques to data classification systems to assign data classification labels with probable certainty. Examples enable the tracking, monitoring, storage, sorting, and retrieval of labeled data objects. Examples provide for access control configuration of services to restrict or allow access to data objects based on data classifications and other service parameters.


