Automated Data Classification via Vector Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for data classification in multi-layer service-oriented platforms are inefficient, prone to human errors, and struggle to keep pace with the rapid growth of new applications and services, leading to compliance issues and misclassification of data objects.

Innovation Solution

A data classification system that retrieves data objects from a repository, parses them into text-based elements, converts them into vector data objects, and maps these objects to a trained data classification vector set to determine accurate labels, enabling automated and real-time classification and access control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data classification methods are used, then human review and judgment can be applied, but human errors occur and efficiency is low

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical classification processes with an automated machine learning system. The system uses trained classifiers, vector space models, and automated parsing to substitute human reviewers, thereby eliminating human errors while maintaining high classification accuracy and improving efficiency through automated batch processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service classification by automatically parsing data objects, generating vector representations, and applying classification labels without human intervention. The automated pipeline includes self-training capabilities where the system can learn from labeled examples and improve its classification performance independently.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If existing classification systems are used, then they can handle current data, but they struggle to keep pace with rapid growth of new applications and services

Engineering Contradiction:
Improvesystem adaptabilityVSAvoidcompliance reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The classification system is designed to be dynamic and adaptable to changing data landscapes. The system can be retrained with new labeled data as new applications and services are introduced, allowing it to adapt its classification models and vector spaces to maintain reliability and compliance with evolving data governance requirements.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary classification and labeling of data objects as they are created or imported into the data lake. By applying classification labels early in the data lifecycle, the system ensures compliance is maintained from the outset, and the preliminary structured data facilitates faster adaptation to new applications.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated classification is implemented, then efficiency improves, but complexity of the system increases

Engineering Contradiction:
Improveclassification throughputVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The complex automated classification system is segmented into distinct modular components: data parsing modules, vector generation modules, classification engine modules, and labeling modules. Each component handles a specific aspect of the classification process, making the overall system more manageable and easier to maintain while achieving high throughput through parallel processing of these segmented functions.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If comprehensive data parsing is performed, then classification accuracy improves, but processing time increases

Engineering Contradiction:
Improvelabeling accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial parsing and vector generation to data objects, focusing computational resources on the most critical fields and data elements that contribute most to classification accuracy. By selectively processing only the necessary portions of data objects rather than exhaustive parsing, the system achieves high labeling accuracy while minimizing processing time and computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11874937B2Apparatuses, methods, and computer program products for programmatically parsing, classifying, and labeling data objects
Publication Date: 2024.01.16 ATLASSIAN PTY LTD
  • US11874937B2 patent drawing
  • US11874937B2 patent drawing
  • US11874937B2 patent drawing

AI summary

Methods, apparatuses, or computer program products are disclosed providing for the dynamic data classification of data objects. Examples enable prediction of candidate data classification labels for data objects associated with one or more applications, services, or computing devices. Examples enable the assignment of one or more data classification labels to a data object for transmission to one or more computing devices. Examples enable the interactive and progressive application of machine learning techniques to data classification systems to assign data classification labels with probable certainty. Examples enable the tracking, monitoring, storage, sorting, and retrieval of labeled data objects. Examples provide for access control configuration of services to restrict or allow access to data objects based on data classifications and other service parameters.