Dynamic Data Class Generation via Dimension Score Thresholding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data governance systems struggle to classify data assets when no matching data class exists in the system, leading to inefficiencies in data classification processing, particularly when analyzing large datasets with unique data classes not currently defined.
Innovation Solution
A computer-implemented method generates a dimension score for each dimension of a data asset's column attributes through static reference data analysis, determining if the total score exceeds a minimum threshold to identify new static reference data, thereby dynamically creating new data classes that accelerate data classification and enrich the data governance system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the system uses predefined data classes for classification, then classification speed is improved, but the system cannot handle data assets with unique data classes not currently defined
Solution Approach 1:
The system dynamically generates new data classes based on analyzed data characteristics when predefined classes are insufficient. This allows the classification system to adapt to unique data assets while maintaining efficient predefined class matching for standard cases, resolving the contradiction between speed and adaptability.
Solution Approach 2:
The system automatically analyzes data assets and generates new data classes without requiring manual intervention. This self-service capability enables the system to handle unique data classes independently, maintaining both classification speed and adaptability.
2Adaptability or versatility
If the system manually creates new data classes for unique data assets, then adaptability is improved, but classification processing efficiency deteriorates
Solution Approach 1:
The system automatically analyzes data assets and generates new data classes without requiring manual intervention. This self-service capability enables the system to handle unique data classes independently, maintaining both classification speed and adaptability.
Solution Approach 2:
The system performs preliminary analysis of data assets to identify unique characteristics before classification. This preliminary action enables automatic generation of appropriate data classes, avoiding manual intervention and maintaining processing efficiency while improving adaptability.
3Measurement precision
If the system performs comprehensive data analysis to identify new data classes, then data classification accuracy is improved, but processing time increases
Solution Approach 1:
The system performs partial analysis focused on key dimensions sufficient to identify new data classes without conducting exhaustive comprehensive analysis. This approach achieves adequate classification accuracy while minimizing processing time by avoiding excessive analysis.
Solution Approach 2:
The system segments the data analysis process into specific dimensional evaluations rather than performing comprehensive analysis. This segmentation enables identification of new data classes through focused analysis of key attributes, improving processing efficiency while maintaining classification accuracy.
Data Source
AI summary
New data class generation is provided. A dimension score is generated for each respective dimension of a plurality of predefined dimensions as relating to column attributes of a data asset while performing a static reference data analysis of the data asset. The dimension score of each respective dimension is added together to obtain a total dimension score for the data asset. It is determined whether the total dimension score of the data asset is greater than a predefined minimum dimension score threshold level. The data asset is identified as new static reference data in response to determining that the total dimension score of the data asset is greater than the predefined minimum dimension score threshold level. A new data class is generated based on the new static reference data.


