Automatic Data Catalog Tagging for Analysts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data catalog generation methods are inadequate as they require manual tagging, which is not comprehensive, and are limited to industries with standardized data models, making it difficult for analysts without field data knowledge to select and use analysis data effectively.
Innovation Solution
A data catalog automatic generation system that uses a field data reception section and a data management section to extract relationships between objective and explanatory variables, and attach catalog tags based on set classification rules, enabling analysts with limited knowledge to select and analyze data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If manual tagging is performed by crowdsourcing, then a catalog can be generated, but the comprehensiveness is not satisfactory and leakage may occur
Solution Approach 1:
The system enables self-service by allowing the data itself to provide the tagging information through embedded metadata and machine-readable formats. The automated extraction process reads data characteristics directly from the source without requiring external human intervention, thus achieving both automation and reliability simultaneously
Solution Approach 2:
The manual mechanical process of crowdsourcing tagging is replaced with an automated electronic system that uses machine learning algorithms and pattern recognition to extract and assign tags automatically. This substitution eliminates human error while maintaining comprehensive coverage of all data elements
2Ease of operation
If a data model prescribed in an industry standard is used, then automatic conversion can be performed, but it can be used only in an industry in which a data model is prescribed and a catalog cannot be selected without sufficient knowledge
Solution Approach 1:
The system achieves universality by designing a flexible tagging framework that can adapt to multiple industries and data types without requiring industry-specific standardized models. The machine learning algorithms learn patterns across different domains, enabling the same system to generate catalogs for various industries while remaining easy to operate
Solution Approach 2:
The system changes parameters dynamically by adjusting tagging criteria and classification rules based on the characteristics of the input data. Rather than requiring fixed industry standards, the system learns and adapts its parameters automatically, making it both versatile across industries and easy to operate without expert knowledge
3Loss of information
If conventional catalog generation methods are used, then field data can be organized, but analysts without field data knowledge cannot effectively select and use analysis data
Solution Approach 1:
The system introduces an intermediary layer of automatically generated tags and metadata that bridge the gap between raw field data and user needs. These intermediaries encode organizational structure and relationships, allowing users to search and select data without needing to understand the underlying field data complexity
Solution Approach 2:
The data organization process serves itself by automatically generating comprehensive tags and relationships without human intervention. This self-organizing capability ensures high data organization quality while simultaneously making the data easily searchable and selectable for analysts regardless of their domain expertise
Data Source
AI summary
A technology is disclosed that makes it possible even for an analyst, who has poor knowledge relating to field data, to select and use analysis data in analysis. A data catalog automatic generation system that generates a catalog tag to be used to select analysis data from collected field data is configured such that, based on a set classification rule input, a relationship between an objective variable as an analysis perspective relating to field data and an explanatory variable or a causal relationship between a plurality of the explanatory variables is extracted, and based on a result of the extraction, a catalog tag of the objective variable and a catalog tag of the explanatory function are specified and attached.


