Dataset Tag Normalization Ontology Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing dataset tagging, categorization, validation, and indexing techniques face challenges due to differences in tagging conventions, leading to unreliable search results, categorization, and validation, especially as the number of datasets and sources increases, resulting in decreased utility and reliability of searchable indexes.
Innovation Solution
The processing of tags involves tokenization, expansion of abbreviations, and mapping based on alphanumeric characteristics, followed by searching an ontology to determine standardized tags for indexing datasets, and using both ontology-based and machine-learning models to determine categories for improved categorization and validation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If datasets are tagged using different tagging conventions from different sources, then the datasets can be collected from multiple sources, but the search reliability and categorization accuracy decrease
Solution Approach 1:
The patent introduces an intermediary tag normalization layer that translates diverse tagging conventions into a unified tag schema. This intermediary mechanism allows datasets from multiple sources with different tagging conventions to be collected while maintaining search reliability and categorization accuracy through standardization.
Solution Approach 2:
The patent applies parameter changes by transforming tag attributes (such as case sensitivity, spacing, punctuation) into standardized forms. This parameter transformation enables consistent categorization and searching across datasets from different sources without sacrificing the ability to collect from multiple sources.
2Quantity of substance
If the number of datasets and dataset sources increases, then the coverage and versatility of the system improves, but the reliability of search and categorization results decreases
Solution Approach 1:
The tag normalization system acts as an intermediary that maintains reliability even as the quantity of datasets increases. By standardizing tags through the intermediary layer, the system can handle larger volumes of data from diverse sources without compromising search and categorization reliability.
Solution Approach 2:
The patent implements a universal tag schema that serves multiple functions: it standardizes tags from different sources, enables consistent searching, and maintains categorization accuracy. This universal approach allows the system to scale with increasing dataset quantity while preserving reliability.
3Stability of the object's composition
If tags are standardized using only ontology-based methods, then categorization consistency improves, but the system complexity and processing requirements increase
Solution Approach 1:
The patent merges ontology-based categorization with machine-learning-based tag normalization into a unified system. This combination achieves consistent categorization while managing system complexity by integrating complementary approaches rather than relying on a single complex method.
Solution Approach 2:
The system uses a composite approach combining traditional ontology-based methods with modern machine-learning techniques. This composite material of methodologies leverages the strengths of both approaches to achieve consistent categorization with manageable processing requirements.
Data Source
AI summary
Aspects described herein may relate to methods, systems, and apparatuses that determine one or more categories associated with a dataset, or a portion thereof. The determination may be performed based on one or more tags associated with the dataset and/or a description associated with the dataset. Further, the determination may be performed by searching an ontology based on the one or more tags and/or the description. The determination may be performed by using a machine-learning model based on the one or more tags and/or the description. Once the one or more categories associated with the dataset are determined, the one or more categories may be used as a basis for modifying the dataset and/or validating the dataset.


