Dataset Tag Normalization Ontology Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dataset tagging, categorization, validation, and indexing techniques face challenges due to differences in tagging conventions, leading to unreliable search results, categorization, and validation, especially as the number of datasets and sources increases, resulting in decreased utility and reliability of searchable indexes.

Innovation Solution

The processing of tags involves tokenization, expansion of abbreviations, and mapping based on alphanumeric characteristics, followed by searching an ontology to determine standardized tags for indexing datasets, and using both ontology-based and machine-learning models to determine categories for improved categorization and validation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If datasets are tagged using different tagging conventions from different sources, then the datasets can be collected from multiple sources, but the search reliability and categorization accuracy decrease

Engineering Contradiction:
Improveability to collect datasets from multiple sourcesVSAvoidsearch reliability and categorization accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary tag normalization layer that translates diverse tagging conventions into a unified tag schema. This intermediary mechanism allows datasets from multiple sources with different tagging conventions to be collected while maintaining search reliability and categorization accuracy through standardization.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by transforming tag attributes (such as case sensitivity, spacing, punctuation) into standardized forms. This parameter transformation enables consistent categorization and searching across datasets from different sources without sacrificing the ability to collect from multiple sources.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If the number of datasets and dataset sources increases, then the coverage and versatility of the system improves, but the reliability of search and categorization results decreases

Engineering Contradiction:
Improvenumber of datasets and sourcesVSAvoidsearch and categorization reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The tag normalization system acts as an intermediary that maintains reliability even as the quantity of datasets increases. By standardizing tags through the intermediary layer, the system can handle larger volumes of data from diverse sources without compromising search and categorization reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a universal tag schema that serves multiple functions: it standardizes tags from different sources, enables consistent searching, and maintains categorization accuracy. This universal approach allows the system to scale with increasing dataset quantity while preserving reliability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Stability of the object's composition

If tags are standardized using only ontology-based methods, then categorization consistency improves, but the system complexity and processing requirements increase

Engineering Contradiction:
Improvecategorization consistencyVSAvoidsystem complexity and processing requirements
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent merges ontology-based categorization with machine-learning-based tag normalization into a unified system. This combination achieves consistent categorization while managing system complexity by integrating complementary approaches rather than relying on a single complex method.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses a composite approach combining traditional ontology-based methods with modern machine-learning techniques. This composite material of methodologies leverages the strengths of both approaches to achieve consistent categorization with manageable processing requirements.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20240346077A1Determining data categorizations based on an ontology and a machine-learning model
Publication Date: 2024.10.17 CAPITAL ONE SERVICES LLC
  • US20240346077A1 patent drawing
  • US20240346077A1 patent drawing
  • US20240346077A1 patent drawing

AI summary

Aspects described herein may relate to methods, systems, and apparatuses that determine one or more categories associated with a dataset, or a portion thereof. The determination may be performed based on one or more tags associated with the dataset and/or a description associated with the dataset. Further, the determination may be performed by searching an ontology based on the one or more tags and/or the description. The determination may be performed by using a machine-learning model based on the one or more tags and/or the description. Once the one or more categories associated with the dataset are determined, the one or more categories may be used as a basis for modifying the dataset and/or validating the dataset.