Semantic Labeling System for Automated Data Feature Engineering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning tools face challenges in automating feature engineering and data labeling, particularly in handling diverse and noisy data from physical systems, due to the need for domain expertise and lack of standardized meta-descriptors, which hinders efficient data analytics and integration.

Innovation Solution

A method for learning semantic descriptions of data based on physical knowledge using processors, where physical knowledge data and semantic labels are associated with data sources to generate textual descriptors, enabling automated semantic labeling and guiding machine learning model predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine learning tools are used for data analytics, then computational power and data processing capability are improved, but the need for domain expertise and lack of standardized meta-descriptors increases complexity in automating feature engineering and data labeling

Engineering Contradiction:
Improvedata processing capabilityVSAvoidcomplexity in automating feature engineering
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary component - a semantic labeling system that uses ontologies and natural language processing - to bridge the gap between raw data and machine learning models. This intermediary automatically generates meta-descriptors and performs semantic labeling, eliminating the need for manual domain expertise while maintaining high data processing capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by pre-processing data to extract semantic information and generate standardized meta-descriptors before the main machine learning processing. This preliminary semantic analysis automates the feature engineering process, reducing complexity in subsequent data analytics stages.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If diverse data from multiple data sources is integrated, then data variety and analytical insights are improved, but the lack of standardized meta-descriptors and heterogeneous data formats increases processing complexity and time

Engineering Contradiction:
Improvedata integration capabilityVSAvoidprocessing time for heterogeneous data
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements a universal semantic labeling framework that can process diverse data formats from multiple sources through a common ontology-based approach. This multi-functional system automatically adapts to different data types (sensor data, text, images) and generates standardized meta-descriptors, enabling efficient integration of heterogeneous data without increasing processing time.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the parameter representation of diverse data by transforming various data formats into a unified semantic parameter space using ontologies. This parameter transformation allows heterogeneous data to be processed uniformly, reducing the time required for data integration while maintaining adaptability to multiple data sources.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated semantic labeling is implemented, then data labeling efficiency is improved, but the need for physical knowledge data and semantic ontologies increases system complexity

Engineering Contradiction:
Improvedata labeling efficiencyVSAvoidsystem complexity for semantic labeling
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a self-service semantic labeling system that automatically generates labels by extracting physical knowledge directly from the data and matching it with ontology concepts. The system serves itself by autonomously performing semantic analysis without requiring external manual intervention, thereby improving data labeling efficiency while managing system complexity through automation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where the semantic labeling results are continuously refined based on the extracted physical knowledge and ontology matching accuracy. This feedback loop improves data labeling efficiency over time while the system learns to manage its own complexity by identifying patterns in the data that reduce the need for complex ontology queries.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230252310A1Learning semantic description of data based on physical knowledge data
Publication Date: 2023.08.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230252310A1 patent drawing
  • US20230252310A1 patent drawing
  • US20230252310A1 patent drawing

AI summary

Embodiments for learning semantic description of data based on physical knowledge in a computing environment by a processor. Physical knowledge data and semantic labels associated with data from one or more data sources may be learned. Source attributes of the one or more data sources may be associated with one or more classes and concepts of a plurality of ontologies based on the physical knowledge data and the semantic labels to generate textual descriptors of the data.