Ontology Linking Service for Unstructured Text

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing business intelligence systems struggle to utilize unstructured alphanumeric data, such as textual notes, due to its disorganized format, varying spellings, acronyms, and lack of standardization, making it difficult to extract meaningful data for analytics.

Innovation Solution

A service that uses machine learning models to identify entities and relationships in unstructured text and links them to standardized concepts in an ontology, enabling the structured representation of unstructured data for business intelligence applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual tagging is used to structure unstructured data, then data can be organized with well-defined structure, but the process is completely impractical for large amounts of data and produces significant numbers of errors

Engineering Contradiction:
Improvedata structure qualityVSAvoidtagging speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical tagging process with an automated machine learning-based natural language processing system. The system uses trained models to automatically identify entities, extract relationships, and structure unstructured text data without human intervention, thereby achieving both high productivity and acceptable precision at scale.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables unstructured data to be automatically structured through self-service mechanisms where machine learning models autonomously perform entity recognition, relationship extraction, and data organization tasks that would otherwise require manual human effort, allowing the system to scale indefinitely without additional human resources.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated tagging software is used to process unstructured data, then processing speed increases, but the systems tend to introduce many errors and typically only work for specific use cases

Engineering Contradiction:
Improvedata processing speedVSAvoidtagging accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent employs preliminary action by training machine learning models on extensive annotated datasets before deployment. This pre-training phase allows the system to learn patterns, relationships, and contextual nuances in the data, enabling it to perform accurate entity recognition and relationship extraction when processing actual unstructured data at scale.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system achieves universality by designing a flexible machine learning framework that can adapt to multiple use cases and data types. Rather than creating specialized tagging systems for specific applications, the patent develops a general-purpose NLP system that can handle various unstructured data formats and domains through configurable models and parameters.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If unstructured data is used directly in business intelligence applications, then data volume is maintained, but the applications are unable to utilize the data due to lack of explicit data structure and schema

Engineering Contradiction:
Improvedata volumeVSAvoiddata usability
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent introduces an intermediary layer between raw unstructured data and business intelligence applications. This intermediary system uses machine learning to automatically extract structured information from unstructured text, transforming it into a format that BI applications can consume while preserving the original data volume and preventing information loss.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system applies segmentation by breaking down unstructured text data into discrete structured components such as entities, relationships, attributes, and metadata. This segmentation process transforms continuous unstructured text into discrete, queryable data elements that maintain the original information while becoming usable by standard business intelligence tools.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12242525B1Service architecture for ontology linking of unstructured text
Publication Date: 2025.03.04 AMAZON TECH INC
  • US12242525B1 patent drawing
  • US12242525B1 patent drawing
  • US12242525B1 patent drawing

AI summary

Techniques for ontology linking of unstructured text as a service are described. A service may receive a request to link unstructured text to a standardized ontology, and the service may segment and tokenize the unstructured text and send the result to multiple services implementing multiple deep machine learning models trained to identify particular entities and one or more relationships between entities. The service may perform a search of the standardized ontology to identify a set of similar candidates from the standardized ontology for the detected entities and the one or more relationships, and then rank the set of similar candidates from the standardized ontology according to their similarity to the detected entities within the unstructured text. The output from the service may include a result identifying a highest ranked candidate of the set of similar candidates from the standardized ontology for the detected entities within the unstructured text.