Annotation System for Extracting Attributes from Unstructured Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual entry of product information into electronic inventory systems is labor-intensive and prone to errors due to unstructured and unformatted product descriptions, making it difficult to ensure accuracy and proper cataloging, especially when dealing with thousands of products.

Innovation Solution

A computing system that extracts attributes from unstructured description strings and maps them to defined columns in a database using an annotation module, inference module, and structure module, transforming unstructured data into structured entries for accurate and organized product information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual entry of product information is used, then product data can be entered into the database, but the process is labor-intensive and prone to errors

Engineering Contradiction:
Improveproduct data entry efficiencyVSAvoiddata accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces the manual mechanical data entry process with an automated computer vision system that uses image recognition and natural language processing to extract product attributes from images and descriptions, thereby improving both productivity and reliability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically processing product information through multi-stage extraction pipelines, validation rules, and error correction mechanisms without requiring manual intervention for each product entry

Inventive Principle:
Principle #25Self-service

2Productivity

If unstructured product descriptions are copied into the database, then data entry is faster, but the information contains grammatical errors, spelling errors, and other inaccuracies

Engineering Contradiction:
Improvedata entry speedVSAvoidproduct information accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs preliminary processing of unstructured product descriptions through multiple extraction stages, validation rules, and error correction before final database insertion, ensuring accuracy is improved before data is committed to the database

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary processing layer between the unstructured product descriptions and the database, including extraction modules, validation rules, and correction mechanisms that clean and structure the data before storage

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If manual review and entry of product attributes is performed, then accurate data can be entered into separate database columns, but the process is time-consuming

Engineering Contradiction:
Improvedata cataloging accuracyVSAvoidproduct information processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual review and entry of product attributes with automated extraction systems that use image recognition, natural language processing, and validation rules to accurately populate database columns without manual intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system segments the product information processing into distinct automated stages including attribute extraction, validation, error correction, and database insertion, allowing each stage to be optimized independently for both accuracy and speed

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10628403B2Annotation system for extracting attributes from electronic data structures
Publication Date: 2020.04.21 ORACLE INT CORP
  • US10628403B2 patent drawing
  • US10628403B2 patent drawing
  • US10628403B2 patent drawing

AI summary

Systems, methods, and other embodiments associated with extracting attributes from electronic data structures are described. In one embodiment, a method includes correlating tokens from description strings with defined attributes in an electronic inventory database by identifying which of the defined attributes match the tokens to link the tokens with columns of the database associated with the defined attributes. The method includes iteratively updating annotation strings for unidentified ones of the tokens by generating suggested matches for the unidentified tokens according to known correlations between identified tokens and the defined attributes using a conditional random fields model. The method also includes populating the database using the identified tokens from the description strings according to the annotation strings by automatically storing the tokens from the description strings into the columns as identified by the annotation strings.