Annotation System for Extracting Attributes from Unstructured Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual entry of product information into electronic inventory systems is labor-intensive and prone to errors due to unstructured and unformatted product descriptions, making it difficult to ensure accuracy and proper cataloging, especially when dealing with thousands of products.
Innovation Solution
A computing system that extracts attributes from unstructured description strings and maps them to defined columns in a database using an annotation module, inference module, and structure module, transforming unstructured data into structured entries for accurate and organized product information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual entry of product information is used, then product data can be entered into the database, but the process is labor-intensive and prone to errors
Solution Approach 1:
The patent replaces the manual mechanical data entry process with an automated computer vision system that uses image recognition and natural language processing to extract product attributes from images and descriptions, thereby improving both productivity and reliability
Solution Approach 2:
The system enables self-service by automatically processing product information through multi-stage extraction pipelines, validation rules, and error correction mechanisms without requiring manual intervention for each product entry
2Productivity
If unstructured product descriptions are copied into the database, then data entry is faster, but the information contains grammatical errors, spelling errors, and other inaccuracies
Solution Approach 1:
The system performs preliminary processing of unstructured product descriptions through multiple extraction stages, validation rules, and error correction before final database insertion, ensuring accuracy is improved before data is committed to the database
Solution Approach 2:
The patent introduces an intermediary processing layer between the unstructured product descriptions and the database, including extraction modules, validation rules, and correction mechanisms that clean and structure the data before storage
3Reliability
If manual review and entry of product attributes is performed, then accurate data can be entered into separate database columns, but the process is time-consuming
Solution Approach 1:
The patent replaces manual review and entry of product attributes with automated extraction systems that use image recognition, natural language processing, and validation rules to accurately populate database columns without manual intervention
Solution Approach 2:
The system segments the product information processing into distinct automated stages including attribute extraction, validation, error correction, and database insertion, allowing each stage to be optimized independently for both accuracy and speed
Data Source
AI summary
Systems, methods, and other embodiments associated with extracting attributes from electronic data structures are described. In one embodiment, a method includes correlating tokens from description strings with defined attributes in an electronic inventory database by identifying which of the defined attributes match the tokens to link the tokens with columns of the database associated with the defined attributes. The method includes iteratively updating annotation strings for unidentified ones of the tokens by generating suggested matches for the unidentified tokens according to known correlations between identified tokens and the defined attributes using a conditional random fields model. The method also includes populating the database using the identified tokens from the description strings according to the annotation strings by automatically storing the tokens from the description strings into the columns as identified by the annotation strings.


