Automated Unstructured Text Processing via Embedding Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information processing systems face inefficiencies in managing unstructured text data, particularly in large volumes, as they require manual screening and rule customization, making the processing tedious and time-consuming.

Innovation Solution

An automated system that selects and processes unstructured text data in paired data fields by determining embeddings, identifying similar data fields, and providing recommendations based on syntactic differences, reducing the need for manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual screening and rule customization are used to process unstructured text data, then processing accuracy can be maintained, but processing time and effort increase significantly

Engineering Contradiction:
Improveprocessing accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical processing (human reviewers reading and analyzing unstructured text) with an automated electronic system that uses natural language processing, machine learning models, and text analytics to process unstructured data, thereby eliminating the time-consuming manual screening while maintaining processing accuracy through algorithmic analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service processing where the automated platform independently analyzes unstructured text data, generates insights, and produces outputs without requiring continuous human intervention or manual rule customization, allowing the system to serve itself in processing large volumes of unstructured data efficiently

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If manual customization of processing rules is implemented, then processing can be tailored to specific needs, but device complexity and maintenance burden increase

Engineering Contradiction:
Improveprocessing customizationVSAvoidrule set complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic processing rules that automatically adapt to different data types and contexts through machine learning models, eliminating the need for static manual rule customization. The system dynamically adjusts its analysis approach based on the characteristics of the unstructured data being processed, providing versatility without complexity

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The automated processing platform provides universal functionality that can handle multiple types of unstructured data (text, documents, media) and various processing tasks (analysis, extraction, classification) through a single integrated system, eliminating the need for separate customized rule sets for different scenarios

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If automated processing is implemented, then processing speed and productivity improve, but the system requires sophisticated algorithms and computational resources

Engineering Contradiction:
Improvedata processing throughputVSAvoidalgorithm complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the automated processing system into distinct functional modules including natural language processing components, machine learning models, text analytics engines, and output generation systems. Each module handles specific aspects of unstructured data processing, allowing high productivity through parallel processing while managing complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary processing layers including natural language understanding components and text analytics bridges that translate unstructured data into structured formats, enabling automated high-speed processing while managing algorithmic complexity through intermediate representation layers

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12026182B2Automated processing of unstructured text data in paired data fields of a document
Publication Date: 2024.07.02 ARCHER TECHNOLOGIES LLC
  • US12026182B2 patent drawing
  • US12026182B2 patent drawing
  • US12026182B2 patent drawing

AI summary

An apparatus comprises a processing device configured to select a first data field of a first type that is associated with a second data field of a second type in a document, to determine an embedding of terms of unstructured text data in the first data field and to identify a subset of paired data fields from an unstructured text database based at least in part on metrics characterizing similarity between (i) the embedding of terms in the first data field and (ii) embeddings of terms in data fields of the first type in the paired data fields. The processing device is further configured to determine syntactic differences between the unstructured text data in the first data field and the identified subset of paired data fields, and to provide recommendations for unstructured text data to fill the second data field in the document based on the syntactic differences.