Neural Network Data Extraction Using N-gram Tokenization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document processing technologies face challenges in efficiently and accurately extracting data from semi-structured and unstructured documents due to variations in templates, formats, languages, and poor document quality, leading to inefficiencies and errors in data extraction.
Innovation Solution
A system and method for optimized training of a neural network model that generates a predetermined format of documents by extracting words with coordinates, generates N-grams based on threshold criteria, labels them with field names, and tokenizes words relative to named entities to improve data extraction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If template and rule-based document extraction processes are used, then data extraction can be performed with existing methods, but continuous training is required which makes the process time-consuming and inefficient
Solution Approach 1:
The patent replaces traditional rule-based and template-matching mechanical systems with a neural network-based intelligent system. The neural network model learns patterns from training data and automatically performs data extraction without requiring manual rule updates, thereby eliminating continuous training requirements and reducing processing time while maintaining accuracy.
Solution Approach 2:
The patent transforms the extraction system from a static rule-based approach to a dynamic neural network approach that adapts its parameters (weights and biases) through training. This allows the system to automatically adjust to different document formats and templates without manual reconfiguration, reducing both time loss and improving reliability.
2Ease of manufacture
If rule-based techniques with pre-populated dictionaries are used for data extraction, then extraction can be performed, but the dictionary needs constant updating which makes the process error-prone
Solution Approach 1:
The patent replaces the mechanical dictionary-update process with an intelligent neural network system that automatically learns and adapts to new terminology and formats. The neural network processes raw text and automatically identifies entities without requiring manual dictionary maintenance, thereby improving reliability while keeping the process simple.
3Productivity
If template matching techniques are used, then data extraction can be performed, but the techniques cannot determine relationships between text blocks and do not consider features at top and bottom of documents
Solution Approach 1:
The patent extends the extraction process from simple template matching to a multi-dimensional neural network approach that considers spatial relationships, sequential context, and hierarchical document structures. The neural network analyzes features at all positions in the document (top, bottom, middle) and determines relationships between text blocks through learned patterns, thereby improving measurement precision without sacrificing productivity.
4Adaptability or versatility
If conventional document processing methods are used, then processing can be performed, but new and different document types require template changes which makes processing challenging and error-prone
Solution Approach 1:
The patent creates a universal neural network-based extraction system that can handle multiple document types and formats through a single model. The neural network learns to adapt to different document structures, templates, and formats during training, eliminating the need for separate templates for each document type and reducing management complexity while improving adaptability.
Data Source
AI summary
A system and method for optimized training of a neural network model for data extraction is provided. The present invention provides for generating a pre-determined format type of input document by extracting words from input document along with coordinates corresponding to each word. Further, N-grams are generated by analyzing neighboring words associated with entity text present in predetermined format type of document based on threshold measurement criterion and combining extracted neighboring words in pre-defined order. Further, generated N-grams are compared with coordinates corresponding to words for labelling N-grams with field name. Further, each word in N-gram identified by the field name is tokenized in accordance with location of each of the words relative to named entity (NE) for assigning token marker. Lastly, neural network model is trained based on tokenized words in N-gram identified by token marker. The trained neural network model is implemented for extracting data from documents.


