Techniques for targeted data extraction from unstructured sets of documents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual data review processes for identifying compromised information in cybersecurity breaches are inefficient, slow, and prone to errors due to the increasing complexity and volume of modern data ecosystems, particularly in unstructured data formats.
Innovation Solution
A computer-implemented method and system for targeted data extraction using a graphical user interface (GUI) that allows users to select and preserve visual boxes across multiple pages, combined with optical character recognition (OCR) techniques to extract relevant data from unstructured documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual data review processes are used to identify compromised information in unstructured documents, then human judgment and flexibility are maintained, but processing speed and efficiency deteriorate significantly
Solution Approach 1:
The patent introduces an intermediary system comprising OCR engines, natural language processing models, and machine learning classifiers that act as a bridge between unstructured document data and security analysis requirements. This intermediary automatically extracts and structures relevant information from unstructured documents, enabling both high-speed processing and maintained accuracy through multiple verification layers.
Solution Approach 2:
The patent segments the document review process into distinct automated stages: optical character recognition for text extraction, natural language processing for information structuring, machine learning classification for compromise identification, and verification stages. This segmentation allows parallel processing of multiple documents while maintaining comprehensive analysis through specialized sub-processes.
2Measurement precision
If manual review of unstructured data is performed, then contextual understanding can be applied, but processing time and resource consumption increase
Solution Approach 1:
The patent transforms unstructured document data into structured parameters through automated extraction processes. By converting free-text information into standardized data fields and categories, the system enables efficient querying and analysis while maintaining the contextual relationships between different pieces of information through structured data modeling.
Solution Approach 2:
The patent performs preliminary automated extraction and structuring of information from unstructured documents before human review or detailed analysis. OCR and NLP processes pre-process documents to identify and structure key information elements in advance, reducing the time required for subsequent manual verification or deep analysis.
3Reliability
If comprehensive review of all documents is conducted manually, then thoroughness is achieved, but scalability deteriorates with increasing data volume
Solution Approach 1:
The patent implements self-service automated systems that perform document review, extraction, and classification without requiring proportional increases in human resources. The machine learning models continuously process documents autonomously, adapting to new data types and formats while maintaining consistent review thoroughness across expanding document volumes through automated quality assurance mechanisms.
Data Source
AI summary
A computer-implemented method performed by one or more processors for targeted data extraction includes causing a display of a first view of a graphical user interface (GUI) displaying a plurality of graphical elements for selection; detecting a first input on a graphical user interface (GUI) selecting a first page view of a document from a plurality of documents; detecting a second input identifying an area of the first page view of the document; determining coordinates of the area of the first page view; generating a visual box based on the coordinates; causing a display of the visual box in the document on the GUI; and preserving the visual box across an additional page that is one of remaining pages of the document or pages of remaining documents of the plurality of documents such that the display of the visual box persists in the additional page on the GUI.


