Machine Learning Document Organization with OCR for Text-Image Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of efficiently organizing and accessing diverse information across numerous electronic documents, including text, images, and videos, becomes increasingly difficult as data grows in volume and complexity, making it hard to extract valuable information quickly and collectively.

Innovation Solution

A system and method that utilizes machine learning models to segment textual information into chunks, generate numerical representations, associate images with text, and store this information in a memory, enabling quick retrieval and integration of relevant data through natural language processing and optical character recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in numerous electronic documents of different types, then the quantity and variety of information is increased, but the difficulty of collecting and processing relevant information increases

Engineering Contradiction:
Improvequantity of informationVSAvoidcomplexity of data organization
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent merges multiple electronic documents into a single organized structure by extracting text and images from various document types and consolidating them into unified data structures with standardized fields, making the data more accessible and easier to process while preserving the quantity and variety of information

Inventive Principle:
Principle #5Merging (Combining)

2Quantity of substance

If the number of electronic documents increases to accommodate business growth, then the richness of data is improved, but the ease of extracting valuable information deteriorates

Engineering Contradiction:
Improverichness of dataVSAvoidease of information extraction
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments the content of electronic documents into discrete, standardized elements including text chunks, images, and associated metadata fields. This segmentation allows individual pieces of information to be easily extracted, indexed, and retrieved without processing entire documents, thereby improving ease of information extraction while maintaining data richness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms unstructured document content into structured data with defined parameters and fields (e.g., text content, file type, creation date, associated images). This parameterization enables efficient filtering, searching, and extraction of specific information from large volumes of documents

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If text is segmented into numerous chunks for detailed processing, then the precision of information retrieval is improved, but the time and computational resources required increase

Engineering Contradiction:
Improveprecision of information retrievalVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary segmentation and organization of text into chunks during the initial document processing phase, creating indexed and structured data structures in advance. This preliminary action enables rapid retrieval and precise information extraction later without requiring re-processing of entire documents, thus reducing processing time while maintaining retrieval precision

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250258860A1System and method of organizing data
Publication Date: 2025.08.14 HONEYWELL INTERNATIONAL INC
  • US20250258860A1 patent drawing
  • US20250258860A1 patent drawing
  • US20250258860A1 patent drawing

AI summary

A system and a method of organizing data is described. The method comprises extracting a first textual information from electronic documents, and segmenting it into one or more chunks of sentences including at least one word. First numerical representations of the one or more chunks are generated using a machine learning model. Identity and association between the electronic documents, the one or more chunks, and the first numerical representations are stored in a memory. Images are extracted from the electronic documents and a second textual information is extracted from the images. Second numerical representations of keywords present in the second textual information are generated using the machine learning model. The first numerical representations are matched with the second numerical representations for determining an association of the images with the first textual information. The association of the images with the first textual information is also stored in the memory.