Multimodal Table Extraction for Semantic Search in Unstructured Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning systems struggle with processing unstructured data formats such as images, health records, documents, and metadata, as these are not directly compatible with conventional machine learning models, leading to inefficiencies and reduced performance metrics.

Innovation Solution

A machine learning platform that uses source-agnostic models to preprocess unstructured data, optimizing it for machine learning models by segmenting large documents, applying convolutional neural networks (CNN), and utilizing ontologies to handle spelling irregularities, enabling multimodal data extraction and semantic searches across diverse data types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning models are used directly on unstructured data, then model simplicity is maintained, but processing accuracy and performance metrics deteriorate

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary preprocessing layer that converts unstructured data into structured formats suitable for machine learning models. This intermediary processing layer handles the transformation from raw unstructured data (images, documents, metadata) into structured representations, thereby improving model accuracy without requiring the models themselves to become more complex

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the data processing pipeline into distinct stages: data ingestion, preprocessing/standardization, and model processing. By segmenting the processing of different data types (images, text documents, tabular data, metadata) into separate modules, the system achieves high accuracy for each data type while maintaining overall system manageability

Inventive Principle:
Principle #1Segmentation

2Productivity

If unstructured data is processed without standardization, then processing speed is maintained, but data compatibility and extraction efficiency deteriorate

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoiddata processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary standardization actions to transform unstructured data into standardized formats before they are fed into machine learning models. This preliminary processing includes converting various data types into consistent representations, which enables efficient extraction and processing while reducing the time needed for subsequent model processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a universal preprocessing framework that handles multiple data types (images, documents, tabular data, metadata) through a common standardization approach. This multi-functional preprocessing system improves extraction efficiency across all data types while maintaining consistent processing time characteristics

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If source-specific machine learning models are used for different data types, then data processing accuracy is improved, but system adaptability and complexity increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidsource-agnostic capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal machine learning platform that can process multiple data types (images, documents, tabular data, metadata) through a common architecture. The system maintains high accuracy for each data type while providing source-agnostic capability, eliminating the need for separate models for each data source

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses data representations and feature extractsors that can be copied and applied across different data types. By creating standardized representations that can be reused for various purposes, the system achieves both accuracy and adaptability without requiring custom models for each data source

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250384301A1Multimodal table extraction and semantic search in a machine learning platform for structuring data in organizations
Publication Date: 2025.12.18 EXISERVICE HOLDINGS INC
  • US20250384301A1 patent drawing
  • US20250384301A1 patent drawing
  • US20250384301A1 patent drawing

AI summary

Systems, methods, and computer-readable media for computer-assisted output validation in machine learning/artificial intelligence platforms are disclosed. An application instance includes one or more machine learning models used to generate searchable data structures based on multimodal inputs.