AI Workflow Platform for Multi-Format Data Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI platforms face challenges in handling the complexity of multiple data formats, requiring significant time, manpower, and expertise for data preparation, visualization, labeling, and machine learning tasks, especially with the scarcity of tools for end-to-end AI model development and deployment in enterprises.

Innovation Solution

A comprehensive AI workflow platform that receives and indexes structured, semi-structured, and unstructured data, automatically schedules and uploads it to a database, visualizes and cleanses the data, labels and annotates it, and builds AI models based on user input and predefined rules, enabling seamless integration and management of AI models across various formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple data formats (structured, semi-structured, unstructured) are handled manually through separate tools and processes, then data processing can be performed, but the complexity of the system increases and requires significant time, manpower, and expertise

Engineering Contradiction:
Improveability to handle multiple data formatsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal data processing platform that can handle multiple data formats (structured, semi-structured, and unstructured data) through a single integrated system. The platform uses common processing pipelines, visualization tools, and machine learning algorithms that work across all data types, eliminating the need for separate specialized tools for each format and thereby reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If data preparation, cleaning, and integration are performed manually by data scientists, then data quality can be ensured, but 80-90% of time is spent on these tasks

Engineering Contradiction:
Improvedata qualityVSAvoidtime for data preparation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements automated self-service mechanisms for data preparation, cleaning, and integration. The system automatically performs data validation, cleaning, and transformation operations without requiring manual intervention from data scientists. This includes automated detection of data quality issues, automatic correction of common errors, and self-organizing of data into appropriate formats, thereby ensuring data quality while dramatically reducing the time spent on preparation tasks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-processing and pre-validating data as it enters the system. Data is automatically cleaned, transformed, and validated before reaching the analysis stage, preparing it in advance for machine learning operations. This preliminary data preparation reduces the need for extensive manual cleaning and integration work later in the process.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If deep learning models are used for image and text processing, then accuracy improves, but large quantities of labeled training data are required

Engineering Contradiction:
Improvemodel accuracyVSAvoidamount of training data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses semi-structured data as an intermediary representation that bridges unstructured data (images, text) and structured data formats. By converting unstructured data into semi-structured representations with embedded metadata and relationships, the system enables machine learning models to process complex data types without requiring vast amounts of manually labeled training data, thus maintaining accuracy while reducing data requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If specialized knowledge in specific fields (Computer Vision, NLP, time-series) is required for data science tasks, then task performance improves, but the barrier to entry increases and requires multiple software products and experts

Engineering Contradiction:
Improvetask performanceVSAvoidease of use
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent segments specialized data science tasks into modular, reusable components that can be independently developed and combined. Each data type (images, text, time-series) has dedicated processing modules with specialized algorithms, but these modules expose standardized interfaces that can be orchestrated through a unified platform. This allows specialized functionality to be maintained while providing an easier-to-use integrated system that reduces the need for multiple separate software products and deep specialized knowledge across all areas.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12079737B1Data-mining and AI workflow platform for structured and unstructured data
Publication Date: 2024.09.03 THINKTRENDS LLC
  • US12079737B1 patent drawing
  • US12079737B1 patent drawing
  • US12079737B1 patent drawing

AI summary

A method of data mining and AI workflow platform for structured and unstructured data is described. The method comprises receiving data from a data source, wherein the data comprises at least one data format of at least one of structured data, semi-structured data and unstructured data; indexing and analyzing the received data; scheduling and uploading automatically the data to a database as per the indexing; visualizing the data and determining at least one of the structured data, the semi-structured data, and the unstructured data from the data uploaded; cleansing and filtering the data based on at least one of an input from a user, and a predefined rule; labeling and annotating seamlessly the data available in the database; and building an artificial intelligence (AI) model based on at least one of the data available in the database, the input from the user, and the predefined rule.