ROI-Based Data Graph for Unstructured Data Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current database management systems are inadequate for effectively managing unstructured data due to the limited number of system-level attributes available for describing and retrieving unstructured data types, such as images, audio, and social media postings, which require content-level attributes for accurate retrieval.

Innovation Solution

A data management system that identifies regions of interest (ROIs) in unstructured data items, encodes them into ROI vectors, creates a data graph to represent the data subjects and vectors, and stores this information in a graph database, allowing for the grouping of vectors into clusters and integration with structured data for comprehensive data management across formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If system-level attributes are used to describe unstructured data, then the data can be stored in existing database systems, but the retrieval accuracy is insufficient due to limited attributes

Engineering Contradiction:
Improveretrieval accuracyVSAvoiddata structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments unstructured data into structured components by identifying and extracting regions of interest (ROIs) from the original data. Each ROI is encoded into a vector representation, transforming the unstructured data into a structured format that can be efficiently queried while preserving retrieval accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces ROI vectors as an intermediary representation between the original unstructured data and the database storage system. These vectors serve as a bridge that enables accurate retrieval without requiring the database to directly handle complex unstructured data formats.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If content-level attributes are extracted from unstructured data, then retrieval accuracy improves, but the data management complexity increases

Engineering Contradiction:
Improveretrieval accuracyVSAvoiddata management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical data management approaches with AI-based automated ROI identification and encoding systems. Machine learning models automatically extract meaningful regions from unstructured data and convert them into vectors, eliminating the need for manual attribute extraction while improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If a graph database structure is created to represent data subjects and ROI vectors, then data retrieval efficiency improves, but the initial processing time and computational resources increase

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidinitial processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary processing by pre-extracting ROIs and pre-encoding them into vectors during data ingestion. The graph database structure is pre-built with these vectors, so that when retrieval operations occur, the heavy computational work has already been completed, enabling fast query responses.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11899754B2ROI-based data content graph for wide data management
Publication Date: 2024.02.13 DELL PROD LP
  • US11899754B2 patent drawing
  • US11899754B2 patent drawing
  • US11899754B2 patent drawing

AI summary

This disclosure provides systems, methods, and media for creating a data graph database from various unstructured and unstructured data items for use by various services. The method comprises the operations of identifying unstructured data items in data subjects; recognizing regions of interest (ROIs) in the unstructured data items; and extracting the ROIs from the unstructured data items. The method further comprises encoding the extracted ROIs into ROI vectors; creating a data graph to represent the data subjects, the data items, and the ROI vectors; and storing the data graph into a graph database. The various embodiments can manage data items of different data formats together rather than separately, thus creating a data management system for managing data across data formats. The data management system can also store structured data items into the graph database, thus complementing the existing ETL procedure for structured data items.