Compressed Data Representation Matching for Tabular Graph Interoperability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage and computing technologies face challenges in interoperability among disparate datasets due to differences in computing platforms, database technologies, and data formats, leading to isolated 'data silos' that are difficult to scale and match data across large datasets.

Innovation Solution

A collaborative dataset consolidation system that uses a dataset ingestion controller to compute compressed data representations for tabular data and match them against reference compressed data representations in graph-based data arrangements, facilitating the identification and linking of equivalent data subsets across different data formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional data storage and computing technologies are used to store disparate datasets, then data can be stored in different formats and platforms, but data interoperability deteriorates and isolated data silos are formed

Engineering Contradiction:
Improvedata format compatibilityVSAvoiddata interoperability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces a graph-based data arrangement as an intermediary layer between disparate tabular datasets. This graph structure serves as a universal mediator that can represent relationships between data from different sources, formats, and platforms, enabling interoperability without requiring direct compatibility between the original data silos.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The graph data arrangement is designed to perform multiple functions: storing data from various tabular formats, representing relationships between datasets, enabling pattern matching, and supporting different query types. This universal structure replaces the need for multiple specialized storage systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If relational database architectures are used to store large datasets, then data can be organized in tabular formats, but scalability and data matching efficiency deteriorate

Engineering Contradiction:
Improvedataset sizeVSAvoiddata matching efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent transitions from static tabular data arrangements to dynamic graph-based structures that can adaptively represent relationships as data grows. The graph structure allows for efficient traversal and pattern matching even as the quantity of data increases, unlike rigid relational schemas that become progressively slower.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent adds a new dimension to data storage by introducing graph-based relationships alongside traditional tabular structures. This additional dimensional layer enables efficient pattern matching and data discovery that is not possible with conventional two-dimensional spreadsheet or database tables alone.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If manual intervention is used to identify related datasets, then data can be accurately matched between formats, but time consumption and operational complexity increase

Engineering Contradiction:
Improvedata matching accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service data matching through pattern recognition algorithms that operate on the graph-based data arrangement. The structure itself facilitates automatic identification of related datasets through inherent relationship representations, eliminating the need for manual intervention while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The graph-based arrangement provides feedback mechanisms through pattern matching that automatically identify and validate relationships between datasets. This feedback loop enables the system to continuously improve data matching accuracy without manual input, learning from the structural relationships inherent in the graph data.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11669540B2Matching subsets of tabular data arrangements to subsets of graphical data arrangements at ingestion into data-driven collaborative datasets
Publication Date: 2023.06.06 SERVICENOW INC
  • US11669540B2 patent drawing
  • US11669540B2 patent drawing
  • US11669540B2 patent drawing

AI summary

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to interface among repositories of disparate datasets and computing machine-based entities configured to access datasets, and, more specifically, to a computing and data storage platform to identify and match equivalent subsets of data between an ingested dataset, such as in a tabular data arrangement, and one or more graph-based data arrangements, according to at least some examples. For example, a method may include identifying a tabular data arrangement including a subset of data as a column, computing a compressed data representation for a column of data, correlating a compressed data representation to a reference compressed data representations, detecting a link between a column of data associated with a correlated compressed data representation to a dataset stored in a graph data arrangement, and forming an expanded tabular data arrangement.