Graph Data Subset Classification for Dataset Interoperability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage and computing technologies face challenges in interoperability and scalability, particularly with graph-based data structures, leading to inefficiencies in identifying and joining relevant datasets due to the complexity of matching data across disparate formats and systems, resulting in data silos that hinder effective data operations and analysis.

Innovation Solution

A computerized toolset is implemented to analyze and classify data subsets within graph data arrangements, using compressed data representations and similarity matrices to determine degrees of similarity and joinability, facilitating the linking of tabular and graph-based datasets through a collaborative dataset consolidation system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional data storage and computing technologies are used to store and access datasets, then data can be stored in various formats (CSV, TSV, HTML, JSON, XML), but data interoperability and accessibility are hindered due to data silos and incompatible access techniques

Engineering Contradiction:
Improvedata interoperabilityVSAvoiddata access complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal data access layer that provides a standardized interface for accessing multiple data formats (CSV, TSV, HTML, JSON, XML) through a single unified mechanism. This access layer translates various data formats into a common representation, enabling interoperability across different data sources without requiring format-specific access techniques for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary data access layer positioned between the diverse data sources and the analysis tools. This intermediary layer handles format conversion, data validation, and access coordination, serving as a mediator that reconciles the incompatibilities between different data formats and access methods while presenting a unified interface to users.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If graph-based data structures are used to represent complex relationships among datasets, then data relationships can be modeled more effectively, but computational complexity increases making it difficult to identify and join relevant datasets

Engineering Contradiction:
Improvedata relationship accuracyVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex graph-based data structure into hierarchical layers, organizing data relationships from general to specific. This layered segmentation divides the computational task of traversing and analyzing relationships into manageable stages, reducing the overall computational complexity while preserving the accuracy of relationship modeling through structured organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary indexing and pre-processing of graph data structures before actual data analysis operations. By pre-computing relationship paths, indexing connected components, and organizing graph data in advance, the system reduces the computational burden during query execution and dataset joining operations, making complex graph traversals more efficient.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual techniques are used to identify and join relevant datasets across disparate formats, then data can be accurately matched, but the process is time-consuming and inefficient for large-scale data operations

Engineering Contradiction:
Improvedata matching accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements automated data matching and joining capabilities that perform self-service operations without requiring manual intervention. The system automatically identifies relevant datasets, matches data elements across different formats using standardized comparison algorithms, and executes joining operations based on predefined criteria, thereby maintaining high matching accuracy while dramatically improving processing efficiency for large-scale data operations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the data matching process by changing parameters from manual inspection to automated computational comparison. By converting qualitative matching criteria into quantifiable parameters and applying algorithmic comparison methods, the system achieves both high accuracy in data matching and improved efficiency, processing large numbers of datasets automatically with consistent results.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12008050B2Computerized tools configured to determine subsets of graph data arrangements for linking relevant data to enrich datasets associated with a data-driven collaborative dataset platform
Publication Date: 2024.06.11 SERVICENOW INC
  • US12008050B2 patent drawing
  • US12008050B2 patent drawing
  • US12008050B2 patent drawing

AI summary

Various embodiments relate generally to data science and data analysis, computer software and systems, and wired and wireless network communications to interface among repositories of disparate datasets and computing machine-based entities configured to access datasets, and, more specifically, to a computing and data storage platform to implement computerized tools to identify data classifications and similar subsets of graph-based data arrangements with which to join, according to at least some examples. For example, a method may include determining a classification type for a subset of data based on a graph data arrangement, generating presentation data as a first user input to detect selection of the first user input. The method may include predicting a classification type for data. The method may also include generating other presentation data configured to join datasets, such as a column of tabular-formatted data with one or more portions of a graph data arrangement.