Collaborative Dataset Consolidation System for Interoperability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage and computing technologies face challenges in managing and analyzing disparate datasets due to incompatibility and isolation among data silos, requiring significant effort for data scientists to assess and utilize datasets effectively.

Innovation Solution

A collaborative dataset consolidation system that facilitates interoperability among datasets through a programmatic interface, enabling the creation, analysis, and sharing of datasets across disparate platforms, using atomized data points and metadata to link and correlate datasets, and providing user interfaces for data discovery and visualization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional data storage technologies are used to store increasing amounts of data, then data capacity is improved, but data accessibility and interoperability deteriorate due to data silos

Engineering Contradiction:
Improvedata capacityVSAvoiddata interoperability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal data access interface that can interact with multiple different data storage systems and formats through a single standardized API. This interface layer provides multi-functional capabilities, allowing the same interface to access diverse data sources including structured databases, unstructured file systems, and cloud storage, thereby resolving the interoperability issue while maintaining data capacity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary data access layer that sits between the user applications and the underlying data storage systems. This intermediary layer translates various data formats and storage mechanisms into a unified access model, enabling seamless interoperability across different data silos without requiring changes to the underlying storage infrastructure

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data scientists manually download and analyze datasets using ad hoc approaches, then data analysis flexibility is improved, but time and effort required for data assessment deteriorate

Engineering Contradiction:
Improvedata analysis flexibilityVSAvoidtime for data assessment
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements preliminary data summarization and metadata generation that occurs automatically when datasets are uploaded to the system. Key statistics, data quality metrics, and descriptive information are pre-calculated and stored, allowing data scientists to quickly assess datasets without performing manual exploratory analysis, thereby reducing assessment time while maintaining analytical flexibility

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides self-service data assessment capabilities through automated generation of data summaries, quality reports, and contextual information. The dataset metadata includes pre-computed statistics and descriptors that enable data scientists to evaluate datasets independently without requiring extensive manual intervention or expert knowledge

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If contextual information is absent from ad hoc datasets, then data creation simplicity is improved, but data understanding and assessment difficulty deteriorate

Engineering Contradiction:
Improvedata creation simplicityVSAvoiddata understanding difficulty
Core Design Contradiction:
Ease of manufactureVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces an intermediary metadata layer that automatically captures and stores contextual information about datasets including data quality metrics, statistical summaries, and descriptive attributes. This metadata acts as a mediator between the raw data and the user, providing necessary contextual information without complicating the data creation process

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system automatically generates and updates data summary parameters and metadata attributes that describe key characteristics of the dataset. These parameters include data quality metrics, statistical measures, and contextual descriptors that are dynamically computed and stored, making contextual information readily available without requiring manual annotation during data creation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220374448A1Interactive interfaces to present data arrangement overviews and summarized dataset attributes for collaborative datasets
Publication Date: 2022.11.24 SERVICENOW INC
  • US20220374448A1 patent drawing
  • US20220374448A1 patent drawing
  • US20220374448A1 patent drawing

AI summary

Various embodiments relate generally to data science and data analysis, and computer software and systems, to provide an interface between repositories of disparate datasets and computing machine-based entities that seek access to the datasets, and, more specifically, to a computing and data storage platform that facilitates consolidation of one or more datasets, whereby user interfaces may be implemented as computerized tools for presenting summarization of dataset attributes to facilitate discovery, formation, and analysis of interrelated collaborative datasets. In some examples, a method may include receiving data resulting from insight calculations. Insight calculations may be based on a derived dataset attribute. Also, the method may include presenting a data arrangement overview summarizing the data attributes as an aggregation of data attributes in a portion of the user interface. The data arrangement overview may include an interactive display of a distribution associated with a collaborative atomized dataset.