Metadata Broker for Online Library Data Dependency Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Curators of online library systems face challenges in identifying and managing duplicative or derivative data sets, which can lead to functional bottlenecks and reliability issues, as they often lack the means to verify the quality and variety of information provided by multiple data providers, resulting in users accessing limited and potentially inaccurate information.

Innovation Solution

A processor-based apparatus and method that receives and compares normalized metadata portions from multiple vendor devices to identify identical or dependent data sets, generating visualizations to depict relationships among data sets and transmit them to client devices for presentation, enabling users to understand data dependencies and inclusions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If curators enter into licensing agreements with multiple data providers to broaden data sources, then the variety and completeness of information is improved, but duplicative or derivative data sets are introduced creating functional bottlenecks

Engineering Contradiction:
Improvedata source varietyVSAvoidsystem functionality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system implements automated feedback mechanisms by continuously monitoring and analyzing metadata from multiple data providers. The broker device receives metadata portions, normalizes them against a controlled vocabulary, and automatically identifies duplicative or derivative relationships. This feedback loop enables curators to detect and manage data redundancies in real-time, preventing functional bottlenecks while maintaining diverse data sources.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces a broker device as an intermediary between multiple data providers and the online library system. This intermediary normalizes metadata from various providers using a controlled vocabulary, compares metadata portions to identify duplicative relationships, and manages the complexity of multi-provider agreements. The broker acts as a mediator that harmonizes data from diverse sources while preventing redundancy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If curators rely on data providers to perform curation and separate reliable information from questionable information, then information quality is improved, but curators lose the ability to verify and evaluate the quality and variety of information

Engineering Contradiction:
Improveinformation qualityVSAvoidcurator verification capability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system provides curators with feedback through visualizations that display the provenance and relationships of data sets. By showing duplicative relationships, source dependencies, and data overlaps, the system enables curators to verify information quality and provider performance without manually examining raw data. This feedback mechanism preserves curator oversight while leveraging provider curation efforts.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The broker device performs preliminary actions by automatically analyzing and normalizing metadata before data is fully integrated into the system. By pre-processing metadata and identifying potential duplicative relationships in advance, the system prepares verification information for curators, enabling them to efficiently evaluate data quality and provider reliability without exhaustive manual review.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If the system stores and processes data from multiple providers to ensure users have access to diverse sources, then information completeness is improved, but storage and processing capabilities of individual sources may become insufficient creating access bottlenecks

Engineering Contradiction:
Improveinformation completenessVSAvoiddata access speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The system performs preliminary actions by pre-processing and normalizing metadata from multiple data providers before data access requests are made. The broker device prepares metadata portions, establishes controlled vocabulary mappings, and identifies relationships in advance. This preliminary preparation enables rapid data access and retrieval operations without overwhelming individual source systems during peak usage.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transitions from a single-dimension data storage model to a multi-dimensional metadata structure. By organizing data through normalized metadata portions, controlled vocabularies, and relationship graphs, the system creates additional dimensions for data organization and access. This dimensional transformation enables efficient querying and retrieval across diverse data sources without proportionally increasing storage or processing demands on individual sources.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10380214B1Identification and visualization of data set relationships in online library systems
Publication Date: 2019.08.13 SAS INSTITUTE INC
  • US10380214B1 patent drawing
  • US10380214B1 patent drawing
  • US10380214B1 patent drawing

AI summary

An apparatus includes a processor to: receive multiple normalized metadata portions based on metadata portions originating from vendor devices storing data sets of a distributed online library system; compare the multiple pieces of information between pairs of normalized metadata portions to identify at least one pair of identical portions of data; analyze the pieces of information of normalized metadata portions corresponding to an identified pair of identical portions of data to determine if there is a dependency relationship between each portion of data of the pair and another identical portion of data stored within another device; and in response to there being such a pair of dependency relationships, generate a visualization that includes a combination of graphical elements depicting the pair of dependency relationships, and transmit the visualization to the client device to enable a visual presentation of the visualization.