Keyword Model Visual Interface for Data Source Relationship Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing large data structures are inefficient, as users manually search for and relate data structures, and existing automated methods often fail to capture relevant relationships between data sources, leading to wasted time and resources.

Innovation Solution

A system generates a keyword model with weighted keywords for multiple data sources, providing a visual representation and allowing users to select relevant data sources, generating relational models and a combined dataset model based on user input to efficiently associate and link data structures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users manually search and identify data structures, then they can understand and control the data relationships, but the process is time-consuming and inefficient

Engineering Contradiction:
Improveaccuracy of identifying data relationshipsVSAvoidtime spent on manual data analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automated analysis to generate candidate data relationships and visualizations before user review. The automated method pre-processes the data structures, identifies potential relationships using algorithms, and creates visual representations, so that users only need to review and confirm pre-prepared options rather than manually analyzing everything from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary automated analysis layer between the raw data structures and the user. This intermediary component uses algorithms to generate candidate relationships and visualizations, which then serve as intermediates for user verification. The user interacts with the intermediary output rather than directly analyzing raw data, combining automated efficiency with human judgment.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated methods are used to identify data relationships, then time efficiency is improved, but the accuracy and relevance of identified relationships deteriorates

Engineering Contradiction:
Improvespeed of data analysisVSAvoidaccuracy of data relationship identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback loops where user interactions with generated visualizations and relationship suggestions are fed back into the automated analysis process. User confirmations, corrections, and selections provide feedback signals that refine and improve the automated relationship identification algorithms over time, allowing the system to learn from user expertise while maintaining automated speed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The automated method serves as an intermediary that generates candidate relationships quickly, which are then refined through user review. The intermediary automated component provides initial high-speed analysis, and the user acts as a secondary intermediary to filter and validate the results, combining the speed of automation with the accuracy of human judgment.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If users review all data structures manually, then comprehensive understanding is achieved, but the complexity and resource requirements increase

Engineering Contradiction:
Improvecompleteness of data relationship understandingVSAvoidcomplexity of data analysis process
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the large set of data structures into manageable groups or categories based on preliminary automated analysis. Instead of presenting all data structures at once, the system divides them into segments that can be reviewed more efficiently, organizing them by relevance, type, or relationship strength. This segmentation reduces the perceived complexity while maintaining comprehensive coverage through systematic organization.

Inventive Principle:
Principle #1Segmentation

4Productivity

If existing automated methods are used to relate data structures, then some relationships are identified, but relevant information for related data structures is often missing

Engineering Contradiction:
Improveautomation of data relationship identificationVSAvoidrelevance information in related data structures
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary automated analysis to identify not just direct relationships but also potential indirect relationships and relevant contextual information. Before user review, the system pre-processes data to include comprehensive relationship candidates and relevant metadata, ensuring that when users review the results, complete information is already prepared and organized for their consideration.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11314784B2Data model proposals
Publication Date: 2022.04.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11314784B2 patent drawing
  • US11314784B2 patent drawing
  • US11314784B2 patent drawing

AI summary

Relating data in various distributed data sources for use in data analysis is described. The data sources are generally related by first generating a keyword model for a plurality of data sources, which includes a plurality of weighted keywords, and providing a visual representation of the keyword model, such as a word cloud, to a user. The user interacts with the visual representation to modify, update, and select various aspects of the visual representation. The user also identifies keywords and data sources of interest such that a plurality of relational models are generated based on the user interest. Relating the data sources also includes providing the plurality of relational models to the user, receiving a user selection of the plurality of relational models, and generating a combined dataset model which relates one or more of the data sources according to the selected relational models.