ML Data Profiling Visualizations for Distributed Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data scientists face challenges in familiarizing themselves with new data sources, particularly in large organizations with numerous data sets and legacy systems, where understanding data sources, their limitations, and relationships is time-consuming and complex, especially in telecommunication networks with vast and distributed data systems.

Innovation Solution

The method involves generating annotated visual objects that represent data sets with interrelated data objects, using a descriptive classifier to select visual objects based on data object properties and relationships, creating abstract visual representations that facilitate human understanding and enhance machine-understandable data associations, and utilizing these representations for efficient data processing and classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data scientists manually examine and analyze data sources to understand their properties and relationships, then they can gain comprehensive knowledge of the data, but the time required for data familiarization increases significantly

Engineering Contradiction:
Improvedata understandingVSAvoiddata familiarization time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary data profiling and generates visual representations of data sources, their properties, and relationships before data scientists need to analyze them. This advance preparation includes automatically examining data characteristics, identifying relationships between data objects, and creating visual summaries that are ready when scientists need to understand new data sources.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary system that acts as a mediator between raw data sources and data scientists. This intermediary automatically profiles data, generates visual representations, and presents synthesized information about data properties and relationships, eliminating the need for scientists to manually examine raw data while still providing comprehensive understanding.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If the system stores and processes detailed profiles of all data objects and their relationships, then the quality and completeness of data understanding improves, but the complexity of the data processing system increases

Engineering Contradiction:
Improvedata object relationshipsVSAvoiddata processing system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments data profiling into distinct components: data object profiles, relationship profiles, and visual representation generation. Each component handles a specific aspect of data understanding, allowing the complex task of data analysis to be divided into manageable, independent modules that can be processed and stored separately.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of storing and processing all raw data, the system creates simplified copies in the form of data profiles and visual representations. These copies capture the essential properties and relationships of data objects without requiring the full complexity of the original data, reducing processing demands while maintaining understanding quality.

Inventive Principle:
Principle #26Copying

3Ease of operation

If the system generates detailed visual representations with annotations for each data set, then human understanding of data relationships is enhanced, but the computational resources required for generation and storage increase

Engineering Contradiction:
Improvehuman understandingVSAvoidcomputational resources
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system applies local quality by generating visual representations with annotations focused on specific data objects and their relationships rather than attempting to visualize entire data sets uniformly. Each visual representation is tailored to highlight locally relevant properties and relationships, providing human-understandable insights without the computational overhead of comprehensive universal visualization.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240370454A1Machine learning-based data set profiling and visualization
Publication Date: 2024.11.07 AT&T INTELLECTUAL PROPERTY I L P
  • US20240370454A1 patent drawing
  • US20240370454A1 patent drawing
  • US20240370454A1 patent drawing

AI summary

A processing system may obtain at least a first data object of a data set and obtain a profile of the at least the first data object, where the profile defines at least one property of the first data object and at least one relationship between the at least the first data object and at least a second data object of the data set. The processing system may then select at least a first component of a visual object for the at least the first data object based upon the at least one property and the at least one relationship, label the at least the first component of the visual object in accordance with the at least the first data object, in response to the selecting, and present the visual object with the at least the first component labeled in accordance with the at least the first data object.