Semantic Dataset Synopsis Generation for Faster Data Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in selecting appropriate datasets due to the large number of available datasets with limited documentation, making it a tedious and error-prone task, especially in data science and machine learning applications.

Innovation Solution

A computer-implemented method generates a customized semantic synopsis of datasets using semantic and spatio-temporal analysis, tailored to user roles, contexts, and search queries, providing natural language descriptions of dataset content, locations, and time periods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual documentation of datasets is required, then accuracy and completeness of dataset descriptions are improved, but time consumption and labor effort increase significantly

Engineering Contradiction:
Improveaccuracy of dataset descriptionVSAvoidtime consumption for documentation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables datasets to document themselves automatically through semantic analysis of data contents, schema, and metadata. The semantic synopsis generator processes dataset information autonomously without requiring manual intervention, producing accurate descriptions that reflect the actual data structure and content.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual documentation processes with automated semantic analysis mechanisms. Using natural language processing and semantic analysis algorithms, the system automatically generates dataset descriptions, replacing the mechanical manual writing process with computational methods that achieve comparable or superior accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If detailed dataset documentation is provided, then ease of dataset selection and understanding is improved, but the complexity of data catalog management increases

Engineering Contradiction:
Improveease of dataset selectionVSAvoidcomplexity of data catalog management
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system segments dataset documentation into structured components including semantic synopsis, schema information, and metadata. This segmentation allows detailed information to be organized in manageable sections, making the data catalog easier to navigate and understand while maintaining comprehensive documentation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The semantic synopsis is generated in advance as a preliminary summary that captures essential dataset characteristics. This pre-computed information is readily available for users to review without requiring them to delve into complex underlying data structures, simplifying the selection process.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated semantic analysis is applied to all datasets, then productivity of dataset discovery is improved, but computational resources and processing time are consumed

Engineering Contradiction:
Improveproductivity of dataset discoveryVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system applies semantic analysis selectively rather than uniformly to all datasets. It prioritizes analysis for datasets that lack documentation or require detailed descriptions, while using lighter processing methods for well-documented datasets, optimizing the balance between productivity and resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12554932B2Dynamic semantic synopsis generation for datasets in data catalog
Publication Date: 2026.02.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12554932B2 patent drawing
  • US12554932B2 patent drawing
  • US12554932B2 patent drawing

AI summary

Generating a semantic synopsis is provided. A dataset comprising structured data is received and semantic analysis of the dataset is performed to determine categories of information included in the datasets. A spatio-temporal analysis of the dataset is also performed to determine time periods and locations to which the dataset applies. A static synopsis of the dataset is then generated in natural language sentences that describes semantic information in the dataset and time periods and locations covered by the dataset. The synopsis can be customized according to a user role and search query data.