Semantic Versioning for Data Products With Automated Change Categorization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data versioning systems lack reliable and standardized methods for detecting the magnitude and significance of changes in data products, leading to subjective manual analysis and inconsistent versioning conventions, which can cause downstream errors and inefficiencies.

Innovation Solution

An automated semantic versioning system that uses machine learning algorithms to detect and categorize changes in data assets, generating new semantic versions based on predefined thresholds and notifying clients of updates, ensuring all downstream systems use the most updated versions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual analysis is used to evaluate changes to data products, then flexibility in decision-making is maintained, but consistency and reliability of versioning deteriorates

Engineering Contradiction:
Improveversioning consistencyVSAvoidautomation system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces manual mechanical analysis with an automated machine learning-based system that automatically detects, evaluates, and categorizes changes in data products. The system uses AI models to substitute human judgment, ensuring consistent and reliable versioning decisions without subjective variability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically monitoring and evaluating its own data products for changes. The machine learning models continuously analyze data drift, schema modifications, and other variations, then autonomously determine appropriate version updates without requiring manual intervention.

Inventive Principle:
Principle #25Self-service

2Reliability

If automated machine learning algorithms are used to detect changes, then versioning consistency and reliability improve, but system complexity increases

Engineering Contradiction:
Improvechange detection accuracyVSAvoidautomation system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the change detection process into distinct components: data drift detection, schema change detection, and category assignment. Each component handles a specific aspect of change analysis, making the overall complex system more manageable and maintainable through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system manages complexity by dynamically adjusting detection parameters and thresholds based on the specific characteristics of each data product. The machine learning models adapt their sensitivity and categorization criteria to match the unique structure and requirements of different datasets.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If manual evaluation of changes is performed, then system complexity remains low, but productivity and efficiency deteriorate

Engineering Contradiction:
Improveversioning speedVSAvoidautomation infrastructure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements continuous monitoring of data products where the machine learning system operates continuously to detect changes as they occur. This eliminates the need for periodic manual checks, ensuring that versioning decisions are made continuously and rapidly in response to actual data changes.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary detection and categorization of changes before they impact downstream systems. By proactively identifying and classifying changes in advance, the system prepares versioning decisions ahead of time, enabling rapid deployment without last-minute manual intervention.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If subjective manual analysis is used for versioning, then adaptability to unique cases is maintained, but measurement precision and standardization deteriorate

Engineering Contradiction:
Improvechange magnitude measurementVSAvoidoperational simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent uses categorical classification (analogous to color changes) to transform complex change detections into distinct, easily interpretable categories. Each category represents a specific type of change with predefined implications for versioning, making precise measurement of change magnitude both accurate and operationally simple.

Inventive Principle:
Principle #32Color changes

Solution Approach 2:

The system incorporates feedback loops where the machine learning models continuously learn from actual versioning decisions and their outcomes. This feedback mechanism refines the measurement precision over time by adapting to edge cases and improving the accuracy of change magnitude assessment while maintaining operational simplicity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250355846A1Semantic versioning calculator for data products
Publication Date: 2025.11.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250355846A1 patent drawing
  • US20250355846A1 patent drawing
  • US20250355846A1 patent drawing

AI summary

A computer-implemented method for receiving evaluation criteria comprising rules for evaluating changes to a data file, where the data file comprises a plurality of assets. The method may further include detecting at least one change to one or more assets of the plurality of assets and identifying a category for the at least one change, based on the received evaluation criteria. In response to identifying a category for a plurality of changes, the method may aggregate the identified category of each of the plurality of changes of each asset of the data file. In response to the aggregated categories exceeding one or more predetermined thresholds, the method may further include generating a new semantic version for each changed asset of the plurality of assets and a new semantic version for the data file overall. The method may also push the new semantic version to at least one client computer.