Semantic Versioning for Data Products With Automated Change Categorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data versioning systems lack reliable and standardized methods for detecting the magnitude and significance of changes in data products, leading to subjective manual analysis and inconsistent versioning conventions, which can cause downstream errors and inefficiencies.
Innovation Solution
An automated semantic versioning system that uses machine learning algorithms to detect and categorize changes in data assets, generating new semantic versions based on predefined thresholds and notifying clients of updates, ensuring all downstream systems use the most updated versions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual analysis is used to evaluate changes to data products, then flexibility in decision-making is maintained, but consistency and reliability of versioning deteriorates
Solution Approach 1:
The patent replaces manual mechanical analysis with an automated machine learning-based system that automatically detects, evaluates, and categorizes changes in data products. The system uses AI models to substitute human judgment, ensuring consistent and reliable versioning decisions without subjective variability.
Solution Approach 2:
The system performs self-service by automatically monitoring and evaluating its own data products for changes. The machine learning models continuously analyze data drift, schema modifications, and other variations, then autonomously determine appropriate version updates without requiring manual intervention.
2Reliability
If automated machine learning algorithms are used to detect changes, then versioning consistency and reliability improve, but system complexity increases
Solution Approach 1:
The patent segments the change detection process into distinct components: data drift detection, schema change detection, and category assignment. Each component handles a specific aspect of change analysis, making the overall complex system more manageable and maintainable through modular architecture.
Solution Approach 2:
The system manages complexity by dynamically adjusting detection parameters and thresholds based on the specific characteristics of each data product. The machine learning models adapt their sensitivity and categorization criteria to match the unique structure and requirements of different datasets.
3Productivity
If manual evaluation of changes is performed, then system complexity remains low, but productivity and efficiency deteriorate
Solution Approach 1:
The patent implements continuous monitoring of data products where the machine learning system operates continuously to detect changes as they occur. This eliminates the need for periodic manual checks, ensuring that versioning decisions are made continuously and rapidly in response to actual data changes.
Solution Approach 2:
The system performs preliminary detection and categorization of changes before they impact downstream systems. By proactively identifying and classifying changes in advance, the system prepares versioning decisions ahead of time, enabling rapid deployment without last-minute manual intervention.
4Measurement precision
If subjective manual analysis is used for versioning, then adaptability to unique cases is maintained, but measurement precision and standardization deteriorate
Solution Approach 1:
The patent uses categorical classification (analogous to color changes) to transform complex change detections into distinct, easily interpretable categories. Each category represents a specific type of change with predefined implications for versioning, making precise measurement of change magnitude both accurate and operationally simple.
Solution Approach 2:
The system incorporates feedback loops where the machine learning models continuously learn from actual versioning decisions and their outcomes. This feedback mechanism refines the measurement precision over time by adapting to edge cases and improving the accuracy of change magnitude assessment while maintaining operational simplicity.
Data Source
AI summary
A computer-implemented method for receiving evaluation criteria comprising rules for evaluating changes to a data file, where the data file comprises a plurality of assets. The method may further include detecting at least one change to one or more assets of the plurality of assets and identifying a category for the at least one change, based on the received evaluation criteria. In response to identifying a category for a plurality of changes, the method may aggregate the identified category of each of the plurality of changes of each asset of the data file. In response to the aggregated categories exceeding one or more predetermined thresholds, the method may further include generating a new semantic version for each changed asset of the plurality of assets and a new semantic version for the data file overall. The method may also push the new semantic version to at least one client computer.


