Data Management System Snapshot Acquisition for AI Model Relearning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data management systems for AI applications fail to store and retrieve past data effectively, making it impossible for AI users to relearn or perform evidence searches using historical data, as necessary data is not stored or coordinated with AI analysis applications.
Innovation Solution
A data management system with a data analysis server, storage areas for raw, curated, and learned data models, and a storage management server that manages data configurations and acquires snapshots of curated and raw data to enable AI users to relearn or search past data models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is updated on the data management system over time, then the data management system can be optimized and improved, but the original data and learning data sets are lost making relearning impossible
Solution Approach 1:
The system performs preliminary actions by acquiring snapshots of the second storage area (curated data) and third storage area (raw data) at the moment before data updates occur. This preliminary capture of data states enables future relearning and evidence search while allowing the system to proceed with data optimization and updates without losing historical information.
Solution Approach 2:
The invention creates copies of the original data and curated data through snapshot acquisition. These snapshots serve as preserved copies that can be used for relearning and evidence search, while the original system continues to update and optimize the main data stores without being constrained by the need to preserve every historical version.
2Reliability
If backup images are stored using conventional techniques, then data can be restored, but coordination with AI analysis applications is not considered making past data inaccessible for relearning
Solution Approach 1:
The snapshot acquisition mechanism serves multiple functions: it enables data restoration (conventional backup function) and simultaneously supports AI analysis applications by preserving the correspondence between learned data models, curated data, and raw data. This multi-functional approach allows the same backup mechanism to serve both traditional data recovery needs and AI-specific relearning requirements.
Solution Approach 2:
The data management server acts as an intermediary that coordinates between the backup system and AI analysis applications. It manages the correspondence information that links learned data models with their associated curated data and raw data snapshots, enabling AI applications to access and utilize past data effectively while maintaining conventional backup and restoration capabilities.
3Ease of operation
If multiple storage areas are managed for different data types, then data organization is improved, but system complexity increases
Solution Approach 1:
The data management server serves as an intermediary that manages the complexity of coordinating multiple storage areas. It maintains correspondence information that links the first storage area (learned data models), second storage area (curated data), and third storage area (raw data), abstracting away the complexity from users while providing organized data access and snapshot management across all storage areas.
Data Source
AI summary
The data analysis server manages first data configuration information for managing a correspondence of the learned data model, a raw data bucket that stores the raw data for generating the learned data model, and a curated data bucket that stores the curated data for generating the learned data model and second data configuration information for managing correspondences of the learned data model and the first storage area, the curated data bucket and the second storage area, and the raw data bucket and the third storage area, and gives an instruction to acquire snapshots of the second storage area that stores the curated data for generating the learned data model and the third storage area that stores the raw data for generating the learned data model to the data storage via the storage management server when the learned data model is generated.


