Machine Learning Hot-Cold Data Processing for Database Storage Engines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database systems face low accuracy and recall rates in task execution due to reliance on fixed rules for hot and cold data identification, which fail to adapt to varying loads.
Innovation Solution
Integrate a machine learning model with the database storage engine to dynamically adjust data samples and training resources based on load pressure, enabling online updating and improving task execution accuracy and recall rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If fixed rules (LRU/LFU) are used for hot and cold data identification, then implementation is simple, but accuracy rate and recall rate of task results are low
Solution Approach 1:
The patent replaces the mechanical rule-based system (LRU/LFU algorithms) with a machine learning model system. The ML model learns optimal data identification strategies from historical database operation data, substituting fixed mechanical rules with adaptive intelligent decision-making that achieves higher accuracy and recall rates while maintaining reasonable implementation complexity.
Solution Approach 2:
The patent dynamically adjusts identification parameters by using the machine learning model to determine optimal thresholds and criteria for hot/cold data classification based on current database load conditions. This allows the system to change parameters adaptively rather than using fixed rule-based thresholds, improving measurement precision while managing complexity through automated parameter optimization.
2Measurement precision
If machine learning model is integrated for dynamic adjustment, then accuracy and recall rates improve, but device complexity increases
Solution Approach 1:
The patent implements self-service through automated model training and updating mechanisms. The system automatically collects historical operation data, trains the machine learning model, and updates it without requiring manual intervention. This self-service approach handles the complexity internally while presenting a simplified interface to users, maintaining high accuracy without proportionally increasing operational complexity.
Solution Approach 2:
The patent performs preliminary actions by pre-training the machine learning model using historical database operation data before actual hot/cold data identification tasks. This preliminary training phase prepares the model in advance, allowing it to handle complex identification tasks efficiently during runtime without adding real-time computational complexity to the main database operations.
3Measurement precision
If online updating is performed continuously, then model accuracy improves, but resource consumption increases
Solution Approach 1:
The patent implements periodic action through scheduled model updating mechanisms. Instead of continuous updating, the system updates the machine learning model at predetermined intervals or when specific triggers are met (e.g., accumulation of sufficient training data, time-based schedules). This periodic approach maintains model accuracy while significantly reducing resource consumption compared to continuous updating.
Solution Approach 2:
The patent applies dynamics by making the model updating frequency and intensity adaptive based on system conditions. The system dynamically adjusts updating behavior according to available resources, data changes, and performance requirements, allowing flexible balance between maintaining model accuracy and managing resource consumption in varying operational contexts.
Data Source
AI summary
Embodiments of this application provide a database task processing method, a hot and cold data processing method, a storage engine, a device, and a storage medium. In the embodiments of this application, a database storage engine is integrated with a machine learning model, involved data samples and used resources in training of the machine learning model are dynamically adjusted by sensing a load pressure of the database storage engine, and through data collection at a storage engine level and a design of a lightweight model, online updating is performed on a model based on a background task. In this way, effectiveness and availability that the database storage engine is integrated with the machine learning model are ensured, and an accuracy rate and a recall rate of a task execution result are improved.


