用于在数据库系统中估计不同值个数的方法
By integrating machine learning models into the database system and introducing a model fine-tuning mechanism, the problem of high time cost in NDV estimation in existing technologies is solved, achieving fast and accurate NDV estimation and improving the performance and efficiency of the database system.
CN118093544BActive Publication Date: 2026-07-17DOUYIN VISION CO LTD +1
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-01-05
- Publication Date
- 2026-07-17
AI Technical Summary
Technical Problem
Existing techniques require scanning the entire table when estimating the number of distinct values (NDV) of a column in a large database system, resulting in high time costs and impacting system performance.
Method used
The machine learning model is integrated into the database system kernel, trained offline and loaded into a distributed file system for fast and accurate NDV estimation, and a model fine-tuning mechanism is introduced to deal with anomalies.
Benefits of technology
It enables fast and accurate estimation of NDV without affecting system performance, reduces the time overhead of full table scan, and improves the accuracy and efficiency of estimation.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN118093544B_ABST
Abstract
本公开的实施例提供了用于在数据库系统中估计不同值个数的方法、装置、设备和介质。方法包括:从分布式文件系统读取并加载用于NDV估计的模型到数据库系统的内核;生成数据库系统中的目标列的采样数据的特征化数据;使用已经加载到内核的模型,基于特征化数据确定针对目标列的估计的NDV;以及存储目标列的估计的NDV以供数据库系统的内核使用。以此方式,本公开的技术方案可以灵活地将基于机器学习的NDV估计模型集成到数据库系统内核,在不影响系统性能的情况下提供更准确和高效的NDV估计功能。
Need to check novelty before this filing date? Find Prior Art