用于在数据库系统中估计不同值个数的方法

By integrating machine learning models into the database system and introducing a model fine-tuning mechanism, the problem of high time cost in NDV estimation in existing technologies is solved, achieving fast and accurate NDV estimation and improving the performance and efficiency of the database system.

CN118093544BActive Publication Date: 2026-07-17DOUYIN VISION CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2024-01-05
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing techniques require scanning the entire table when estimating the number of distinct values ​​(NDV) of a column in a large database system, resulting in high time costs and impacting system performance.

Method used

The machine learning model is integrated into the database system kernel, trained offline and loaded into a distributed file system for fast and accurate NDV estimation, and a model fine-tuning mechanism is introduced to deal with anomalies.

Benefits of technology

It enables fast and accurate estimation of NDV without affecting system performance, reduces the time overhead of full table scan, and improves the accuracy and efficiency of estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118093544B_ABST
    Figure CN118093544B_ABST
Patent Text Reader

Abstract

本公开的实施例提供了用于在数据库系统中估计不同值个数的方法、装置、设备和介质。方法包括:从分布式文件系统读取并加载用于NDV估计的模型到数据库系统的内核;生成数据库系统中的目标列的采样数据的特征化数据;使用已经加载到内核的模型,基于特征化数据确定针对目标列的估计的NDV;以及存储目标列的估计的NDV以供数据库系统的内核使用。以此方式,本公开的技术方案可以灵活地将基于机器学习的NDV估计模型集成到数据库系统内核,在不影响系统性能的情况下提供更准确和高效的NDV估计功能。
Need to check novelty before this filing date? Find Prior Art