A method and device for estimating the cardinality of an unstructured data function (UDF) query
Patent Information
- Application Number
- CN202610997169.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本发明的目的在于针对非结构化数据UDF查询基数估计中查询结构表达不足、数据域差异难以刻画、UDF属性未被充分利用以及小样本估计不稳定的问题,提供一种非结构化数据UDF查询基数估计方法及装置
[0018]本发明的有益效果包括:第一,通过查询编码、数据概要、UDFSig和采样信息四类特征共同建模,能够同时表征查询结构、数据集合背景、UDF算子静态属性和样本命中证据;第二,通过采样命中统计、区间信息和平滑估计构造采样信息特征,能够降低低选择率查询在小样本条件下的估计波动;第三,通过UDFSig特征引入UDF算子的输入输出属性、标签词表规模和执行属性,使模型能够区分不同UDF算子对基数估计的影响;第四,采用梯度提升决策树模型对融合特征进行非线性回归,能够刻画不同特征块之间的组合关系,从而在不完整执行待估计查询的情况下得到基数估计值。
Smart Images

Figure CN122817271A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of database query optimization, machine learning systems, and unstructured data processing, and particularly to a method and apparatus for estimating the cardinality of unstructured data UDF queries. UDF, as referred to herein, is a user-defined function, which can be a classifier, detector, recognizer, extractor, or other function capable of outputting label results for unstructured records. Background Technology
[0002] With the widespread application of unstructured data such as text, images, and videos in intelligent search, recommendation, content moderation, data analysis, and multimedia management systems, more and more data management systems are encapsulating machine learning models or rule processing logic into UDF operators, allowing users to express query conditions through multiple UDF operators and their labeled results. A typical query may require a record to match any of several candidate labels under a certain UDF, or it may require a record to simultaneously satisfy multiple UDF clauses.
[0003] Cardinality is the number of records that satisfy a query condition. It is a crucial factor for the query optimizer in selecting the execution order, estimating execution costs, allocating computational resources, and determining whether pre-filtering is necessary. For traditional structured data, database systems typically use column histograms, sampling, or learned estimators to estimate predicate selectivity. However, for unstructured data UDF queries, the input is text, images, or videos, and the output is usually a set of labels. Furthermore, UDFs have high execution costs, large label spaces, and query structures involving intra-clause OR relations and inter-clause AND relations. Therefore, traditional cardinality estimation methods for structured data cannot be simply applied.
[0004] Existing methods, relying solely on the independence assumption, struggle to express the combined relationships between query structures and different UDFs, and are prone to significant biases in scenarios with a large number of labels, low query selectivity, or strong correlation between UDF results. While estimation methods dependent on small samples are simple to implement, they may suffer from zero hits, excessively wide intervals, or large estimation fluctuations when the sample hit rate is low. Schemes relying solely on model learning, lacking interpretable query structure features, data domain summaries, and UDF attributes, also struggle to achieve stable generalization across different data domains, different UDF sets, and different query formats.
[0005] Therefore, there is a need for a cardinality estimation method for unstructured data UDF queries that models the query itself, the dataset, the static attributes of the UDF, and the sample hit evidence together, so that the system can obtain stable cardinality estimation results with strong generalization ability without having to fully execute the query to be estimated. Summary of the Invention
[0006] The purpose of this invention is to address the problems of insufficient expression of query structure, difficulty in characterizing data domain differences, underutilization of UDF attributes, and instability of small sample estimation in UDF query cardinality estimation of unstructured data, and to provide a method and apparatus for estimating the query cardinality of unstructured data UDF.
[0007] The objective of this invention is achieved through the following technical solution: a method for estimating the cardinality of unstructured data UDF queries, comprising: (1) Obtain unstructured data set, UDF operator information and UDF query to be estimated, parse the UDF query to be estimated into one or more clauses, each clause includes a UDF operator and one or more labels, multiple labels in the same clause are matched according to the OR relation, and different clauses are matched according to the AND relation; (2) Generate query encoding features based on the UDF query to be estimated, wherein the query encoding features characterize the clause structure, operator structure, label structure and logical complexity of the UDF query to be estimated; (3) Generate data summary features based on the unstructured data set and its query workload, wherein the data summary features characterize the data domain type, record size, UDF size, tag vocabulary size and query workload overview; (4) Generate UDFSig features based on the UDF operator information involved in the UDF query to be estimated. The UDFSig features represent the static attributes of the UDF operators and the aggregation result of the static attributes of multiple UDF operators. (5) Determine a sample record set from the unstructured data set according to the preset sampling strategy, and perform matching on the sample record set according to the query semantics of the UDF query to be estimated to generate sampling information features; (6) The query encoding features, data summary features, UDFSig features and sampling information features are fused into cardinality estimation features, and the cardinality estimation features are input into the cardinality estimation model to obtain the cardinality estimate value corresponding to the UDF query to be estimated.
[0008] Furthermore, the unstructured data set includes text data, image data, video data, or combinations thereof, and the UDF operator is used to perform classification, detection, recognition, extraction, or label generation on the unstructured records, and output one or more label results.
[0009] Further, in step (1), the UDF query to be estimated is parsed into one or more clauses, including: reading the list of operators and the list of predicates in the UDF query to be estimated, pairing each UDF operator with the corresponding set of predicate labels to form a clause, deduplicating and sorting the labels in the same clause, and generating a query specification signature based on the UDF operator name and the sorted set of labels in the clause.
[0010] Furthermore, the query encoding features are used to characterize the query structure itself, including at least one of the following: number of query clauses, number of UDF operators, number of unique UDF operators, information on repeated operators, number of predicate tags, statistics on the number of tags in each clause, information on single-tag clauses, information on multi-tag clauses, OR relation complexity, AND relation complexity, query logic width, query logic depth, query shape, statistics on tag text length, tag vocabulary hit information, information on unknown tags, valid markers for each clause, and operator identifiers corresponding to each clause; the query encoding features do not depend on sample hit results and can be generated directly after receiving the query.
[0011] Furthermore, the data summary features are used to characterize the overall context of the dataset and its workload, including at least one of the following: total number of records in the dataset, data domain identifiers, number of UDF operators, total size of the tag vocabulary, statistics on the size of each UDF tag vocabulary, query workload size, average number of clauses per workload, average number of predicates per workload, proportion of queries with different number of clauses, statistics on the size of original records, and proportion of missing original records; the data summary features are used to enable the model to distinguish cardinality ranges under different sizes, different data types, and different query workloads.
[0012] Furthermore, the UDFSig feature is used to characterize the static attributes of the UDF operators involved in the query, including at least one of the following: UDF label vocabulary size, whether the UDF vocabulary is closed, whether the UDF outputs multiple labels, UDF input type, UDF output type, whether the UDF is executed deterministically, UDF call cost information, UDF model type, and UDF complexity information; for a UDF query to be estimated that contains multiple clauses, the UDFSig field (UDFSig feature) of the multiple UDF operators involved in the UDF query to be estimated is aggregated by mean, maximum, minimum or summation to form a query-level UDFSig feature with fixed dimensions.
[0013] Furthermore, the sampling information features are used to characterize direct hit evidence on the sample records and are generated through the following preset sampling strategy: one or more sample record sets are determined according to a preset sampling ratio and sampling number; for each sample record, the label result under the UDF operator involved in the UDF query to be estimated is read; labeling or matching is performed on each clause; the matching results of multiple clauses are performed and combined; and the number of sample hits, sample selection rate, and amplified sample cardinality are statistically analyzed based on the sampled matching results. The sampling information features provide the model with observational evidence directly related to the UDF query to be estimated and improve the stability in small sample scenarios through smoothing and interval statistics.
[0014] Furthermore, the sampling information features include at least one of the following: sampling ratio, sample size, number of sampling repetitions, sample hit count statistics, sample selection rate statistics, zero hit ratio, any hit identifier, binomial proportional standard error, confidence interval information, relative interval width, smoothed selection rate, smoothed cardinality estimate, selection rate of each clause, number of hits of each clause, combined estimate based on clause selection rate, ratio of actual sample selection rate to combined estimate, and difference between actual sample selection rate and combined estimate.
[0015] Furthermore, the cardinality estimation model is a gradient boosting decision tree model. The gradient boosting decision tree model takes the cardinality estimation features obtained by fusing the query encoding features, data summary features, UDFSig features, and sampling information features as input, and outputs the cardinality estimate of the UDF query to be estimated. When training the gradient boosting decision tree model, the cardinality of the query truth or the transformed value of the query truth cardinality is used as the supervision target, and the cardinality estimation features are used as the model input. When estimating the UDF query to be estimated, the output of the gradient boosting decision tree model is converted into a non-negative value and used as the cardinality estimate.
[0016] Furthermore, a sampling mean model, a smoothed sampling model, and a combined estimation model based on clause selectivity can be set as a control or auxiliary estimation model. The smoothed sampling model can include the Laplace smoothed sampling model and the Jeffreys smoothed sampling model, used to provide smoothed sampling estimation results when the number of sample hits is small. The aforementioned control estimation models can be used for model performance comparison, feature validity verification, or as supplementary input to the cardinality estimation model.
[0017] The present invention also provides a cardinality estimation device for unstructured data UDF queries, comprising: The query parsing module is used for unstructured data sets, UDF operator information, and UDF queries to be estimated. It parses the UDF queries to be estimated into one or more clauses. Each clause includes a UDF operator and one or more labels. Multiple labels within the same clause are matched according to an OR relationship, and different clauses are matched according to an AND relationship. The query encoding module is used to generate query encoding features based on the UDF query to be estimated. The query encoding features characterize the clause structure, operator structure, label structure and logical complexity of the UDF query to be estimated. The data summarization module is used to generate data summary features based on the unstructured data set and its query workload. The data summary features characterize the data domain type, record size, UDF size, tag vocabulary size, and query workload overview. The UDFSig module is used to generate UDFSig features based on the UDF operator information involved in the UDF query to be estimated. The UDFSig features represent the static attributes of the UDF operators and the aggregation result of the static attributes of multiple UDF operators. The sampling execution module is used to determine a sample record set from the unstructured data set according to a preset sampling strategy, perform matching on the sample record set according to the query semantics of the UDF query to be estimated, and generate sampling information features. The feature fusion module is used to fuse the query encoding features, data summary features, UDFSig features, and sampling information features into cardinality estimation features; The cardinality estimation module is used to input the cardinality estimation features into the cardinality estimation model to obtain the cardinality estimate value corresponding to the UDF query to be estimated.
[0018] The beneficial effects of this invention include: First, by jointly modeling four types of features—query encoding, data summary, UDFSig, and sampling information—it can simultaneously characterize the query structure, dataset background, static attributes of UDF operators, and sample hit evidence. Second, by constructing sampling information features through sampling hit statistics, interval information, and smoothing estimation, it can reduce the estimation fluctuation of low-selectivity queries under small sample conditions. Third, by introducing the input-output attributes, label vocabulary size, and execution attributes of UDF operators through UDFSig features, the model can distinguish the impact of different UDF operators on cardinality estimation. Fourth, by using a gradient boosting decision tree model to perform nonlinear regression on the fused features, it can characterize the combination relationship between different feature blocks, thereby obtaining cardinality estimates without fully executing the query to be estimated. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall process of the unstructured data UDF query cardinality estimation method provided in the embodiments of the present invention; Figure 2 This is a schematic diagram of UDF query parsing and query object generation provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the construction of four types of features—query code, data summary, UDFSig, and sampling information—provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the sampling information generation process provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the cardinality estimation model training and estimation process provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the unstructured data UDF query cardinality estimation device module provided in an embodiment of the present invention. Detailed Implementation
[0020] The present invention will now be described in detail with reference to the embodiments. It should be understood that the following embodiments are used to illustrate the technical solutions of the present invention, and not to limit the scope of protection of the present invention. Where there is no conflict, the technical features in the following embodiments can be combined with each other.
[0021] like Figure 1 As shown, this invention provides a method for estimating the cardinality of unstructured data UDF queries, including: (1) Obtain unstructured data set, UDF operator information and UDF query to be estimated, and parse the UDF query to be estimated into one or more clauses. Each clause includes a UDF operator and one or more labels. Multiple labels in the same clause are matched according to the OR relationship, and different clauses are matched according to the AND relationship.
[0022] (2) Generate query encoding features based on the UDF query to be estimated, wherein the query encoding features characterize the clause structure, operator structure, label structure and logical complexity of the UDF query to be estimated.
[0023] (3) Generate data summary features based on the unstructured data set and its query workload. The data summary features characterize the data domain type, record size, UDF size, tag word size and query workload overview.
[0024] (4) Generate UDFSig features based on the UDF operator information involved in the UDF query to be estimated. The UDFSig features represent the static attributes of the UDF operators and the aggregation result of the static attributes of multiple UDF operators.
[0025] (5) Determine a sample record set from the unstructured data set according to the preset sampling strategy, and perform matching on the sample record set according to the query semantics of the UDF query to be estimated to generate sampling information features.
[0026] (6) The query encoding features, data summary features, UDFSig features and sampling information features are fused into cardinality estimation features, and the cardinality estimation features are input into the cardinality estimation model to obtain the cardinality estimate value corresponding to the UDF query to be estimated.
[0027] Example 1: UDF Query Parsing In this embodiment, the unstructured data set is denoted as R, and the total number of records is denoted as N. Each record may include a record identifier, original content, and label results from several UDF operators. The original content may be text, images, video clips, or their metadata. The UDF operators may output a single label or multiple labels.
[0028] The UDF query Q to be estimated consists of one or more clauses, preferably one to three clauses. Each clause C is denoted as C=(o,L), where o is the UDF operator and L is the set of labels in the clause. For record r, if the output label set of UDF operator o on record r has a non-empty intersection with L, record r is determined to satisfy clause C; if record r satisfies all clauses in query Q, record r is determined to satisfy query Q. Thus, it is possible to simultaneously express multiple labels or relations within the same UDF and AND relations between multiple UDFs.
[0029] like Figure 2 As shown, when reading the query to be estimated, the naming of fields from different sources is normalized. For example, some query files use the `operators` and `predicates` fields to represent the list of operators and the list of predicates, while other query files use the `ml_operator_name_list` and `predicate_list` fields to represent the same meaning. These fields are unified into a list of operators and a list of predicates, and each operator is paired with its corresponding label list to form a clause.
[0030] Preferably, tags within the same clause are deduplicated and sorted to eliminate the impact of tag order on the query representation. Further, a query canonical signature is generated based on the UDF operator names of each clause and the sorted tag set. This query canonical signature can be used for query deduplication, workload statistics, model evaluation grouping, or similar query management.
[0031] Example 2: Construction of Query Encoding Features like Figure 3 As shown, the query encoding features only describe the query itself. For query Q, we first count the number of query clauses, the number of UDF operators, the number of unique UDF operators, whether it contains duplicate UDF operators, and the number of duplicate UDF operators. These features are used to characterize the range of UDFs involved in the query and whether there are multiple conditions with the same UDF.
[0032] Then, further statistical analysis was conducted on predicate tag-related characteristics, including the total number of predicate tags in the query, the mean, standard deviation, minimum and maximum number of tags in each clause, the number of single-tag clauses, the proportion of single-tag clauses, and the number of multi-tag clauses. For tags or relations within clauses, a higher number of tags usually indicates a wider OR width; for AND relations between clauses, a higher number of clauses usually indicates stronger query constraints.
[0033] Furthermore, query logic structure features are constructed, including OR complexity, AND complexity, query logic width, query logic depth, whether it is a single-clause query, whether it is a two-clause query, whether it is a three-clause query, whether all clauses are single-labeled clauses, and whether multi-labeled clauses exist. These features can map queries with different logical forms to a unified numerical representation.
[0034] To characterize the tag text and vocabulary information, we calculate the mean tag text length, standard deviation of tag text length, maximum tag text length, tag vocabulary hit rate, number of unknown tags, and proportion of unknown tags. For the implementation with a fixed maximum number of clauses, we construct the number of tags, valid markers, and operator identifiers for the 0th, 1st, and 2nd clauses, enabling the model to recognize structural differences in clauses at different positions.
[0035] Example 3: Data Summary Feature Construction like Figure 3 As shown, the data summary feature describes the overall characteristics of the dataset and query workload. The data summary is obtained from data configuration, data schema, UDF registration information, and historical workload metadata. The data summary includes at least the total number of records N, data domain identifiers, the number of UDF operators, and the total size of the tag vocabulary. Data domain identifiers can take the form of text fields, image fields, video fields, or combinations thereof.
[0036] The size of the tag vocabulary for each UDF was statistically analyzed to obtain the mean, maximum, and minimum tag vocabulary sizes for each UDF. The tag vocabulary size reflects the size of the UDF's output space; different output space sizes typically correspond to different hit probability ranges and estimation difficulties.
[0037] Statistical analysis of query workload yields metrics such as query workload size, average number of clauses, average number of predicate tags, proportion of single-clause queries, proportion of two-clause queries, and proportion of three-clause queries. This workload analysis helps the model identify common query patterns in the current application, making the cardinality estimation results more consistent with the specific system environment.
[0038] For raw unstructured records, we can calculate the mean size, standard deviation, and missing percentage of raw records. For example, the size of a text record can be the number of characters or words, the size of an image record can be the number of pixels, the file size, or the feature length, and the size of a video record can be the number of frames, the duration, or the file size.
[0039] The above features (query encoding features and data summary features) are used to characterize the size and record complexity of different datasets.
[0040] Example 4: UDFSig Feature Construction like Figure 3 As shown, UDFSig features represent the static signature information of a UDF operator, used to describe the input, output, and operational attributes of the UDF operator itself. Unlike query encoding, UDFSig does not describe the combination of labels in a specific query; unlike sampling information, UDFSig does not depend on sample hit results.
[0041] In one implementation, the UDFSig field includes the UDF label vocabulary size, whether the UDF vocabulary is closed, whether the UDF outputs multiple labels, the UDF input type, the UDF output type, the UDF version identifier, whether the UDF is deterministically executed, the runtime resource type, the UDF call cost level, the UDF model type, and the UDF parameter size. These fields can be derived from the UDF registry, model description, system configuration, or operator metadata.
[0042] When query Q contains only one clause, the UDFSig field of the corresponding UDF for that clause is directly read as the query-level UDFSig representation. When query Q contains multiple clauses, the UDFSig field of the corresponding UDF for each clause is read separately, and mean, maximum, minimum, and sum aggregations are performed on each field. For example, if the query contains three UDFs, the mean, maximum, minimum, and sum of the tag vocabulary sizes of the three UDFs can be calculated to obtain four aggregated features.
[0043] For categorical UDFSig fields, they can be mapped to discrete identifiers or one-hot encodings before aggregation. For Boolean fields, 0 and 1 can be used to represent and perform mean, maximum, minimum, and sum aggregations. For numeric fields, they can be directly aggregated. UDFSig aggregation can convert variable-number UDF clauses into fixed-dimensional model inputs.
[0044] Example 5: Construction of Sampling Information Features like Figure 3 As shown, the sampling information features are used to provide sample-level observational evidence for the query to be estimated. For example... Figure 4 As shown, a sample record set is extracted from an unstructured dataset according to a preset sampling ratio, sampling number, and random seed. The sampling ratio can be set according to estimation accuracy requirements, response time requirements, and data size, such as 0.01, 0.05, or other ratios.
[0045] For each sample, the label results of the sample record under the UDF operator involved in the query are read, and matching is performed according to the query semantics. For clause C=(o,L), if the output label set of the sample record under the UDF operator o contains any label in L, then the sample record matches clause C; for query Q, if the sample record matches all clauses in Q, then the sample record matches query Q.
[0046] Count the number of hits and the selection rate for each sampling, and multiply the selection rate by the total number of records N in the dataset to obtain the amplified sample cardinality estimate. For multiple samplings, calculate the mean, standard deviation, minimum, and maximum of the number of hits, the mean, standard deviation, minimum, and maximum of the selection rate, and the mean, standard deviation, minimum, and maximum of the amplified cardinality estimate.
[0047] To improve stability under small sample conditions, we further construct zero-hit ratio, any-hit flag, binomial proportion standard error, Wilson interval lower bound, Wilson interval upper bound, and relative interval width. These interval-type features can reflect the uncertainty of sampling estimation, enabling the model to make more robust adjustments when sample evidence is weak.
[0048] Preferably, the Laplace smoothed selectivity and Jeffreys smoothed selectivity can also be calculated, and a smoothed cardinality estimate can be obtained accordingly. For queries with low selectivity, the smoothing feature can alleviate the estimation instability caused by zero or extremely low sample hits.
[0049] In addition, clause-level sampling features are generated, including the selectivity of each clause, the number of hits for each clause, the mean, standard deviation, minimum, and maximum selectivity of the clauses. The selectivity of each clause is multiplied to obtain a combination estimate based on the clause selectivity, and the ratio and difference between the actual sample selectivity and the combination estimate are calculated to characterize the combinational relationships between clauses.
[0050] Example 6: Feature Fusion and Cardinality Estimation like Figure 5 As shown, the query encoding features, data summary features, UDFSig features, and sampling information features are concatenated to form the cardinality estimation feature vector. The query encoding features are used to describe the structure of the query to be estimated, the data summary features are used to describe the background of the dataset and the query workload, the UDFSig features are used to describe the static properties of the UDF operators involved in the query, and the sampling information features are used to describe the hit rate and estimation uncertainty on the sample records.
[0051] During the training phase, historical UDF queries with truth cardinality are used as training samples. For each training query, the four types of features are constructed following the same process as in the estimation phase, with the query truth cardinality or the transformed value of the query truth cardinality used as the supervision target.
[0052] Preferably, log(1+y) is used as the supervision target, where y is the cardinality of the query truth value.
[0053] In a preferred embodiment, the cardinality estimation model employs a gradient boosting decision tree model. The gradient boosting decision tree model comprises multiple regression trees, each of which continues to fit based on the residuals of the preceding regression trees. The final model output is obtained by a weighted combination of the outputs from the multiple regression trees. Since the cardinality of unstructured data UDF queries is influenced by the query structure, data domain, UDF attributes, and sampling hit rate, the gradient boosting decision tree model can model the nonlinear relationships between these various features.
[0054] During the estimation phase, the system generates query encoding features, data summary features, UDFSig features, and sampling information features for the UDF query to be estimated, and forms a cardinality estimation feature vector according to the feature pattern of the training phase. The cardinality estimation feature vector is then input into the trained gradient boosting decision tree model to obtain the model output. If the model output lies in the logarithmic space, an inverse exponential transform is performed on the model output; if the result of the inverse transform is less than 0, it is truncated to 0 and used as the cardinality estimate of the UDF query to be estimated.
[0055] In another implementation, a sample mean model, a Laplace smoothing sampling model, a Jeffreys smoothing sampling model, and a combined estimation model based on clause selectivity can be used as contrastive estimation models. The sample mean model obtains a cardinality estimate based on the mean of the sample selectivity and the total number of records in the dataset; the Laplace smoothing sampling model and the Jeffreys smoothing sampling model obtain cardinality estimates based on the smoothed selectivity; and the combined estimation model based on clause selectivity obtains a cardinality estimate based on the combined result of the selectivity of each clause sample.
[0056] The above model can be used to verify the effectiveness of the sampled information features, and can also be used as part of the input features for the gradient boosting decision tree model.
[0057] Example 7: Device Structure like Figure 6 As shown, the present invention also provides a cardinality estimation device for unstructured data UDF queries. The device includes a query parsing module, a query encoding module, a data summarization module, a UDFsig module, a sampling execution module, a feature fusion module, a cardinality estimation module, and a result output module.
[0058] The query parsing module reads the UDF query to be estimated and generates clause objects; the query encoding module generates query encoding features based on the query objects; the data summarization module generates data summary features; the UDFSig module reads the static signatures of the UDFs involved in the query and performs aggregation; the sampling execution module extracts sample records and performs query matching; the feature fusion module unifies the four types of features into model input; the cardinality estimation module outputs cardinality estimates; and the result output module provides the estimates to the query optimizer, execution plan selector, or upper-layer applications.
[0059] Example 8: Electronic Devices and Storage Media The present invention also provides an electronic device, including one or more processors, a memory, and a program stored in the memory. When the program is executed by the processor, it implements the above-described method for estimating the cardinality of unstructured data UDF queries. The electronic device can be a server, workstation, database node, query optimization node, or computing instance in a cloud computing platform.
[0060] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the above-described method for estimating the cardinality of unstructured data UDF queries. The computer-readable storage medium may be a hard disk, a solid-state drive, a read-only memory, a random access memory, flash memory, or other media capable of storing program code.
[0061] The above embodiments are merely preferred embodiments of the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the number of query clauses, UDF types, sampling ratios, feature fields, aggregation methods, or model types without departing from the core ideas of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for estimating the cardinality of unstructured data UDF queries, characterized in that, Includes the following steps: (1) Obtain unstructured data set, UDF operator information and UDF query to be estimated, parse the UDF query to be estimated into one or more clauses, each clause includes a UDF operator and one or more labels, multiple labels in the same clause are matched according to the OR relation, and different clauses are matched according to the AND relation; (2) Generate query encoding features based on the UDF query to be estimated, wherein the query encoding features characterize the clause structure, operator structure, label structure and logical complexity of the UDF query to be estimated; (3) Generate data summary features based on the unstructured data set and its query workload, wherein the data summary features characterize the data domain type, record size, UDF size, tag vocabulary size and query workload overview; (4) Generate UDFSig features based on the UDF operator information involved in the UDF query to be estimated. The UDFSig features represent the static attributes of the UDF operators and the aggregation result of the static attributes of multiple UDF operators. (5) Determine a sample record set from the unstructured data set according to the preset sampling strategy, and perform matching on the sample record set according to the query semantics of the UDF query to be estimated to generate sampling information features; (6) The query encoding features, data summary features, UDFSig features and sampling information features are fused into cardinality estimation features, and the cardinality estimation features are input into the cardinality estimation model to obtain the cardinality estimate value corresponding to the UDF query to be estimated.
2. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, The unstructured data set includes one or more combinations of text data, image data, and video data. The UDF operator is used to perform classification, detection, recognition, extraction, or label generation on the unstructured records and output one or more label results.
3. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, In step (1), the UDF query to be estimated is parsed into one or more clauses, including: reading the list of operators and the list of predicates in the UDF query to be estimated, pairing each UDF operator with the corresponding set of predicate labels to form a clause, deduplicating and sorting the labels in the same clause, and generating a query specification signature based on the UDF operator name and the sorted set of labels in the clause.
4. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, The query encoding features include at least one of the following: number of query clauses, number of UDF operators, number of unique UDF operators, information on repeated operators, number of predicate tags, statistics on the number of tags in each clause, information on single-tag clauses, information on multi-tag clauses, OR relation complexity, AND relation complexity, query logic width, query logic depth, query shape, statistics on tag text length, tag vocabulary hit information, information on unknown tags, valid markers for each clause, and operator identifiers corresponding to each clause.
5. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, The data summary features include at least one of the following: total number of records in the dataset, data domain identifier, number of UDF operators, total size of the tag vocabulary, statistics on the size of each UDF tag vocabulary, query workload size, average number of clauses per workload, average number of predicates per workload, proportion of queries with different number of clauses, statistics on the size of original records, and proportion of missing original records.
6. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, The UDFSig features include at least one of the following: UDF label vocabulary size, whether the UDF vocabulary is closed, whether the UDF outputs multiple labels, UDF input type, UDF output type, whether the UDF is executed deterministically, UDF call cost information, UDF model type, and UDF complexity information; for a UDF query to be estimated that contains multiple clauses, the UDFSig features of the multiple UDF operators involved in the UDF query to be estimated are aggregated by mean, maximum, minimum, or summation.
7. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, The sampling information features are generated through the following preset sampling strategy: one or more sample record sets are determined according to the preset sampling ratio and sampling number; for each sample record, the label result under the UDF operator involved in the UDF query to be estimated is read; labeling or matching is performed on each clause; the matching results of multiple clauses are performed and combined; and the sample hit number, sample selection rate and amplified sample base are statistically analyzed based on the sampling matching results.
8. The method for estimating the cardinality of unstructured data UDF queries according to claim 7, characterized in that, The sampling information features include at least one of the following: sampling ratio, sample size, number of sampling repetitions, sample hit count statistics, sample selection rate statistics, zero hit ratio, any hit identifier, binomial proportional standard error, confidence interval information, relative interval width, smoothed selection rate, smoothed cardinality estimate, selection rate of each clause, number of hits of each clause, combined estimate based on clause selection rate, ratio of actual sample selection rate to combined estimate, and difference between actual sample selection rate and combined estimate.
9. The method for estimating the cardinality of unstructured data UDF queries according to claim 1, characterized in that, The cardinality estimation model is a gradient boosting decision tree model. The gradient boosting decision tree model takes the cardinality estimation features obtained by fusing the query encoding features, data summary features, UDFSig features and sampling information features as input, and outputs the cardinality estimate of the UDF query to be estimated. When training the gradient boosting decision tree model, the cardinality of the query truth or the transformed value of the query truth cardinality is used as the supervision target, and the cardinality estimation feature is used as the model input. When estimating the UDF query to be estimated, the output of the gradient boosting decision tree model is converted into a non-negative value and used as the cardinality estimate.
10. A device for estimating the cardinality of unstructured data UDF queries, characterized in that, include: The query parsing module is used to obtain unstructured data sets, UDF operator information, and UDF queries to be estimated. It parses the UDF queries to be estimated into one or more clauses. Each clause includes a UDF operator and one or more labels. Multiple labels within the same clause are matched according to an OR relationship, and different clauses are matched according to an AND relationship. The query encoding module is used to generate query encoding features based on the UDF query to be estimated. The query encoding features characterize the clause structure, operator structure, label structure and logical complexity of the UDF query to be estimated. The data summarization module is used to generate data summary features based on the unstructured data set and its query workload. The data summary features characterize the data domain type, record size, UDF size, tag vocabulary size, and query workload overview. The UDFSig module is used to generate UDFSig features based on the UDF operator information involved in the UDF query to be estimated. The UDFSig features represent the static attributes of the UDF operators and the aggregation result of the static attributes of multiple UDF operators. The sampling execution module is used to determine a sample record set from the unstructured data set according to a preset sampling strategy, perform matching on the sample record set according to the query semantics of the UDF query to be estimated, and generate sampling information features. The feature fusion module is used to fuse the query encoding features, data summary features, UDFSig features, and sampling information features into cardinality estimation features; The cardinality estimation module is used to input the cardinality estimation features into the cardinality estimation model to obtain the cardinality estimate value corresponding to the UDF query to be estimated.