A query plan result dataset fast generation method
By dividing the data range in the database system and building an overhead prediction model, the problems of high cost and low accuracy in generating query plan evaluation result datasets are solved, enabling fast and accurate generation of simulated data and supporting efficient training and optimization of learning-based query optimizers.
Patent Information
- Application Number
- CN202511740589.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-11-25
AI Technical Summary
Existing technologies suffer from high computational costs and poor estimation accuracy when generating query plan evaluation result datasets. In particular, when dealing with the nonlinear and uncertain overhead of machine learning models, traditional methods cannot quickly generate massive amounts of evaluation data, hindering the development of learning-based query optimizers.
By dividing the full dataset into multiple sample data intervals, recording the total time cost and data characteristics of each interval, constructing randomly combined data intervals, and using a cost prediction model to estimate query execution results, a simulated dataset is generated, including specific feature fitting models for text, image, and video data, adapting to different system configurations.
It enables the low-cost and efficient generation of large amounts of simulated data that can be used to evaluate query plans, improves estimation accuracy, supports rapid training and optimization of learning optimizers, and is suitable for query optimization scenarios of various machine learning models.
Smart Images

Figure CN121188092B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of database technology and query optimization, and in particular to a method for rapidly generating query plan result datasets. Background Technology
[0002] In modern data analytics and database systems, complex queries often involve calls to built-in or user-defined machine learning models. The query optimizer needs to select the most efficient execution plan (query plan) for such queries. The key to evaluating the merits of different query plans lies in the ability to quickly and accurately predict their execution costs (such as time costs) and results on a specific data distribution.
[0003] Traditional methods typically require either actually executing the query plan or estimating based on coarse statistical information, which has significant drawbacks: 1) The actual execution method is computationally very expensive, requiring expensive machine learning models to process large amounts of data for each evaluation, making it impossible to quickly generate massive amounts of evaluation data; 2) Traditional statistical information methods cannot effectively handle the nonlinearity and uncertainty overhead of machine learning models, which are "black box" functions, resulting in poor estimation accuracy.
[0004] Furthermore, while machine learning-based query optimizers have shown the potential to surpass traditional optimizers in recent years, their performance heavily relies on large amounts of high-quality training data. This data needs to include the execution overhead and results of different query plans across diverse data distributions. However, as mentioned earlier, collecting this data through real-world executions is costly and impractical, becoming a major bottleneck hindering the development and implementation of learning-based optimizers.
[0005] Therefore, there is an urgent need in this field for a method that can generate result datasets for query plan evaluation in a low-cost, efficient and accurate manner. Summary of the Invention
[0006] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for rapidly generating query plan result datasets, which can quickly generate a large amount of simulated data that can be used to evaluate different query plans with extremely low computational cost.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0008] A method for quickly generating query plan result datasets includes the following steps:
[0009] (1) Sample data collection steps: Divide the full dataset into multiple sample data intervals, actually execute the target query on multiple sample data intervals, record the total time cost of each sample data interval, the data characteristics of the interval, and the processing result of each data; the data characteristics are used to characterize the overall attributes of the data in the interval;
[0010] (2) Data interval construction steps: Based on the full dataset, multiple data intervals with different data distributions are constructed by random combination;
[0011] (3) Query plan execution result estimation step: For each constructed data interval, estimate the execution result of the target query on the data interval; wherein, the execution result includes the total processing result and the total time cost; the total processing result is obtained by aggregating the model processing results of all data in the data interval; the total time cost is estimated by inputting the data characteristics of the current data interval into an cost prediction model; the cost prediction model is obtained based on the total time cost and data characteristics of multiple sample data intervals collected in the sample data collection step.
[0012] Furthermore, in the sample data acquisition step, the total time overhead and data characteristics are acquired under at least one system configuration.
[0013] Furthermore, the data preprocessing step also includes an overhead prediction model generation step: when the data is text data, a first overhead prediction model is generated by fitting the total time cost of multiple sample data intervals and the corresponding text length.
[0014] When the data is image data, a second cost prediction model is generated by fitting the total time cost of multiple sample data intervals and the corresponding image features. The image features include one or more of the following: image resolution, image file size, number of color channels, and content complexity features, but at least the image resolution is included. The content complexity features include edge complexity and the initially estimated number of targets.
[0015] When the data is video data, the video data is processed according to a preset sampling frequency to obtain the effective number of processed frames. A model is then built based on the effective number of processed frames and the average cost per frame, and dynamic features are introduced as correction factors to generate a third cost prediction model. The average cost per frame is determined according to the second cost prediction model. The dynamic features include the overall motion intensity of the video and the scene switching frequency.
[0016] Furthermore, the data preprocessing step is executed under various different system configurations, and for each system configuration, the total time cost and data characteristics of each sample data interval are recorded, and a corresponding cost prediction model is generated.
[0017] Furthermore, the query plan execution result dataset generated by the method is used to train a learning-based query optimizer, which takes query plan characteristics and / or data range characteristics as input and the estimated execution cost and results as output.
[0018] Furthermore, in the data interval construction step, the random combination method is random sampling with replacement, random sampling without replacement, or sampling that conforms to the data distribution pattern.
[0019] Furthermore, the data distribution includes the normal distribution.
[0020] The present invention also provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs stored in the memory, the one or more computer programs including instructions that, when executed by the electronic device, cause the electronic device to perform the above-described method for rapidly generating a query plan result dataset.
[0021] The present invention also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to execute the above-described method for rapidly generating a query plan result dataset.
[0022] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for rapidly generating a query plan result dataset.
[0023] Compared with the prior art, the beneficial effects of the present invention include:
[0024] Significantly improves efficiency: By bringing forward and solidifying the expensive model computation overhead, the subsequent data generation stage is almost costless, thus enabling the generation of massive amounts of data in a short period of time.
[0025] Ensuring data validity: The generated data is based on the processing results of the real model on real data. Its data distribution and result characteristics are highly similar to those of real query execution, and can be used for effective query plan evaluation.
[0026] High versatility: This method is decoupled from specific machine learning model types and database systems, and can be widely applied to various query optimization scenarios involving model inference.
[0027] Easy to use: The solution has a clear process and is easy to integrate into existing database query optimizers.
[0028] Directly supports the learning optimizer: It can automatically and cost-effectively generate massive amounts of high-quality training sample pairs (query plan, data environment, execution overhead, results), providing a key data foundation for the effective training of the learning query optimizer.
[0029] High estimation accuracy and wide scenario coverage: By introducing data features (such as text length) and system configuration as estimation factors, the time cost prediction is more accurate, and the execution performance of the same query under different hardware environments can be easily simulated. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart illustrating the method for rapidly generating query plan result datasets provided in this embodiment of the invention.
[0032] Figure 2 This is a flowchart illustrating the steps for estimating the execution result of a query plan provided in an embodiment of the present invention.
[0033] Figure 3 This is a schematic diagram of the structure of an electronic device applicable to this method, provided in an embodiment of the present invention. Detailed Implementation
[0034] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.
[0035] The present invention provides a method for rapidly generating query plan result datasets, comprising the following steps:
[0036] (1) Sample data collection steps: Divide the full dataset into multiple sample intervals, actually execute the target query on multiple sample data intervals, record the total time cost of each sample data interval, the data characteristics of the interval, and the processing result of each data; wherein, the target query includes the query that calls the machine learning model;
[0037] The data features are used to characterize the overall attributes of the data within the interval; if the data is text data, the data features include at least the average text length; if the data is image data, the data features include at least the average image resolution; if the data is video data, the data features include at least the average number of video frames or the duration of the video.
[0038] Specifically, the model processing results include classification labels, regression values, etc.; the total time cost and data features are collected under at least one system configuration.
[0039] In one embodiment, the sample data acquisition step further includes an overhead prediction model generation step: when the data is text data, a first overhead prediction model is generated by fitting the total time cost of multiple sample data intervals and the corresponding text length.
[0040] When the data is image data, a second cost prediction model is generated by fitting the total time cost of multiple sample data intervals and the corresponding image features. The image features include one or more of the following: image resolution, image file size, number of color channels, and content complexity features, but at least image resolution is included. The content complexity features include edge complexity and the initially estimated number of targets.
[0041] When the data is video data, the video data is processed according to a preset sampling frequency to obtain the effective number of processed frames. A model is then built based on the effective number of processed frames and the average cost per frame, and dynamic features are introduced as correction factors to generate a third cost prediction model. The average cost per frame is determined according to the second cost prediction model. The dynamic features include the overall motion intensity of the video and the scene switching frequency.
[0042] The above cost prediction model can be generated by fitting linear regression or multiple regression, or by training a deep learning model for fitting.
[0043] In one embodiment, the data preprocessing step is performed under a variety of different system configurations, and for each system configuration, the total time cost and data characteristics of each sample data interval are recorded, and a corresponding cost prediction model is generated.
[0044] (2) Data interval construction steps: Based on the full dataset, multiple data intervals with different data distributions are constructed by random combination;
[0045] Specifically, in the data interval construction step, the random combination method is random sampling with replacement, random sampling without replacement, or sampling that conforms to the data distribution pattern (such as normal distribution).
[0046] Specifically, this step aims to simulate and generate a large number of different query data environments. It starts with the metadata (or index) of the full dataset and performs numerous random samplings using a random algorithm (such as the Monte Carlo method). Each sampling generates a data interval description, which may include: randomly selected start and end rows, a randomly selected set of data IDs, or a subset of data conforming to a certain random distribution (such as a normal distribution). These data intervals represent the various different data distributions and scales that the query might encounter.
[0047] (3) Query plan execution result estimation step: For each constructed data interval, estimate the execution result of the target query on the data interval; wherein, the execution result includes the total processing result and the total time cost; the total processing result is obtained by aggregating the model processing results of all data in the data interval; the total time cost is estimated by inputting the data characteristics of the current data interval into an cost prediction model; the cost prediction model is obtained based on the total time cost and data characteristics of multiple sample data intervals collected in the sample data collection step.
[0048] Specifically, for the overall processing result, there is no need to recalculate. Just find the model processing results corresponding to all data in the data interval from the results recorded in step (1), and then perform aggregation operations on these results according to the requirements of the query statement (such as summation, average, group counting, etc.) to obtain the estimated execution result of the query in this interval.
[0049] For example, query plan evaluation and optimizer training data generation based on user review sentiment analysis.
[0050] 1. System Environment and Data Preparation
[0051] Database: ecommerce_db
[0052] The data table, product_reviews, contains tens of millions of user reviews of various products.
[0053] Key fields: review_id (primary key), product_id (product ID), review_text (review text), rating (rating 1-5), create_time (review time).
[0054] Machine learning model: A trained sentiment analysis model ML_Sentiment, which takes review text as input and outputs a sentiment label (e.g., 'positive', 'negative', 'neutral') and a sentiment confidence score (a floating-point number between 0 and 1).
[0055] Target query: The business analyst wants to run the following query to analyze the trend of positive review rates for different products:
[0056] SQL
[0057] SELECT product_id,
[0058] COUNT(*) as total_reviews,
[0059] SUM(CASE WHEN ML_Sentiment(review_text) = 'Positive' THEN 1 ELSE 0 END)as positive_reviews,
[0060] (SUM(CASE WHEN ML_Sentiment(review_text) = 'Positive' THEN 1 ELSE 0 END)* 1.0 / COUNT(*)) as positive_rate
[0061] FROM product_reviews
[0062] WHERE create_time BETWEEN '2023-01-01' AND '2023-12-31'
[0063] GROUP BY product_id
[0064] HAVING COUNT(*)>100;
[0065] 2. See also Figure 1 Steps of applying the method of the present invention
[0066] S101: Data Preprocessing Stage
[0067] At this stage, the system receives the target query and the full dataset.
[0068] Query parsing: The system parses the query and locates the part that calls the machine learning model ML_Sentiment.
[0069] Sample Interval Selection and Actual Execution: The system randomly selects or divides the entire dataset into M representative sample data intervals (e.g., M=100) based on a certain strategy (such as uniform distribution of data volume). For each selected sample data interval Sample_Interval_i, the system performs the following operations:
[0070] Actual query execution: The target query (or the part that calls ML_Sentiment) is actually run on this interval, and its total execution time is precisely measured as Total_Cost_i.
[0071] Record interval features: Simultaneously, record the feature vector Feature_i for this sample interval. This vector serves as the input to the subsequent prediction model and must contain:
[0072] data_count_i: The total number of data entries contained in this interval.
[0073] avg_text_length_i: The average character length of all comment texts within this range.
[0074] Storage of processing results: The system will persistently store the model processing results (such as sentiment_label and confidence_score) of all data within the interval to a preprocessing result table precomputed_sentiment for subsequent result estimation.
[0075] Multi-environment sample collection: To simulate different deployment environments, the above sample collection process was executed under various system configurations (or extrapolated using performance models). For example:
[0076] Config_A: 2 CPU Cores
[0077] Config_B: 4 CPU Cores
[0078] Config_C: 8 CPU Cores
[0079] For each configuration Config_X, collect a separate sample set { (Feature_i, Total_Cost_i)}_X.
[0080] Construct an overhead prediction model:
[0081] For each system configuration Config_X, the system uses its corresponding sample set { (Feature_i, Total_Cost_i)}_X, takes the interval feature Feature_i as input, and the measured total time cost Total_Cost_i as output label, and performs fitting through a regression algorithm (such as linear regression) to generate an independent cost prediction model Cost_Model_X(Feature).
[0082] Specifically, for the system configuration Config_B, its model form is: Cost_B(L) = α * L + β.
[0083] Where L is the average length of the input text, and Cost_B(L) is the predicted processing time per line. The regression coefficients α and intercept β are the best-fit parameters obtained by minimizing the error between the predicted values and the actual time costs for all samples.
[0084] The model training process is executed only once in this stage, and the trained model parameters will be persistently stored for repeated use in subsequent estimation stages.
[0085] S102: Data Range Construction Phase
[0086] At this stage, the system aims to simulate and generate a large number of different query data environments. Starting with the metadata of the full dataset, it performs numerous random samplings using a random algorithm (such as the Monte Carlo method). Each sampling generates a data interval descriptor, which contains the data volume, a logical reference to the data ID set, and the data characteristics of that interval (such as the average text length L_avg). This process is repeated thousands of times to generate a large number of data intervals {Interval_1, Interval_2, ...}.
[0087] S103: Query plan execution result estimation stage
[0088] See Figure 2 This stage involves a rapid estimation of each generated data interval.
[0089] S201: For the "processing result", the system finds the sentiment_label corresponding to all data in the data range from the results stored in S101, and then performs aggregation operation according to the requirements of the query statement to obtain the estimated execution result.
[0090] S202: For "time overhead," perform a refined estimate. For example, for the data interval Interval_k (containing N_k = 500,000 comments, with an average text length L_avg_k = 350 characters), estimate its overhead on Config_B (4 CPU Cores):
[0091] The linear cost prediction model Cost_B(L) = α * L + β fitted for Config_B in stage S101 is invoked. Then the total time cost can be estimated as: Total_Model_Cost_k = N_k * (α * L_avg_k + β).
[0092] 3. Generate data for training the learning optimizer.
[0093] The dataset generated by this invention ultimately takes the form of a series of records in the following format:
[0094] (Query_Plan_Fingerprint, Data_Interval_Descriptor, System_Config,Estimated_Cost, Estimated_Result)
[0095] Query_Plan_Fingerprint: Feature vector of the query plan (such as the join order, index type, aggregation algorithm, etc.).
[0096] Data_Interval_Descriptor: Characteristics of the data interval (such as data volume N_k, average text length L_avg_k, proportion of positive comments, etc.).
[0097] System_Config: System configuration identifier (e.g., Config_B).
[0098] Estimated_Cost and Estimated_Result: The execution cost and result estimated by this invention.
[0099] These records constitute a massive training set covering a wide range of scenarios. A learning-based query optimizer (such as a deep neural network) can use this training set for supervised learning. After training, the optimizer can take a new query plan and new data environment features as input, and directly and quickly predict the execution cost and result of the plan, thus selecting the optimal plan within milliseconds.
[0100] Effect Comparison
[0101] Traditional method: In order to obtain evaluation data for thousands of data intervals, it is necessary to actually call the ML_Sentiment model interval average data volume * thousands of times, which is a huge overhead.
[0102] The method of this invention calls the model only a relatively small number of times in the first step, and the subsequent thousands of evaluations are almost cost-free. The total time to generate massive training data is reduced to the "minute level," paving the way for the practical application of learning-based optimizers.
[0103] Extended Implementation: Cost Estimation for Image and Video Data
[0104] The technical solution of this invention is also applicable to query scenarios involving unstructured data such as images and videos.
[0105] For image data, in the data preprocessing step (S101), in addition to recording the processing results, the following features related to processing overhead also need to be extracted and recorded:
[0106] Core feature: Image resolution (width W and height H).
[0107] Optional features: image file size, number of color channels, and content complexity features estimated by a lightweight model (e.g., by quickly scanning the image with a lightweight object detection model (e.g., YOLO-Tiny) that is computationally significantly more efficient than the machine learning model in the query to estimate the number of identifiable objects and edge complexity it contains).
[0108] Subsequently, using the aforementioned features (including at least the resolution W * H) and the measured processing time, a cost prediction model for the image model is fitted through linear regression or multiple regression analysis. Its form can be: Cost_picture = α * (W * H) + β * Estimated_Objects + γ.
[0109] For video data, the preprocessing step (S101) can be expanded as follows:
[0110] Key features: total number of frames F, frame rate (fps), and resolution per frame. The duration of the video can be calculated using T = F / fps.
[0111] Dynamic features: the overall motion intensity and scene transition frequency of the video. These features are determined as follows:
[0112] Overall motion intensity: quantified by calculating the average optical flow amplitude between consecutive frames in a video sequence. Specifically, the Farneback or Lucas-Kanade optical flow algorithms can be used for calculation.
[0113] Scene switching frequency: This is calculated by dividing the number of scene switches in the video by the total video duration. Scene switching is detected by calculating the difference in color histograms between consecutive frames. A switch is defined as a difference exceeding a preset threshold. This preset threshold is dynamically determined by statistically analyzing the distribution of differences between all consecutive frames in the video. Specifically, the mean μ and standard deviation σ of the differences between all consecutive frames are calculated, and the threshold is set as T = μ + n * σ, where n is an empirical coefficient (usually between 3 and 5). A scene switch is defined as a difference between any two frames exceeding this threshold.
[0114] Considering performance, not every frame is typically processed; therefore, a specific sampling strategy (e.g., sampling 1 frame per second) needs to be defined. Ultimately, the total video cost can be modeled based on the effective number of processed frames F_effective and the average cost per frame Cost_frame (which can be estimated by the image cost model described above), with dynamic features introduced as a correction factor. The formula is: Cost_video = F_effective * Cost_frame * (1 + λ * Motion_Intensity), where λ is a weighting coefficient determined through regression analysis, and Motion_Intensity is the scene switching frequency or overall motion intensity.
[0115] By introducing the aforementioned features closely related to data types, this invention can construct a more accurate execution time cost estimation model for queries that include complex machine learning models such as image recognition and video content analysis, demonstrating the versatility and strong adaptability of this method.
[0116] See Figure 3 An electronic device 300 of the present invention includes, but is not limited to: one or more processors 301; a memory 302 for storing one or more programs; when the one or more programs are executed by the one or more processors 301, the one or more processors 301 implement a method for rapidly generating a query plan result dataset as described above. The electronic device 300 may also include an input / output interface 303 for data interaction with external devices.
[0117] It should be noted that, in addition to Figure 3 In addition to the memory and processor shown, electronic devices may include other hardware depending on their actual functions, which will not be elaborated further.
[0118] The present invention also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to execute the above-described method for rapidly generating a query plan result dataset.
[0119] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for rapidly generating a query plan result dataset.
[0120] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0122] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0123] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions should all be covered within the scope of protection of the present invention.
Claims
1. A method for fast generation of query plan result dataset, characterized in that, The method comprises the following steps: (1) sample data collection step: divide the full data set into multiple sample data intervals, actually execute the target query on the multiple sample data intervals, record the total time cost of each sample data interval and the data characteristics of the interval and the processing result of each data; (2) data interval construction step: based on the full data set, a plurality of data intervals with different data distribution are constructed by random combination; (3) query plan execution result estimation step: for each data interval constructed, the execution result of the target query on the data interval is estimated; wherein, the execution result includes total processing result and total time cost; the total processing result is obtained by aggregating the machine learning model processing result of all data in the data interval; the total time cost is obtained by inputting the data characteristics of the current data interval into an overhead prediction model; the overhead prediction model is obtained based on the total time cost and data characteristics of the multiple sample data intervals collected in the sample data collection step.
2. The method of claim 1, wherein, In the sample data collection step, the total time cost and data characteristics are collected under at least one system configuration.
3. The method of claim 1, wherein, The sample data collection step further comprises an overhead prediction model generation step: when the data is text data, a first overhead prediction model is generated by fitting the total time cost and corresponding text length of the multiple sample data intervals; When the data is picture data, a second overhead prediction model is generated by fitting the total time cost and corresponding picture characteristics of the multiple sample data intervals; The picture characteristics include one or more of the resolution of the image, the file size of the image, the number of color channels, and the content complexity characteristics, but at least include the resolution of the image; the content complexity characteristics include edge complexity, initial estimated target number; When the data is video data, the video data is processed according to a preset sampling frequency to obtain the effective processing frame number, a third overhead prediction model is generated based on the effective processing frame number and the single frame average overhead, and the dynamic characteristics are introduced as a correction factor; wherein, the single frame average overhead is determined according to the second overhead prediction model; the dynamic characteristics include the overall motion intensity of the video and the scene switching frequency.
4. The method of claim 1, wherein, Further comprising a data preprocessing step, the data preprocessing step is executed under multiple different system configurations, and the total time cost of each sample data interval and the data characteristics of the interval corresponding to each system configuration are recorded respectively, and the corresponding overhead prediction model is generated.
5. The method of claim 1, wherein, The query plan execution result data set generated by the method is used to train a learning type query optimizer, which takes query plan characteristics and / or data interval characteristics as input and estimated execution overhead and result as output.
6. The method of claim 1, wherein, In the data interval construction step, the random combination method is random sampling with replacement, random sampling without replacement, or sampling according to the data distribution rule.
7. The method of claim 6, wherein, The data distribution includes normal distribution.
8. An electronic device, comprising: Comprise: One or more processors; Memory; and one or more computer programs stored in the memory, the one or more computer programs comprising instructions that, when executed by the electronic device, cause the electronic device to perform a method of fast generation of query plan result dataset as claimed in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which, when running on a computer, causes the computer to perform a method of fast generation of query plan result dataset as claimed in any one of claims 1-7.
10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements a method of fast generation of query plan result dataset as claimed in any one of claims 1-7.
Citation Information
Patent Citations
Query optimization method, statistical information prediction model training method and equipment
CN119066088A
Time-factored performance prediction
US20190378048A1