File Merging Method, Device, and Storage Medium for Real-Time Data Lake

By obtaining the query task characteristics and file attribute information of the real-time data lake table, predicting resource overhead, and scientifically and reasonably determining the merge timing, the problem of small and medium-file merging in the real-time data lake table affects query performance, and achieving cost optimization and performance improvement.

CN119052226BActive Publication Date: 2025-07-22BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411155688.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-07-22
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

In the prior art, the fixed-cycle merging method of small files in real-time data lake tables will affect query performance or increase unnecessary resource overhead, resulting in increased costs.

Method used

By obtaining the query task feature information of the real-time data lake table and the attribute information of the files to be merged, the resource overhead in the case of non-merging and merging is predicted, the merge time is scientifically and reasonably determined, and the solution with the least overhead is selected for file merging.

Benefits of technology

Reduces user costs, improves query performance, reduces resource waste, and optimizes the timing selection of file merging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052226B_ABST
    Figure CN119052226B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, device, and storage medium for file merging in a real-time data lake. The method includes: obtaining feature information of a query task for a real-time data lake table and attribute information of files to be merged in the real-time data lake table; determining whether to start a merging task based on the feature information of the query task and the attribute information of the files to be merged; and if it is determined to start the merging task, merging the files to be merged. Embodiments of the present disclosure can scientifically and reasonably determine the file merging timing for a real-time data lake table, reduce user costs, and improve query performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer and network communication technologies, and in particular, to a method, device, and storage medium for file merging in a real-time data lake. Background Art

[0002] A real-time data lake is a data storage and processing architecture designed to address the management and analysis of massive amounts of data. Based on distributed storage and computing technologies, it can receive and process large amounts of data generated in real time.

[0003] In the prior art, for small files in a real-time data lake table, file merging is usually performed at fixed intervals, but this may affect query performance or impose unnecessary resource overhead on users, increasing costs. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, device, and storage medium for file merging in a real-time data lake to determine the timing of file merging for a real-time data lake table, reduce user costs, and improve query performance.

[0005] In a first aspect, embodiments of the present disclosure provide a method for file merging in a real-time data lake, including:

[0006] Obtaining characteristic information of a query task for a real-time data lake table and attribute information of files to be merged in the real-time data lake table;

[0007] Judging whether to start a merging task according to the characteristic information of the query task and the attribute information of the files to be merged;

[0008] If it is determined to start the merging task, merging the files to be merged.

[0009] In a second aspect, embodiments of the present disclosure provide a device for file merging in a real-time data lake, including:

[0010] An obtaining unit, configured to obtain characteristic information of a query task for a real-time data lake table and attribute information of files to be merged in the real-time data lake table;

[0011] A judging unit, configured to judge whether to start a merging task according to the characteristic information of the query task and the attribute information of the files to be merged;

[0012] An execution unit, configured to merge the files to be merged if it is determined to start the merging task.

[0013] In a third aspect, embodiments of the present disclosure provide an electronic device, including: at least one processor and a memory;

[0014] The memory stores computer-executable instructions;

[0015] At least one processor executes computer-executable instructions stored in a memory, such that the at least one processor executes a file merging method for a real-time data lake as described in the first aspect above and various possible designs of the first aspect.

[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement a file merging method for a real-time data lake as described in the first aspect above and various possible designs of the first aspect.

[0017] In a fifth aspect, an embodiment of the present disclosure provides a computer program product including computer-executable instructions, which, when executed by a processor, implement a file merging method for a real-time data lake as described in the first aspect above and various possible designs of the first aspect.

[0018] The file merging method, device, and storage medium for a real-time data lake provided by the embodiments of the present disclosure obtain the characteristic information of a query task of a real-time data lake table and the attribute information of files to be merged in the real-time data lake table, predict the overhead of the query task in the case of not merging the files to be merged, and the overhead of the query task and the merging task in the case of merging the files to be merged, and then select a scheme with less overhead. The embodiments of the present disclosure can scientifically and reasonably determine the file merging timing for the real-time data lake table, reduce user costs, and improve query performance. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the drawings required for description in the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 An example diagram of file merging for a real-time data lake in the prior art;

[0021] Figure 2 A schematic flowchart of a file merging method for a real-time data lake provided by an embodiment of the present disclosure;

[0022] Figure 3 A schematic flowchart of a file merging method for a real-time data lake provided by another embodiment of the present disclosure;

[0023] Figure 4 An example diagram of file merging provided by an embodiment of the present disclosure;

[0024] Figure 5 A structural block diagram of a file merging device for a real-time data lake provided by an embodiment of the present disclosure;

[0025] Figure 6 Schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. Detailed implementation manners

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0027] A real-time data lake is a data storage and processing architecture designed to address the management and analysis of massive data. Based on distributed storage and computing technologies, it can receive and process a large amount of data generated in real time.

[0028] In the scenario of a real-time data lake, the data written to a real-time data lake table is usually first written into memory and then committed to disk incrementally, that is, the incremental data is stored in disk files incrementally. For example, if the data of a certain target object in the data lake table is modified for the first time, the data after the first modification of the target object is written into memory and then committed to disk, and is saved in the incremental file 1 on the disk; then the data of the target object is modified for the second time, the data after the second modification of the target object is written into memory and then committed to disk, and is saved in the incremental file 2 on the disk. As time goes by, there are more and more incremental files, and the data volume of the incremental files is generally small. In particular, the incremental files and the stock files may include different versions of data of the same object. If these incremental files are not merged with the stock files in time (Compact), it will affect the performance of our subsequent real-time query operations and lead to performance degradation.

[0029] In the prior art, for the incremental files in a real-time data lake table, an asynchronous task is usually configured to perform file merging at a fixed period, as Figure 1 shown, in the prior art, for the stock files and incremental files to be merged in a real-time data lake table: the files to be merged 1, 2, 3... n, an asynchronous task is used to merge them according to a preset fixed time period.

[0030] However, performing file merging at a fixed period through an asynchronous task may affect query performance or bring unnecessary resource overhead to users, increasing costs. The specific reasons are as follows:

[0031] 1) The data stream writing in the real-time data lake is not uniform. In scenarios where the data stream writing speed is relatively high, if file merging is performed at fixed intervals and the fixed interval is too long, it may lead to an excessive number of small files piling up, affecting query performance;

[0032] 2) In scenarios where the data stream writing speed is relatively low, if file merging is performed at fixed intervals and the fixed interval is too short, it may result in not many small files needing to be merged each time, and the merging tasks being too frequent. Moreover, each merging task also incurs resource overhead, increasing unnecessary costs;

[0033] 3) Performing file merging can improve query performance, reduce the resource overhead of query tasks, and lower the cost of query tasks; however, at the same time, file merging also increases resource overhead and the cost of merging tasks. For customers, how to optimize costs to the greatest extent is an urgent problem to be solved.

[0034] To address the above problems, the present disclosure provides a file merging method for a real-time data lake. By using the characteristic information of query tasks and the attribute information of files to be merged, it is determined whether to initiate a merging task for the files to be merged. The judgment basis is the overhead of query tasks in the case of predicting not to merge the files to be merged, and the overhead of query tasks and merging tasks in the case of merging the files to be merged. Then, the scheme with lower overhead is selected, thereby reducing user costs and improving query performance.

[0035] The following will introduce in detail the file merging method for the real-time data lake of the present disclosure with specific embodiments.

[0036] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the file merging method for the real-time data lake provided by an embodiment of the present disclosure. The method of this embodiment can be applied to electronic devices such as servers. The file merging method for the real-time data lake includes:

[0037] S201. Obtain the characteristic information of the query tasks of the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table.

[0038] In this embodiment, to optimize costs, it is necessary to scientifically and reasonably determine the timing of initiating the merging task, and the timing of initiating the merging task is related to the specific situations of query tasks and merging tasks. In this embodiment, the characteristic information of the query tasks of the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table can be obtained to measure the specific situations of query tasks and merging tasks. The files to be merged may include the stock files and incremental files of the real-time data lake table in the disk.

[0039] In this embodiment, a query task may refer to a task of querying target data from the files to be merged in a real-time data lake table based on specific query conditions. For example, querying data of a certain target object. The query content of the query task in the embodiments of the present disclosure is not limited. The query task may be directly and actively initiated by a user, or may be automatically initiated during the process of providing business function services for the user, or may also be initiated based on the requirements of the background system of the business function service.

[0040] The merging task may refer to merging the files to be merged according to preset rules, so as to reduce the number and data volume of the files to be merged. The preset rules may be merging according to categories, sources or timeliness, and this is not limited in this embodiment.

[0041] Optionally, the characteristic information of the query task may refer to the task information of the historical query tasks of the real-time data lake table. For example, it may include but is not limited to the query frequency, resource configuration amount, and duration of any historical query task. Among them, the resource configuration amount may refer to the computer resources occupied during the query task process, and the computer resources may include CPU occupancy, memory occupancy, disk occupancy, etc. Optionally, the characteristic information of the query task may be obtained from the task management center, that is, the task management center will save the characteristic information of each query task after creating the query task for subsequent use; of course, it may also be obtained through any other feasible method, which is not limited here.

[0042] Optionally, the attribute information of the files to be merged may include but is not limited to the size and type of the files to be merged. Optionally, in this embodiment, the attribute information of the files to be merged may be counted by scanning the files to be merged, or other methods may also be used to obtain the attribute information of the files to be merged.

[0043] Optionally, the attribute information of the files to be merged may be obtained through a preset metadata center, where the metadata center stores the attribute information of the files to be merged in advance, and the attribute information of these files to be merged may be written into the metadata center as metadata when the files to be merged are stored on the disk. Of course, the metadata center may also determine the locations of each file to be merged, and then count the number and data volume of the files to be merged.

[0044] Optionally, the metadata center in this embodiment may support different forms of real-time data lakes, and may provide storage and management of metadata for different forms of real-time data lakes. When small files, that is, files to be merged, are generated in any implemented data lake, their file attribute information is written into the metadata center as metadata. Then, when the attribute information of the files to be merged needs to be obtained, it can be directly obtained from the metadata center.

[0045] In specific implementation, any real-time data lake, such as a real-time data lake based on Iceberg (an open table format), can support multiple forms, such as Hive Catalog (providing access to and management of Hive metadata), Hadoop Catalog (metadata management in the Hadoop ecosystem), and Rest Catalog (providing a unified API for metadata management). Considering that the places where metadata is stored are different in different forms, for example, Hive Catalog is defaultly stored in HMS (Hive Metastore, Hive metadata management), Hadoop Catalog is defaultly stored in file storage, and Rest Catalog is stored in a user-defined backend storage, resulting in metadata being scattered in different places, which may lead to inaccurate acquisition of metadata. Therefore, in this embodiment, a metadata center is provided, which can provide storage and management of metadata for different forms of real-time data lakes, be compatible with the Hive protocol, and provide HTTP & Thrift (network communication protocols) services externally. Each time data is written, the electronic device will be notified asynchronously, so as to maintain a view, thereby ensuring the accuracy of metadata.

[0046] S202. Determine whether to start a merging task according to the characteristic information of the query task and the attribute information of the file to be merged.

[0047] In this embodiment, based on the characteristic information of the query task for obtaining the real-time data lake table and the attribute information of the file to be merged in the real-time data lake table, the specific situations of the query task and the merging task can be measured, and the cost prediction and comparison in different situations can be carried out. With the goal of cost optimization, it is determined whether to start a merging task.

[0048] Optionally, the electronic device can predict the overhead of the query task in the case of not merging the file to be merged and the overhead of the query task and the merging task in the case of merging the file to be merged according to the characteristic information of the query task and the attribute information of the file to be merged, so as to determine whether to start a merging task.

[0049] Specifically, determining whether to start a merging task according to the characteristic information of the query task and the attribute information of the file to be merged includes:

[0050] Predict a first resource overhead index of the query task in the case of not merging the file to be merged and a second resource overhead index of the query task and the merging task in the case of merging the file to be merged according to the characteristic information of the query task and the attribute information of the file to be merged;

[0051] Determine whether to start a merging task according to the first resource overhead index and the second resource overhead index.

[0052] In this embodiment, the electronic device predicts the first resource overhead metric of the query task in the case where the files to be merged are not merged, and the second resource overhead metrics of the query task and the merging task in the case where the files to be merged are merged, based on the characteristic information of the query task and the attribute information of the files to be merged. Thus, the first resource overhead metric and the second resource overhead metrics can be compared to better determine whether to start the merging task, achieving the purpose of reducing costs and improving query performance.

[0053] Among them, the first resource overhead metric and the second resource overhead metrics can refer to parameters that quantify the overhead, facilitating the comparison between the first resource overhead metric and the second resource overhead metrics.

[0054] S203. If it is determined to start the merging task, then merge the files to be merged.

[0055] In this embodiment, if it is determined to start the merging task, the execution of the merging task can be started, that is, the files to be merged are merged using a preset rule, thereby reducing the number and data volume of the files to be merged. The preset rule can be merging by category, source, or timeliness, and the present disclosure does not limit this.

[0056] In specific implementation, the files to be merged can be extracted from the disk to the memory, merged in the memory to obtain the merged files, and then the merged files are returned to the disk.

[0057] The file merging method for the real-time data lake provided by the present disclosure predicts the overhead of the query task in the case where the files to be merged are not merged, and the overhead of the query task and the merging task in the case where the files to be merged are merged, by obtaining the characteristic information of the query task of the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table. Furthermore, the scheme with a smaller overhead is selected, and by scientifically and reasonably determining the file merging timing for the real-time data lake table, the user cost is reduced and the query performance is improved.

[0058] Reference Figure 3 , Figure 3 is a schematic flowchart of the file merging method for the real-time data lake provided by another embodiment of the present disclosure. The method of this embodiment can be applied to an electronic device or a server. The file merging method for the real-time data lake includes:

[0059] S301. Obtain the characteristic information of the query task of the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table.

[0060] S302. Predict the first predicted resource amount and the first predicted duration required for the query task in the case where the files to be merged are not merged, based on the characteristic information of the query task.

[0061] Among them, in the real-time data lake scenario, since data streams are continuously written into the implementation data lake table, small files, that is, files to be merged, will be continuously generated, resulting in an increase in the resource configuration amount and duration of the next query task. Assuming that the generation speed of the files to be merged does not change suddenly, the changes in the resource configuration amount and duration of each query task can be considered to conform to a certain change rule. Therefore, without file merging, the first predicted resource amount and the first predicted duration of this query task can be predicted based on the resource configuration amount and duration of historical query tasks.

[0062] For example, without file merging, the resource configuration amount of the Nth historical query task is a, and the query duration is b; the resource configuration amount of the (N + 1)th historical query task is a + 1, and the query duration is b + 2; the resource configuration amount of the (N + 2)th historical query task is a + 2, and the query duration is b + 4. According to the above change trend, it can be predicted that the resource configuration amount of the (N + 3)th historical query task is a + 3, and the query duration is b + 6.

[0063] Optionally, when predicting the first predicted resource amount and the first predicted duration of this query task without merging the files to be merged based on the resource configuration amount and duration of historical query tasks without file merging, it can be assumed that the change trends of the first predicted resource amount and the first predicted duration are linear. As in the above example, a linear regression algorithm can be used to determine the linear change rule of the resource configuration amount and duration of historical query tasks based on the characteristic information of the query task, especially the resource configuration amount and duration of historical query tasks, and then predict the first predicted resource amount and the first predicted duration.

[0064] Optionally, a preset prediction model can also be used to learn the characteristic information of the query task to learn the linear change rule of the resource configuration amount and duration of historical query tasks, and then predict the first predicted resource amount and the first predicted duration. The preset prediction model can be any possible model, such as a neural network model, etc.

[0065] Of course, in this embodiment, any other possible method can also be used to predict the first predicted resource amount and the first predicted duration required for the query task without merging the files to be merged based on the characteristic information of the query task, and no limitation is imposed here.

[0066] S303. Predict the second predicted resource amount and the second predicted duration required for the query task in the case of merging the files to be merged, as well as the third predicted resource amount and the third predicted duration required for the merging task, according to the characteristic information of the query task and the attribute information of the files to be merged.

[0067] In this embodiment, according to the characteristic information of the query task and the attribute information of the files to be merged, the number and data volume of the merged files after the current files to be merged are merged can be determined first, and then, according to the number and data volume of the merged files and the characteristic information of the query task, the second predicted resource volume and the second predicted duration required for the query task on the merged files in the case of merging the files to be merged this time can be predicted.

[0068] Among them, the characteristic information of the query task may further include the resource configuration volume and duration of the historical query task in the case of file merging, and may also include the number and data volume of the historical merged files, that is, the number and data volume of the files when the historical query task is executed in the case of file merging.

[0069] In the case of file merging, the changes in the resource configuration volume and duration of the historical query task under specific numbers and data volumes of files can be considered to conform to a certain change rule. Therefore, in the case of file merging, by determining the change rule among the number and data volume of the historical merged files, the resource configuration volume of the historical query task, and the duration, it can be further used to predict the second predicted resource volume and the second predicted duration required for the query task on the merged files in the case of merging the files to be merged this time.

[0070] For example, in the case of file merging, the resource configuration volume of the Nth historical query task is a, the query duration is b, the number of files when the Nth historical query task is executed is h, and the data volume is i; the number of merged files in the case of merging the files to be merged this time is g1, and the data volume is g2; thus, the second predicted resource volume and the second predicted duration required for the query task on the merged files in the case of merging the files to be merged this time can be predicted.

[0071] Optionally, it can be assumed that in the case of file merging, the change rule among the number and data volume of the historical merged files, the resource configuration volume of the historical query task, and the duration is a linear rule. Therefore, a linear regression algorithm can be used to predict the second predicted resource volume and the second predicted duration required for the query task on the merged files in the case of merging the files to be merged this time based on the number and data volume of the historical merged files, the resource configuration volume of the historical query task, the duration, and the number and data volume of the merged files this time.

[0072] Optionally, a preset prediction model can also be used to learn the variation rules among the number and data volume of historical merged files, the resource allocation volume of historical query tasks, and the duration, and then predict the second predicted resource volume and the second predicted duration required for the query task of the merged files in the case of merging the files to be merged this time. The preset prediction model can be any possible model, such as a neural network model, etc.

[0073] Of course, in this embodiment, any other possible method can also be used to predict the second predicted resource volume and the second predicted duration required for the query task of the merged files in the case of merging the files to be merged this time, and no limitation is imposed here.

[0074] In addition, when the number and data volume of the merged files after the current files to be merged are merged, the number and data volume of the merged files after the current files to be merged are merged can be predicted based on the correspondence between the first number and the first data volume before the historical merging task and the second number and the second data volume after the historical merging task.

[0075] For example, for historical merging task M1, the first number of files before merging is d1, the first data volume is d2, the second number of files after merging is e1, and the second data volume is e2; for historical merging task M2, the first number of files before merging is d3, the first data volume is d4, the second number of files after merging is e3, and the second data volume is e4; the number of the current files to be merged is f1, and the data volume of the files to be merged is f2; the number of the merged files after the files to be merged are merged and to be predicted is g1, and the data volume is g2. In an optional solution, it can be assumed that the variation rules of the number and data volume of the files before and after the merging task show a linear relationship. Therefore, g1 and g2 can be obtained according to d1 / e1 = f1 / g1 and d2 / e2 = f2 / g2; alternatively, according to the interpolation algorithm, (d1 - d3) / (e1 - e3) = (d3 - f1) / (e3 - g1) and (d2 - d4) / (e2 - e4) = (d4 - f2) / (e4 - g2) can be constructed to obtain g1 and g2.

[0076] In addition, the third predicted resource volume and the third predicted duration required for the merging task can also be predicted according to the attribute information of the files to be merged. On the basis of knowing the attribute information of the files to be merged (including the number and data volume of the files to be merged), optionally, the third predicted resource volume and the third predicted duration required for this merging task can be predicted according to preset rules or preset experience information.

[0077] Optionally, the third predicted resource amount and the third predicted duration required for the current merging task can also be predicted based on the resource amount and duration of historical merging tasks (where the number of files and the amount of data before and after merging are known). Specifically, it can be assumed that the variation law between the resource amount and duration of historical merging tasks, and the number of files and the amount of data before and after merging is a linear law. Therefore, a linear regression algorithm can be used with the resource amount and duration of historical merging tasks, the number of files and the amount of data before and after merging, as well as the number of files and the amount of data of the files to be merged currently, to predict the third predicted resource amount and the third predicted duration required for the current merging task.

[0078] Optionally, a preset prediction model can also be used to learn the variation law between the resource amount and duration of historical merging tasks, and the number of files and the amount of data before and after merging. Then, based on the number of files and the amount of data of the files to be merged currently, the third predicted resource amount and the third predicted duration required for the current merging task can be predicted. The preset prediction model can be any possible model, such as a neural network model, etc.

[0079] S304. Determine a first resource overhead metric based on the first predicted resource amount and the first predicted duration, and determine a second resource overhead metric based on the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration.

[0080] In this embodiment, after determining the first predicted resource amount and the first predicted duration required for the query task in the case of not merging the files to be merged, the first resource overhead metric for the query task in the case of not merging the files to be merged can be determined based on the first predicted resource amount and the first predicted duration; after determining the second predicted resource amount and the second predicted duration required for the query task in the case of merging the files to be merged, as well as the third predicted resource amount and the third predicted duration required for the merging task, the second resource overhead metric for the query task and the merging task in the case of merging the files to be merged can be determined based on the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration. It should be noted that in this embodiment, the specific forms of the first resource overhead metric and the second resource overhead metric are not limited, as long as the first resource overhead metric and the second resource overhead metric are comparable. For example, the first resource overhead metric and the second resource overhead metric can be represented in combination with the resource price.

[0081] In a possible implementation manner, the methods for determining the first resource overhead metric and the second resource overhead metric can include:

[0082] Obtain the first product between the first predicted resource amount and the first predicted duration, and determine the first product as the first resource overhead metric;

[0083] Obtain the second product between the second predicted resource amount and the second predicted duration, and the third product between the third predicted resource amount and the third predicted duration, and determine the second resource overhead metric as the sum of the second product and the third product.

[0084] In the embodiments of the present disclosure, when the amount of resources used is larger, the resource overhead is larger, and when the query time is longer, the resource overhead is also larger. Therefore, using the product of the resource amount and the duration as the resource overhead metric can more accurately quantify the first resource overhead metric and the second resource overhead metric.

[0085] Its specific formula can be as follows:

[0086] First resource overhead metric = C1 * T1;

[0087] Second resource overhead metric = C2 * T2 + C3 * T3;

[0088] Wherein, C1 is the first predicted resource amount required for the query task without merging the file to be merged, T1 is the first predicted duration required for the query task without merging the file to be merged; C2 is the second predicted resource amount required for the query task when merging the file to be merged, T2 is the second predicted duration required for the query task when merging the file to be merged; C3 is the third predicted resource amount required for the merging task, and T3 is the third predicted duration required for the merging task.

[0089] In another possible implementation, the method for determining the first resource overhead metric and the second resource overhead metric may include:

[0090] Obtain the fourth product between the first predicted resource amount, the first predicted duration, and the query frequency of the query task, and determine the fourth product as the first resource overhead metric;

[0091] Obtain the fifth product between the second predicted resource amount, the second predicted duration, and the query frequency of the query task, and the sixth product between the third predicted resource amount and the third predicted duration, and determine the second resource overhead metric as the sum of the fifth product and the sixth product.

[0092] In the embodiments of the present disclosure, when the query frequency is greater than 1, since the overhead of each query task increases successively, the query frequency is multiplied by the first product and the second product respectively, and the first product and the second product are amplified according to the query frequency, so that the obtained first resource overhead metric and second resource overhead metric are more accurate (of course, when the query frequency is 1, it is equivalent to the above embodiment).

[0093] Its specific formula can be as follows:

[0094] First resource overhead metric = C1 * T1 * F;

[0095] Second resource overhead metric = C2 * T2 * F + C3 * T3;

[0096] where F is the query frequency of the query task.

[0097] In the case of a periodic query task, by comparing the magnitudes of the first resource overhead metric and the second resource overhead metric, the user cost can be reduced. On the other hand, after merging, the number of files and the amount of data during query of the query task can be reduced, thereby improving the query performance.

[0098] S305. Determine whether to start a merging task according to the first resource overhead metric and the second resource overhead metric.

[0099] Specifically, if the first resource overhead metric is greater than the second resource overhead metric, it is determined to start the merging task; or if the ratio of the first resource overhead metric to the second resource overhead metric is greater than a preset threshold, it is determined to start the merging task.

[0100] The preset threshold can be adjusted according to user requirements, thereby adjusting the merging frequency of the merging task.

[0101] S306. If it is determined to start the merging task, merge the files to be merged.

[0102] Reference Figure 4 , Figure 4 is an example diagram of file merging provided by an embodiment of the present disclosure. According to the requirements of the real-time data lake table, the user configures the file merging method of the real-time data lake and the preset merging frequency into the electronic device, where the preset merging frequency is the frequency for determining whether to start the merging task, for example, once per second. The electronic device triggers the merging task judgment according to the preset merging frequency. If it is determined to start the merging task, the small files in the real-time data lake table are merged. If it is determined not to start the merging task, the small files in the real-time data lake table are not merged.

[0103] Another embodiment of the present disclosure provides a method for merging files in a real-time data lake. By obtaining the characteristic information of the query task for the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table, the first predicted resource amount and the first predicted duration required for the query task in the case of not merging the files to be merged, the second predicted resource amount and the second predicted duration required for the query task in the case of merging the files to be merged, and the third predicted resource amount and the third predicted duration required for the merging task are predicted. Then, a first resource overhead index is determined based on the first predicted resource amount and the first predicted duration, a second resource overhead index is determined based on the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration. Finally, it is determined whether to start the merging task based on the magnitudes of the first resource overhead index and the second resource overhead index, comprehensively considering the resource overhead and duration before and after the merging task, improving the flexibility of the merging task, reducing user costs, and improving query performance.

[0104] Corresponding to the method for merging files in the real-time data lake in the above embodiment, Figure 5 is a structural block diagram of a device for merging files in a real-time data lake provided by an embodiment of the present disclosure. For ease of explanation, only parts related to the embodiment of the present disclosure are shown. Referring to Figure 5 the device 50 includes: an obtaining unit 501, a judging unit 502, and an executing unit 503, where:

[0105] The obtaining unit 501 is configured to obtain the characteristic information of the query task for the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table;

[0106] The judging unit 502 is configured to judge whether to start the merging task according to the characteristic information of the query task and the attribute information of the files to be merged;

[0107] The executing unit 503 is configured to merge the files to be merged if it is determined to start the merging task.

[0108] In one or more embodiments of the present disclosure, the judging unit 502 is further configured to:

[0109] Predict a first resource overhead index of the query task in the case of not merging the files to be merged and a second resource overhead index of the query task and the merging task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged;

[0110] Judge whether to start the merging task according to the first resource overhead index and the second resource overhead index.

[0111] In one or more embodiments of the present disclosure, the judging unit 502 is further configured to:

[0112] Predict the first predicted resource amount and the first predicted duration required for the query task in the case where the files to be merged are not merged, according to the characteristic information of the query task;

[0113] Predict the second predicted resource amount and the second predicted duration required for the query task in the case where the files to be merged are merged, according to the characteristic information of the query task and the attribute information of the files to be merged, as well as the third predicted resource amount and the third predicted duration required for the merging task;

[0114] Determine the first resource overhead metric according to the first predicted resource amount and the first predicted duration, and determine the second resource overhead metric according to the second predicted resource amount, the second predicted duration, the third predicted resource amount and the third predicted duration.

[0115] In one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0116] Obtain the first product between the first predicted resource amount and the first predicted duration, and determine the first product as the first resource overhead metric;

[0117] Obtain the second product between the second predicted resource amount and the second predicted duration, and the third product between the third predicted resource amount and the third predicted duration, and determine the sum of the second product and the third product as the second resource overhead metric.

[0118] In one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0119] Obtain the fourth product among the first predicted resource amount, the first predicted duration and the query frequency of the query task, and determine the fourth product as the first resource overhead metric;

[0120] Obtain the fifth product among the second predicted resource amount, the second predicted duration and the query frequency of the query task, and the sixth product between the third predicted resource amount and the third predicted duration, and determine the sum of the fifth product and the sixth product as the second resource overhead metric.

[0121] In one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0122] Adopt a linear regression algorithm or a preset prediction model to predict the first predicted resource amount and the first predicted duration required for the query task in the case where the files to be merged are not merged, according to the characteristic information of the query task;

[0123] Adopt a linear regression algorithm or a preset prediction model to predict the second predicted resource amount and the second predicted duration required for the query task in the case where the files to be merged are merged, according to the characteristic information of the query task and the attribute information of the files to be merged, as well as the third predicted resource amount and the third predicted duration required for the merging task.

[0124] In one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0125] If the first resource consumption metric is greater than the second resource consumption metric, determine to start the merging task; or

[0126] If the ratio of the first resource consumption metric to the second resource consumption metric is greater than a preset threshold, determine to start the merging task.

[0127] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0128] Refer to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The electronic device 600 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0129] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0130] Typically, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0131] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0132] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0133] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately and not be assembled into the electronic device.

[0134] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0135] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0137] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0138] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0139] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0140] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a method for merging files in a real-time data lake, including:

[0141] Obtaining characteristic information of a query task of a real-time data lake table and attribute information of files to be merged in the real-time data lake table;

[0142] Judging whether to start a merging task according to the characteristic information of the query task and the attribute information of the files to be merged;

[0143] If it is determined to start the merging task, then merge the files to be merged.

[0144] According to one or more embodiments of the present disclosure, judging whether to start a merging task according to the characteristic information of the query task and the attribute information of the files to be merged includes:

[0145] Predicting a first resource overhead metric of the query task in the case of not merging the files to be merged and a second resource overhead metric of the query task and the merging task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged;

[0146] Judging whether to start the merging task according to the first resource overhead metric and the second resource overhead metric.

[0147] According to one or more embodiments of the present disclosure, predicting a first resource overhead metric of the query task in the case of not merging the files to be merged and a second resource overhead metric of the query task and the merging task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged includes:

[0148] Predicting a first predicted resource amount and a first predicted duration required for the query task in the case of not merging the files to be merged according to the characteristic information of the query task;

[0149] According to the characteristic information of the query task and the attribute information of the files to be merged, predict the second predicted resource amount and the second predicted duration required for the query task when the files to be merged are merged, as well as the third predicted resource amount and the third predicted duration required for the merging task;

[0150] Determine a first resource overhead metric according to the first predicted resource amount and the first predicted duration, and determine a second resource overhead metric according to the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration.

[0151] According to one or more embodiments of the present disclosure, determining a first resource overhead metric according to the first predicted resource amount and the first predicted duration includes:

[0152] Obtain a first product between the first predicted resource amount and the first predicted duration, and determine the first product as the first resource overhead metric;

[0153] Determining a second resource overhead metric according to the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration includes:

[0154] Obtain a second product between the second predicted resource amount and the second predicted duration, and a third product between the third predicted resource amount and the third predicted duration, and determine the sum of the second product and the third product as the second resource overhead metric.

[0155] According to one or more embodiments of the present disclosure, determining a first resource overhead metric according to the first predicted resource amount and the first predicted duration includes:

[0156] Obtain a fourth product between the first predicted resource amount, the first predicted duration, and the query frequency of the query task, and determine the fourth product as the first resource overhead metric;

[0157] Determining a second resource overhead metric according to the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration includes:

[0158] Obtain a fifth product between the second predicted resource amount, the second predicted duration, and the query frequency of the query task, and a sixth product between the third predicted resource amount and the third predicted duration, and determine the sum of the fifth product and the sixth product as the second resource overhead metric.

[0159] According to one or more embodiments of the present disclosure, the characteristic information of the query task includes one or more of the following: the resource allocation amounts and durations of multiple historical query tasks, and the query frequency of the historical query tasks;

[0160] The attribute information of the files to be merged includes one or more of the following: the number and data volume of the files to be merged.

[0161] According to one or more embodiments of the present disclosure, predicting a first predicted resource amount and a first predicted duration required for a query task without merging files to be merged based on the characteristic information of the query task includes:

[0162] Using a linear regression algorithm or a preset prediction model, predicting a first predicted resource amount and a first predicted duration required for a query task without merging files to be merged based on the characteristic information of the query task;

[0163] According to the characteristic information of the query task and the attribute information of the files to be merged, predicting a second predicted resource amount and a second predicted duration required for the query task when merging the files to be merged, and a third predicted resource amount and a third predicted duration required for the merging task, including:

[0164] Using a linear regression algorithm or a preset prediction model, predicting a second predicted resource amount and a second predicted duration required for the query task when merging the files to be merged, and a third predicted resource amount and a third predicted duration required for the merging task according to the characteristic information of the query task and the attribute information of the files to be merged.

[0165] According to one or more embodiments of the present disclosure, determining whether to start a merging task based on a first resource overhead metric and a second resource overhead metric includes:

[0166] If the first resource overhead metric is greater than the second resource overhead metric, determining to start the merging task; or

[0167] If the ratio of the first resource overhead metric to the second resource overhead metric is greater than a preset threshold, determining to start the merging task.

[0168] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a file merging device for a real-time data lake, including:

[0169] An obtaining unit, configured to obtain the characteristic information of a query task for a real-time data lake table and the attribute information of files to be merged in the real-time data lake table;

[0170] A determining unit, configured to determine whether to start a merging task according to the characteristic information of the query task and the attribute information of the files to be merged;

[0171] An execution unit, configured to merge the files to be merged if it is determined to start the merging task.

[0172] According to one or more embodiments of the present disclosure, the determining unit 502 is further configured to:

[0173] Predict the first resource consumption metric of the query task in the case where the files to be merged are not merged, and the second resource consumption metrics of the query task and the merging task in the case where the files to be merged are merged, based on the characteristic information of the query task and the attribute information of the files to be merged;

[0174] Determine whether to start the merging task according to the first resource consumption metric and the second resource consumption metric.

[0175] According to one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0176] Predict the first predicted resource amount and the first predicted duration required for the query task in the case where the files to be merged are not merged, based on the characteristic information of the query task;

[0177] Predict the second predicted resource amount and the second predicted duration required for the query task, and the third predicted resource amount and the third predicted duration required for the merging task in the case where the files to be merged are merged, based on the characteristic information of the query task and the attribute information of the files to be merged;

[0178] Determine the first resource consumption metric according to the first predicted resource amount and the first predicted duration, and determine the second resource consumption metric according to the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration.

[0179] According to one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0180] Obtain the first product between the first predicted resource amount and the first predicted duration, and determine the first product as the first resource consumption metric;

[0181] Obtain the second product between the second predicted resource amount and the second predicted duration, and the third product between the third predicted resource amount and the third predicted duration, and determine the sum of the second product and the third product as the second resource consumption metric.

[0182] According to one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0183] Obtain the fourth product among the first predicted resource amount, the first predicted duration, and the query frequency of the query task, and determine the fourth product as the first resource consumption metric;

[0184] Obtain the fifth product among the second predicted resource amount, the second predicted duration, and the query frequency of the query task, and the sixth product between the third predicted resource amount and the third predicted duration, and determine the sum of the fifth product and the sixth product as the second resource consumption metric.

[0185] According to one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0186] Using a linear regression algorithm or a preset prediction model, predict the first predicted resource amount and the first predicted duration required for the query task without merging the files to be merged according to the characteristic information of the query task;

[0187] Using a linear regression algorithm or a preset prediction model, predict the second predicted resource amount and the second predicted duration required for the query task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged, as well as the third predicted resource amount and the third predicted duration required for the merging task.

[0188] According to one or more embodiments of the present disclosure, the determination unit 502 is further configured to:

[0189] If the first resource overhead metric is greater than the second resource overhead metric, determine to start the merging task; or

[0190] If the ratio of the first resource overhead metric to the second resource overhead metric is greater than a preset threshold, determine to start the merging task.

[0191] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;

[0192] The memory stores computer-executable instructions;

[0193] At least one processor executes the computer-executable instructions stored in the memory, so that at least one processor executes the file merging method of the real-time data lake as described in the first aspect above and various possible designs of the first aspect.

[0194] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the file merging method of the real-time data lake as described in the first aspect above and various possible designs of the first aspect is implemented.

[0195] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product, including computer-executable instructions, and when the processor executes the computer-executable instructions, the file merging method of the real-time data lake as described in the first aspect above and various possible designs of the first aspect is implemented.

[0196] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0197] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0198] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A file merging method for a real-time data lake, characterized in that, Including: Obtaining the characteristic information of the query task for the real-time data lake table and the attribute information of the files to be merged in the real-time data lake table; Judging whether to start the merging task according to the characteristic information of the query task and the attribute information of the files to be merged; If it is determined to start the merging task, then merge the files to be merged; The judging whether to start the merging task according to the characteristic information of the query task and the attribute information of the files to be merged includes: Predicting the first resource overhead index of the query task in the case of not merging the files to be merged and the second resource overhead index of the query task and the merging task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged; Judging whether to start the merging task with the goal of minimizing the overhead according to the first resource overhead index and the second resource overhead index; The predicting the first resource overhead index of the query task in the case of not merging the files to be merged and the second resource overhead index of the query task and the merging task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged includes: Predicting the first predicted resource amount and the first predicted duration required for the query task in the case of not merging the files to be merged according to the characteristic information of the query task; Predicting the second predicted resource amount and the second predicted duration required for the query task and the third predicted resource amount and the third predicted duration required for the merging task in the case of merging the files to be merged according to the characteristic information of the query task and the attribute information of the files to be merged; Determining the first resource overhead index according to the first predicted resource amount and the first predicted duration, and determining the second resource overhead index according to the second predicted resource amount, the second predicted duration, the third predicted resource amount and the third predicted duration.

2. The method according to claim 1, wherein The determining the first resource overhead index according to the first predicted resource amount and the first predicted duration includes: Obtaining the first product between the first predicted resource amount and the first predicted duration, and determining the first product as the first resource overhead index; The determining the second resource overhead index according to the second predicted resource amount, the second predicted duration, the third predicted resource amount and the third predicted duration includes: Obtaining the second product between the second predicted resource amount and the second predicted duration and the third product between the third predicted resource amount and the third predicted duration, and determining the sum of the second product and the third product as the second resource overhead index.

3. The method according to claim 1, characterized in that, The determining the first resource overhead index according to the first predicted resource amount and the first predicted duration includes: Obtaining the fourth product among the first predicted resource amount, the first predicted duration and the query frequency of the query task, and determining the fourth product as the first resource overhead index; Determining the second resource overhead metric according to the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration includes: Obtaining a fifth product among the second predicted resource amount, the second predicted duration, and the query frequency of the query task, and a sixth product between the third predicted resource amount and the third predicted duration, and determining the second resource overhead metric as the sum of the fifth product and the sixth product.

4. The method according to any one of claims 1 to 3, characterized in that, The characteristic information of the query task includes one or more of the following: the resource allocation amounts and durations of multiple historical query tasks, and the query frequency of historical query tasks; The attribute information of the file to be merged includes one or more of the following: the number and data volume of the files to be merged.

5. The method according to claim 4, wherein Predicting the first predicted resource amount and the first predicted duration required for the query task without merging the file to be merged according to the characteristic information of the query task includes: Using a linear regression algorithm or a preset prediction model to predict the first predicted resource amount and the first predicted duration required for the query task without merging the file to be merged according to the characteristic information of the query task; Predicting the second predicted resource amount and the second predicted duration required for the query task and the third predicted resource amount and the third predicted duration required for the merging task when merging the file to be merged according to the characteristic information of the query task and the attribute information of the file to be merged includes: Using a linear regression algorithm or a preset prediction model to predict the second predicted resource amount and the second predicted duration required for the query task and the third predicted resource amount and the third predicted duration required for the merging task when merging the file to be merged according to the characteristic information of the query task and the attribute information of the file to be merged.

6. The method according to claim 1, characterized in that, Judging whether to start a merging task according to the first resource overhead metric and the second resource overhead metric includes: If the first resource overhead metric is greater than the second resource overhead metric, determining to start a merging task; or If the ratio of the first resource overhead metric to the second resource overhead metric is greater than a preset threshold, determining to start a merging task.

7. A file merging device for a real-time data lake, characterized in that, Includes: An obtaining unit, configured to obtain the characteristic information of the query task for the real-time data lake table and the attribute information of the file to be merged in the real-time data lake table; A judging unit, configured to judge whether to start a merging task according to the characteristic information of the query task and the attribute information of the file to be merged; An executing unit, configured to merge the file to be merged if it is determined to start a merging task; When judging whether to start a merging task according to the characteristic information of the query task and the attribute information of the file to be merged, the judging unit is configured to: Predict the first resource overhead metric of the query task without merging the file to be merged and the second resource overhead metric of the query task and the merging task when merging the file to be merged according to the characteristic information of the query task and the attribute information of the file to be merged; Based on the first resource overhead metric and the second resource overhead metric, determine whether to start the merging task with the goal of minimizing the overhead. When the determining unit predicts the first resource overhead metric of the query task in the case where the files to be merged are not merged, and the second resource overhead metric of the query task and the merging task in the case where the files to be merged are merged, based on the characteristic information of the query task and the attribute information of the files to be merged, it is used for: Predict the first predicted resource amount and the first predicted duration required for the query task in the case where the files to be merged are not merged, based on the characteristic information of the query task. Predict the second predicted resource amount and the second predicted duration required for the query task, and the third predicted resource amount and the third predicted duration required for the merging task, in the case where the files to be merged are merged, based on the characteristic information of the query task and the attribute information of the files to be merged. Determine the first resource overhead metric based on the first predicted resource amount and the first predicted duration, and determine the second resource overhead metric based on the second predicted resource amount, the second predicted duration, the third predicted resource amount, and the third predicted duration.

8. An electronic device, characterized in that, Comprising: At least one processor and a memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, such that the at least one processor executes the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the method according to any one of claims 1-6 is implemented.

10. A computer program product, characterized in that, Comprising computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Data query method and system

    CN110716900A

  • File merging method, processor and storage medium

    CN114116224A