Bucket dividing method and device for data lake

By dynamically configuring the number of partitioned buckets in the data lake system, the data tilt problem caused by the fixed number of partitioned buckets in the prior art is solved, and the data query and writing efficiency of the data lake are improved.

CN120020750APending Publication Date: 2025-05-20HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410268575.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-03-08
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In the existing data lake bucket writing technology, the number of buckets in partitions is fixed, resulting in uneven distribution of data in different partitions, causing data skew, and reducing query and writing efficiency.

Method used

By obtaining business demand information in the data lake system, predicting the amount of data to be written in each partition, dynamically configure the number of buckets per partition, generating a pre-packet bucket policy, and storing the policy in the metadata area of ​​the data table, so that the number of buckets of partitions is automatically set when writing data.

Benefits of technology

Improve the flexibility of bucketing and data query and writing efficiency, ensure the uniform distribution of data in the partition, and avoid data skew problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020750A_ABST
    Figure CN120020750A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a bucket dividing method and device for a data lake. The bucket dividing method and device are used for improving bucket dividing flexibility and data query and writing efficiency of the data lake. The method comprises the steps that a data lake system obtains service demand information, the service demand information is used for indicating the predicted data volume of to-be-written data in one or more partitions in a data lake, and the to-be-written data is used for being written into data buckets of the one or more partitions. And according to the business demand information, determining a bucket pre-distribution strategy of the one or more partitions, wherein the bucket pre-distribution strategy is used for indicating the number of buckets in different partitions in the data lake. And creating a data table in the data lake based on the pre-bucket distribution strategy, wherein the data table is used for managing data in the data lake.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202311552902.1, titled "A Method and Device for Writing Dynamic Bucketing", filed with the China National Intellectual Property Administration on November 17, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the field of cloud computing, and in particular, to a method and device for bucketing in a data lake. Background Art

[0003] With the continuous development of data lake services, in some scenarios of reading and writing large amounts of data, users have higher and higher requirements for the query and update performance of data lakes. Since the bucketing write technology can bring higher performance data query and update capabilities, more and more users tend to choose to use the bucketing write technology to write data in the data lake when constructing a data lake.

[0004] In the current bucketing write technology, a data lake contains multiple partitions, and each partition contains multiple buckets. During the data writing process, the computing device needs to load the bucketing strategy of each partition, where the bucketing strategy can indicate the number of buckets in that partition. In the current process of configuring the bucketing strategy, the computing device needs to manually configure the bucketing strategy of each partition, and currently, a fixed number of buckets is used for each partition in the bucketing strategy.

[0005] However, when users write data based on actual business, there will be a situation where the written data is unevenly distributed in different partitions. Therefore, using a fixed number of buckets for each partition will result in some buckets having too much data and some buckets having too little data, making the bucketing flexibility poor and prone to data skew, further leading to low data query and write efficiency. Summary of the Invention

[0006] Embodiments of this application provide a method for bucketing in a data lake. The data lake system can configure the number of buckets in different partitions according to business requirements, thereby improving the flexibility of bucketing and the data query and write efficiency. Embodiments of this application also provide a bucketing device for the data lake corresponding to the method for bucketing in the data lake, a computing device, a computing device cluster, a computer-readable storage medium, and a computer program product.

[0007] In a first aspect, an embodiment of the present application provides a method for bucketing a data lake. This method can be executed by a data lake system, or by components of the data lake system, such as a processor, a chip, or a chip system of the data lake system, or can also be implemented by a logic module or software that can implement all or part of the functions of the data lake system. The method provided in the first aspect includes: The data lake system obtains business requirement information, which is used to indicate the predicted data volume of the data to be written in one or more partitions in the data lake, and the data to be written is used to write into the data buckets of one or more partitions. The data lake system determines a pre-bucketing strategy for one or more partitions according to the business requirement information, and the pre-bucketing strategy is used to indicate the number of buckets in different partitions of the data lake, and the number of buckets in different partitions is different. The data lake system creates a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage the data in the data lake. After the data lake system receives the data to be written, it can perform bucketing and writing based on the pre-bucketing strategy in the data table.

[0008] In the embodiment of the present application, the data lake system can configure different pre-bucketing strategies for different partitions according to the business requirement information, and different partitions can flexibly set the number of buckets. Compared with the current data lake system that sets a fixed number of buckets, in the embodiment of the present application, the data lake system flexibly sets the number of buckets in different partitions based on the predicted data volume corresponding to the business requirement, thereby improving the configuration accuracy of the number of buckets in different partitions and the flexibility of bucketing, and further improving the data query and writing efficiency of the data lake.

[0009] In a possible implementation manner, when the data lake system determines the pre-bucketing strategy for each partition according to the business requirement information, the data lake system generates one or more partition strategy key-value pairs based on the business requirement information. The strategy key-value pair is used to indicate the number of buckets in different partitions, and the strategy key-value pair includes two parts: a key and a value. The key is used to specify the partition, and the value is used to specify the number of buckets corresponding to the partition. The types of the strategy key-value pair include one or more of: direct matching, regular expression matching, and range matching. Among them, direct matching means directly specifying the number of buckets in the partition, regular expression matching specifies the number of buckets in the partitions that meet the conditions through a string, and range matching is to specify the number of buckets in the partitions within a range. The data lake system generates a pre-bucketing strategy for one or more partitions based on the strategy key-value pair.

[0010] In the embodiment of the present application, the data lake system can automatically determine the strategy key-value pair according to the business requirement information and generate various types of pre-bucketing strategies based on the strategy key-value pair, thereby improving the configuration efficiency of the number of buckets in different partitions and further improving the data query and writing efficiency of the data lake.

[0011] In a possible implementation, during the process of the data lake system obtaining business requirement information, the data lake system obtains the historical write data volumes of one or more historical partitions, predicts the business requirement information based on the historical write data volumes of the one or more historical partitions, and the business requirement information includes the predicted data volume of the data to be written in a specified partition. During the process of the data lake system determining the pre-bucketing strategy of one or more partitions according to the business requirement information, the number of buckets of the specified partition is determined according to the predicted data volume of the data to be written in the specified partition and the bucket size of the specified partition.

[0012] In the embodiments of the present application, the data lake system can predict the business requirement information according to the historical write data volumes of different partitions, thereby improving the accuracy of the business requirement information and further improving the accuracy of the number of buckets of different partitions.

[0013] In a possible implementation, the data lake system calculates the predicted data volume of the data to be written in one or more partitions according to the historical write data volumes of one or more historical partitions and the weight coefficients corresponding to different historical partitions.

[0014] In the embodiments of the present application, during the process of the data lake system determining the business requirement information according to the historical write data volume, it can calculate the predicted data volume of the data to be written in one or more partitions according to the weight coefficients corresponding to different historical partitions, thereby improving the accuracy of the business requirement information and further improving the accuracy of the number of buckets of different partitions.

[0015] In a possible implementation, the data lake system calculates the predicted data volume of the data to be written in one or more partitions according to the historical write data volumes of one or more historical partitions and the number of historical partitions, that is, determines the predicted data volume of the data to be written according to the average value of the historical write data volumes of the historical partitions.

[0016] In the embodiments of the present application, during the process of the data lake system determining the business requirement information according to the historical write data volume, it can also determine the predicted data volume of the data to be written according to the average value of the historical write data volumes of the historical partitions, thereby improving the prediction efficiency of the business requirement information and further improving the configuration efficiency of the number of buckets of different partitions.

[0017] In a possible implementation, when the data lake system receives the data to be written and determines that there is no pre-bucketing strategy for one or more partitions where the data to be written is to be written, it automatically determines the number of buckets of the one or more partitions based on the partition statistical information, and the partition statistical information includes the write data volume of the data to be written in the one or more partitions.

[0018] In the embodiment of the present application, when the data lake system cannot obtain business requirement information to generate a pre-bucketing strategy, the data lake system can determine partition statistics information based on the data to be written, and determine the number of buckets for each partition based on the partition statistics information, thereby improving the bucketing flexibility and efficiency of the data lake, and further improving the data query and write efficiency of the data lake.

[0019] In a possible implementation, when the data lake system automatically determines the number of buckets for a partition based on partition statistics information, the data lake system performs an aggregation operation on the data to be written to determine the number of data records in each partition, samples the data to be written to determine the average size of each data record, determines partition statistics information based on the number of data records and the average size of each data record, and calculates the number of buckets for each partition in one or more partitions based on the partition statistics information and the bucket size of each partition.

[0020] In the embodiment of the present application, the data lake system can sample the data to be written to determine the average data size, perform partition aggregation on the data to be written to determine the number of data records in each partition, and calculate partition statistics information based on the average data size and the number of data records in each partition, thereby improving the accuracy of the partition statistics information and further improving the accuracy of the configuration of the number of buckets.

[0021] In a possible implementation, when the data lake system creates a data table in the data lake based on a pre-bucketing strategy, the pre-bucketing strategy is written in the metadata area of the data table. When the data lake system writes data into the data lake, the pre-bucketing strategy is loaded from the metadata area of the data table, and the number of buckets for the partition is automatically set according to the pre-bucketing strategy.

[0022] In the embodiment of the present application, the data lake system can directly write the pre-bucketing strategy in the metadata area of the data table, thereby improving the feasibility of configuring the pre-bucketing strategy.

[0023] In a possible implementation, when the data lake system creates a data table in the data lake based on a pre-bucketing strategy, a configuration file is generated based on the pre-bucketing strategy. The configuration file includes the pre-bucketing strategies for one or more partitions, and the loading path of the configuration file is written in the metadata area of the data table. When the data lake system writes data into the data lake, the loading path of the configuration file is obtained from the metadata of the data table, the configuration file of the pre-bucketing strategy is loaded, and the number of buckets for the partition is automatically set according to the pre-bucketing strategy in the configuration file.

[0024] In the embodiment of the present application, the data lake system can also write the loading path of the configuration file corresponding to the pre-bucketing strategy in the metadata area of the data table, thereby improving the feasibility of configuring the pre-bucketing strategy.

[0025] In a possible implementation, when multiple pre-bucketing policies are configured for a partition of the data lake, the number of buckets for a partition is determined based on the intersection of the multiple pre-bucketing policies.

[0026] In the embodiments of the present application, when multiple pre-bucketing policies are configured for a partition of the data lake, the data lake system can determine the number of buckets for the partition based on the intersection of the multiple pre-bucketing policies, thereby improving the configuration accuracy of the number of buckets.

[0027] In a second aspect, the embodiments of the present application provide a method for bucketing a data lake. This method can be executed by the data lake system, or by components of the data lake system, such as the processor, chip, or chip system of the data lake system, or can also be implemented by a logic module or software that can implement all or part of the functions of the data lake system. The method provided in the first aspect includes: the data lake system receives one or more policy key-value pairs input by the user, and the policy key-value pairs are used to indicate the number of buckets for different partitions. The types of policy key-value pairs include one or more of: direct matching, regular matching, and range matching. The data lake system generates pre-bucketing policies for one or more partitions in the data lake based on the one or more policy key-value pairs, and the pre-bucketing policies are used to indicate the number of buckets for different partitions in the data lake. The data lake system creates data tables in the data lake based on the pre-bucketing policies, and the data tables are used to manage the data in the data lake.

[0028] In the embodiments of the present application, the user can also directly configure the pre-bucketing policy through the policy key-value pair. After the data lake system receives the policy key-value pair input by the user, it generates the pre-bucketing policy based on the policy key-value pair, thereby improving the configuration efficiency of the number of buckets for different partitions and further improving the data query and write efficiency of the data lake.

[0029] In a possible implementation, during the process of the data lake system creating a data table in the data lake based on the pre-bucketing policy, the pre-bucketing policy is written into the metadata area of the data table.

[0030] In a possible implementation, during the process of the data lake system creating a data table in the data lake based on the pre-bucketing policy, a configuration file is generated based on the pre-bucketing policy. The configuration file includes the pre-bucketing policies for one or more partitions, and the loading path of the configuration file is written into the metadata area of the data table.

[0031] In a possible implementation, when multiple pre-bucketing policies are configured for a partition of the data lake, the number of buckets for a partition is determined based on the intersection of the multiple pre-bucketing policies.

[0032] In a third aspect, an embodiment of the present application provides a bucketing device for a data lake, the device including an acquisition unit, a processing unit, and a transceiver unit. Among them, the acquisition unit is used to acquire service requirement information, and the service requirement information is used to indicate the predicted data volume of the data to be written in one or more partitions in the data lake, and the data to be written is used to write into the data buckets of one or more partitions. The processing unit is used to determine a pre-bucketing strategy for one or more partitions according to the service requirement information, and the pre-bucketing strategy is used to indicate the number of buckets in different partitions in the data lake. The processing unit is also used to create a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage the data in the data lake.

[0033] In a possible implementation manner, the acquisition unit is specifically used to acquire the historical written data volume of one or more historical partitions, predict the service requirement information according to the historical written data volume of one or more historical partitions, and the service requirement information includes the predicted data volume of the data to be written in the specified partition. The processing unit is specifically used to determine the number of buckets in the specified partition according to the predicted data volume of the data to be written in the specified partition and the bucket size of the specified partition.

[0034] In a possible implementation manner, the transceiver unit is used to receive the data to be written. When the processing unit determines that there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, the processing unit is also used to automatically determine the number of buckets in one or more partitions based on the partition statistical information, and the partition statistical information includes the written data volume of the data to be written in one or more partitions.

[0035] In a possible implementation manner, the processing unit is specifically used to perform an aggregation operation on the data to be written, determine the number of data items in each partition, sample the data to be written, determine the average size of each data item, determine the partition statistical information according to the number of data items and the average size of each data item, and calculate the number of buckets in each of one or more partitions based on the partition statistical information and the bucket size of each partition.

[0036] In a possible implementation manner, the processing unit is specifically used to write the pre-bucketing strategy in the metadata area of the data table.

[0037] In a possible implementation manner, the processing unit is also used to generate a configuration file based on the pre-bucketing strategy, the configuration file includes the pre-bucketing strategy of one or more partitions, and write the loading path of the configuration file in the metadata area of the data table.

[0038] In a possible implementation manner, when multiple pre-bucketing strategies are configured for a partition of the data lake, the processing unit is also used to determine the number of buckets in a partition based on the intersection of the multiple pre-bucketing strategies.

[0039] Fourthly, an embodiment of the present application provides a bucketing device for a data lake. The device includes a transceiver unit and a processing unit. The transceiver unit is configured to receive one or more policy key-value pairs input by a user. The policy key-value pairs are used to indicate the number of buckets for different partitions. The types of the policy key-value pairs include one or more of: direct matching, regular expression matching, and range matching. The processing unit is configured to generate a pre-bucketing policy for one or more partitions in the data lake based on the one or more policy key-value pairs. The pre-bucketing policy is used to indicate the number of buckets for different partitions in the data lake. The processing unit is further configured to create a data table in the data lake based on the pre-bucketing policy. The data table is used to manage the data in the data lake.

[0040] In a possible implementation, the processing unit is specifically configured to write the pre-bucketing policy in the metadata area of the data table.

[0041] In a possible implementation, the processing unit is specifically configured to generate a configuration file based on the pre-bucketing policy. The configuration file includes the pre-bucketing policies for one or more partitions, and write the loading path of the configuration file in the metadata area of the data table.

[0042] In a possible implementation, when multiple pre-bucketing policies are configured for a partition of the data lake, the processing unit is further configured to determine the number of buckets for a partition based on the intersection of the multiple pre-bucketing policies.

[0043] Fifthly, an embodiment of the present application provides a computing device. The computing device includes a processor. The processor is coupled to a memory. The processor is configured to store instructions. When the instructions are executed by the processor, the computing device is caused to execute the method described in the first aspect or any possible implementation of the first aspect, or the computing device is caused to execute the method described in the second aspect or any possible implementation of the second aspect.

[0044] Sixthly, an embodiment of the present application provides a computing device cluster. The computing device cluster includes one or more computing devices. The computing device includes a processor. The processor is coupled to a memory. The processor is configured to store instructions. When the instructions are executed by the processor, the computing device cluster is caused to execute the method described in the first aspect or any possible implementation of the first aspect, or the computing device cluster is caused to execute the method described in the second aspect or any possible implementation of the second aspect.

[0045] Seventhly, an embodiment of the present application provides a computer-readable storage medium, on which instructions are stored. When the instructions are executed, the computer is caused to execute the method described in the first aspect or any possible implementation of the first aspect, or the computer is caused to execute the method described in the second aspect or any possible implementation of the second aspect.

[0046] In an eighth aspect, an embodiment of the present application provides a computer program product. The computer program product includes instructions that, when executed, cause a computer to implement the method described in the first aspect or any possible implementation manner of the first aspect above, or cause a computer to implement the method described in the second aspect or any possible implementation manner of the second aspect above.

[0047] It can be understood that the beneficial effects that can be achieved by any of the above-provided data lake bucketing devices, computing devices, computing device clusters, computer-readable media, or computer program products can refer to the beneficial effects in the corresponding methods, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 FIG. is a schematic diagram of the system architecture of a data lake system provided by an embodiment of the present application;

[0049] Figure 2 FIG. is a schematic flowchart of a method for bucketing a data lake provided by an embodiment of the present application;

[0050] Figure 3 FIG. is a schematic diagram for determining service requirement information based on historical write data volume provided by an embodiment of the present application;

[0051] Figure 4 FIG. is a schematic flowchart of another method for bucketing a data lake provided by an embodiment of the present application;

[0052] Figure 5 FIG. is a schematic flowchart of another method for bucketing a data lake provided by an embodiment of the present application;

[0053] Figure 6 FIG. is a schematic flowchart of another method for bucketing a data lake provided by an embodiment of the present application;

[0054] Figure 7 FIG. is a schematic flowchart of another method for bucketing a data lake provided by an embodiment of the present application;

[0055] Figure 8 FIG. is a schematic diagram of the structure of a data lake bucketing device provided by an embodiment of the present application;

[0056] Figure 9 FIG. is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0057] Figure 10 FIG. is a schematic diagram of the structure of a computing device cluster provided by an embodiment of the present application;

[0058] Figure 11 FIG. is a schematic diagram of the structure of another computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The embodiments of the present application provide a method and device for bucketing a data lake, which are used to improve the bucketing flexibility of the data lake and the data query and writing efficiency.

[0060] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0061] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0062] First, some terms involved in the embodiments of the present application are introduced to facilitate those skilled in the art to understand the technical solutions.

[0063] A data lake is a centralized data repository for storing a large amount of data, and the stored data can be structured or unstructured data. Data lakes usually adopt distributed storage and processing technologies, such as Hadoop, Spark, etc., to support the storage, management and analysis of data.

[0064] A partition is a partition divided in a data lake according to certain rules. For example, it is partitioned and managed according to time, geographical location, business department, etc. to improve query performance and data management flexibility. A partition can include one or more buckets.

[0065] Bucket writing is a data storage and management technology. During the bucket writing process, data will be dispersed and stored in different buckets, and these buckets are usually divided according to certain rules. For example, they are hashed into buckets according to a certain column of the data or bucketed according to a range. By storing data in buckets, the data can be more evenly distributed in physical storage, thereby alleviating performance problems caused by excessive data in a single bucket.

[0066] Static bucketing means that data is written into buckets using a fixed number of buckets during the bucket writing process. Dynamic bucketing means that during the data bucketing writing process, the number of buckets used can vary dynamically rather than being completely fixed.

[0067] The number of buckets refers to the number of buckets into which data is to be written when writing into buckets.

[0068] The bucket size refers to the size of data stored in each bucket.

[0069] To make the technical solution of this application clearer and easier to understand, the system architecture of this application will be introduced below with reference to the accompanying drawings.

[0070] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the system architecture of a data lake system provided by an embodiment of this application. In the Figure 1 illustrated example, the data lake system 10 includes a data lake management subsystem 101, a computing engine subsystem 102, and a centralized storage subsystem 103. Among them, the centralized storage subsystem 103 includes one or more partitions, and each partition contains multiple data buckets. The specific functions of each subsystem in the data lake system 10 will be introduced below.

[0071] The data lake management subsystem 101 is used to implement the data management function of the data lake system 10. The data management function includes a basic management function and an extended management function. Among them, the basic management function includes, for example, data access, access control, task management, and metadata management, etc. The basic management function includes, for example, data governance, quality management, and asset catalog management, etc.

[0072] Among them, in the basic management function, data access means that the data lake management subsystem 101 can interface with various external heterogeneous data sources to implement data extraction and migration, etc. Access control means that the data lake management subsystem 101 can set different data access permissions for different users or user groups. Task management means that the data lake management subsystem 101 can manage and orchestrate tasks in the data lake to achieve automated execution and scheduling of tasks. Metadata management means that the data lake management subsystem 101 can manage and maintain metadata information in the data lake. The metadata information includes the structure, attributes, association relationships, etc. of the data.

[0073] In the extended management function, data governance such as the data lake management subsystem 101 can clean, integrate, and standardize data, and define usage rules for data, etc. Quality management such as the data lake management subsystem 101 can perform quality assessment and management on the data in the data lake, provide data quality reports and monitoring functions, and help users understand the quality status and improvement directions of the data. Asset catalog management such as the data lake management subsystem 101 can establish a data asset catalog, classify and label the data, and facilitate users to quickly find the required data resources.

[0074] The computing engine subsystem 102 is used to process and analyze the data in the data lake system 10. The computing engine subsystem 102 can provide various computing functions, including batch processing, stream computing, interactive query, and machine learning, etc. Among them, for the batch processing function, for example, the computing engine subsystem 102 can perform batch processing and analysis on a large amount of data. For stream computing, for example, the computing engine subsystem 102 can process real-time data streams and perform real-time analysis and processing on the data. For interactive query, for example, the computing engine subsystem 102 can provide an interactive query function, and users can perform data retrieval and analysis through query statements. For the machine learning function, for example, the computing engine subsystem 102 can support the training and prediction tasks of machine learning algorithms and provide the ability for data analysis and prediction.

[0075] The centralized storage subsystem 103 is used to provide the data storage function in the database system 10. The centralized storage subsystem 103 includes one or more devices. Among them, one or more devices can be logically divided into different partitions, and each partition can store a set of data with similar characteristics or business meanings. For example, the centralized storage subsystem 103 partitions according to business attributes such as date, geographical location, product category, etc.

[0076] Among them, the centralized storage subsystem 103 can also set buckets within the partition, that is, the centralized storage subsystem 103 can perform a finer-grained division according to the specified rules for partitioning. Each bucket in the partition contains a part of the data in the partition. Among them, the specified rules include bucketing based on hash values, range values, or sampling values. For example, the centralized storage subsystem 103 performs a hash calculation on the data in the data table according to the value of a certain column, and then puts the data with the same hash value into the same bucket file.

[0077] Based on Figure 1 the data lake system 10 shown, the present application also provides a bucketing method for the data lake. Next, in combination with embodiments, the bucketing method for the data lake provided by the embodiments of the present application will be introduced.

[0078] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a bucketing method for a data lake provided by an embodiment of the present application. InFigure 2 In the example shown, the method includes the following steps:

[0079] 201. Obtain business requirement information, where the business requirement information is used to indicate the predicted data volume of the data to be written in one or more partitions in the data lake, and the data to be written is for writing into the data buckets of one or more partitions.

[0080] The data lake system 10 obtains business requirement information, where the business requirement information is used to indicate the predicted data volume of the data to be written in one or more partitions in the data lake, and the data to be written is for writing into the data buckets of one or more partitions. Specifically, the data lake system 10 may receive the business requirement information input by the user, or the data lake system 10 may also determine the business requirement information according to the historical written data volume, and no specific limitation is made.

[0081] In a possible implementation manner, during the process that the data lake system 10 determines the business requirement information according to the historical written data volume, the data lake system 10 first obtains the historical written data volume of one or more historical partitions, and predicts the business requirement information according to the historical written data volume of one or more historical partitions. The business requirement information includes the predicted data volume of the data to be written in one or more specified partitions.

[0082] Specifically, the data lake system 10 calculates the predicted data volume of the data to be written in the specified partition according to the historical written data volume of one or more historical partitions and the weight coefficients corresponding to different historical partitions, or the data lake system 10 calculates the predicted data volume of the data to be written in the specified partition according to the historical written data volume of one or more historical partitions and the number of historical partitions, that is, determines the predicted data volume of the data to be written according to the average value of the historical written data volume of the historical partitions.

[0083] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a data lake system for determining business requirement information provided by an embodiment of the present application. In Figure 3 the example shown, the data lake system 10 obtains the historical written data volume of one or more historical partitions. For example, the data lake system 10 obtains that the written data volumes of historical partition 1, historical partition 2, historical partition 3,..., historical partition n are S1, S2, S3,..., Sn respectively.

[0084] In Figure 3In the example shown, after the data lake system 10 obtains the historical write data volume of one or more historical partitions, it calculates the predicted data volume of the data to be written in the current partition according to the historical write data volume of one or more historical partitions and the weight coefficients corresponding to different historical partitions. For example, if the weight coefficients corresponding to historical partition 1, historical partition 2, historical partition 3, …, historical partition n are w1, w2, w3, …, wn, where w1, w2, w3, …, wn satisfy the following formula:

[0085]

[0086] In Figure 3 the example shown, the data lake system 10 can calculate the predicted data volume of the data to be written in all specified partitions according to the historical write data volume of historical partitions 1 to n and the weight coefficients w1 to wn corresponding to different historical partitions. The predicted data volume s of the data to be written satisfies the following formula:

[0087]

[0088] In Figure 3 the example shown, when the weight coefficients corresponding to different historical partitions are the same, the predicted write data volume of the data to be written in each partition is s / n.

[0089] 202. Determine the pre-bucketing strategy for one or more partitions according to the service requirement information, where the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake.

[0090] The data lake system 10 determines the pre-bucketing strategy for one or more partitions according to the service requirement information, where the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake. Specifically, the data lake system 10 calculates the number of buckets for one or more partitions based on the predicted data volume of the data to be written in one or more partitions and the bucket size of one or more partitions, where the bucket size is a preset value.

[0091] Please continue to refer to Figure 3 In Figure 3 the example shown, the data lake system 10 can calculate the predicted data volume sn of the data to be written in the specified partition according to the historical write data volume of historical partitions 1 to n and the weight coefficients w1 to wn corresponding to different historical partitions. If the bucket size of the partition is k, the data lake system 10 can calculate the pre-bucketing strategy for this partition based on the predicted data volume sn of the data to be written in this partition and the bucket size k, and the pre-bucketing strategy is that the number of buckets for this partition is sn / k.

[0092] It can be understood that when the business requirement information is the predicted data volume of the data to be written in multiple partitions, the data lake system 10 can determine the corresponding pre-bucketing strategy for each partition in the multiple partitions, that is, the data lake system 10 can determine the number of buckets corresponding to each partition in the multiple partitions.

[0093] In a possible implementation manner, when the data lake system 10 determines the pre-bucketing strategy for each partition according to the business requirement information, the data lake system 10 generates one or more partition policy key-value pairs based on the business requirement information. The policy key-value pairs are used to indicate the number of buckets in different partitions. Among them, the policy key-value pair includes two parts: a key and a value. The key is used to specify the partition, and the value is used to specify the number of buckets corresponding to the partition. The types of policy key-value pairs include one or more of the following: direct matching, regular matching, and range matching. The data lake system 10 generates the pre-bucketing strategy for one or more partitions based on the policy key-value pairs. The following specifically introduces several types of the above policy key-value pairs:

[0094] Among them, direct matching is also called exact matching, which means that the data lake system 10 directly specifies the number of buckets in the partition. For example, "p = 2022-10-11 = 4" is a configuration statement corresponding to the direct matching type, indicating that the number of buckets in the partition corresponding to the day "2022-10-11" is 4. Another example, "p = 2022-11-11 = 10" is also a configuration statement corresponding to the direct matching type, indicating that the number of buckets in the partition corresponding to the day "2022-11-11" is 11.

[0095] Regular matching means that the data lake system 10 specifies the number of buckets in the partitions that meet the conditions through a string. For example, "p = ****-11-11 = 9" is a configuration statement corresponding to the regular matching type, indicating that the number of buckets in the partition corresponding to the day "11-11" of each year is 9.

[0096] Range matching means that the data lake system 10 specifies the number of buckets in a range of partitions. For example, "p > 2023-11-12 = 5" is a configuration statement corresponding to the range matching type, indicating that the number of buckets in the partition corresponding to each day after "2023-11-12" is 5.

[0097] In a possible implementation, if multiple pre-bucketing policies are configured for a partition of the data lake, the data lake system 10 determines the number of buckets for a partition based on the intersection of the multiple pre-bucketing policies. For example, if the partition "2023-12-12" corresponds to two pre-bucketing policies, including "p > 2023-11-12 = 5" and "p = ****-12-12 = 8", then the data lake system 10 determines the number of buckets for the partition "2023-12-12" based on the intersection of "p > 2023-11-12 = 5" and "p = ****-12-12 = 8", and the number of buckets for the partition "2023-12-12" is 8.

[0098] 203. Create a data table in the data lake based on the pre-bucketing policy, and the data table is used to manage the data in the data lake.

[0099] After the data lake system 10 determines the pre-bucketing policy, it creates and saves a data table in the data lake based on the pre-bucketing policy, and the data table is used to manage the data in the data lake. After the data lake system 10 receives the write data for a certain partition, it can read the pre-bucketing policy corresponding to the partition from the data table and perform bucketed writing according to the number of buckets corresponding to the pre-bucketing policy.

[0100] Please refer to Figure 4 , Figure 4 which is a schematic diagram of writing data into the data lake by data bucketing provided by an embodiment of this application. In Figure 4 the illustrated example, before the data lake system 10 receives the write data for one or more partitions, the user needs to configure the pre-bucketing policy for one or more partitions in the data lake system 10 and create a data table.

[0101] In Figure 4 the illustrated example, when the data lake system 10 receives the write data for a certain partition, if the pre-bucketing policy has been configured for this partition, the data lake system 10 loads the pre-bucketing policy corresponding to this partition from the metadata area and performs bucketed writing according to the number of buckets corresponding to the pre-bucketing policy. For example, when the data lake system 10 receives the write data for partition B, since the pre-bucketing policy has been configured for partition B, the data lake system 10 loads the pre-bucketing policy corresponding to partition B from the metadata area and performs bucketed writing according to the number of buckets corresponding to the pre-bucketing policy. The number of buckets corresponding to the pre-bucketing policy of partition B is 2. Therefore, the data lake system 10 writes data into bucket 1 and bucket 2 in partition B respectively.

[0102] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of writing real-time data into the data lake provided by an embodiment of this application. In Figure 5In steps a to c and steps e to f of the illustrated example, the data lake system 10 receives write data for one or more partitions. After that, the data lake system 10 loads and determines whether there is a pre-bucketing policy for the partition corresponding to the write data. If there is a pre-bucketing policy for the partition corresponding to the write data, the data lake system 10 saves the pre-bucketing policy corresponding to the partition to the metadata area of the partition, and at the same time buckets the partition according to the number of buckets in the pre-bucketing policy and performs bucketed writing.

[0103] In a possible implementation, when the data lake system 10 creates a data table in the data lake based on the pre-bucketing policy, the data lake system 10 can write the pre-bucketing policy in the metadata area of the data table. That is, the data lake system 10 can directly write the policy key-value pair corresponding to the pre-bucketing policy in the metadata area of the data table.

[0104] For example, when the data lake system 10 creates a data table, the data lake system 10 directly writes the pre-bucketing policy in the "tblproperties attribute" of the table creation code and saves it to the metadata area of the data table. When writing data in the data lake system 10 later, the pre-bucketing policy is loaded from the metadata of the data table, and the number of buckets for the partition is automatically set according to the pre-bucketing policy. The example code for directly writing the pre-bucketing policy in the table creation code is as follows:

[0105]

[0106] In a possible implementation, the data lake system 10 generates a configuration file based on the pre-bucketing policy. The configuration file includes the pre-bucketing policies for one or more partitions. When the data lake system 10 creates a data table in the data lake based on the pre-bucketing policy, the loading path of the configuration file is written in the metadata area of the data table.

[0107] For example, when the data lake system 10 creates a data table, the data lake system 10 can also write the loading path of the pre-bucketing policy configuration file in the "tblproperties attribute" of the table creation code. When writing data in the data lake system 10 later, the pre-bucketing policy configuration file is loaded from the metadata area of the data table, and the number of buckets for the partition is automatically set according to the configuration file. The example code for writing the loading path of the pre-bucketing policy configuration file in the table creation code is as follows:

[0108] CREATE TABLE LAKETABLE CLUSTED BY(uuid)INTO N BUCKETS

[0109] TBLPROPERTIES(

[0110] bucket.strategy.location = ' / tmp / bucketNumStategy'

[0111] In a possible implementation, after the data lake system 10 receives the data to be written, when it is determined that there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, the data lake system 10 infers the number of buckets for the partitions based on the data to be written, that is, the data lake system 10 automatically determines the number of buckets for one or more partitions based on the partition statistical information, and the partition statistical information includes the amount of data written in one or more partitions of the data to be written.

[0112] Please continue to refer to Figure 4 , in Figure 4 In the example shown, when the data lake system 10 receives the write data for a certain partition, if there is no pre-bucketing strategy configured for this partition, the data lake system 10 infers the number of buckets for the partition based on the data to be written and automatically determines the number of buckets for one or more partitions. For example, when the data lake system 10 receives the write data for partition A, since there is no pre-configured pre-bucketing strategy for partition A, the data lake system 10 infers that the number of buckets for partition A is 2 based on the data to be written for partition A. The data lake system 10 buckets partition A into bucket 1 and bucket 2 according to the number of buckets 2, and performs bucketed writing in bucket 1 and bucket 2 of partition A.

[0113] Please continue to refer to Figure 5 , in Figure 5 In step d of the example shown, when the data lake system 10 determines whether there is a pre-bucketing strategy for the partition corresponding to the write data, if there is no pre-bucketing strategy for the partition corresponding to the data to be written, the data lake system 10 infers the number of buckets based on the amount of the data to be written, buckets the partition according to the inferred number of buckets, and performs bucketed writing.

[0114] In a possible implementation, during the process that the data lake system 10 automatically determines the number of buckets for a partition based on the partition statistical information, the data lake system 10 performs an aggregation operation on the data to be written to determine the number of data items in each partition, and samples the data to be written to determine the average size of each data item. The data lake system 10 determines the partition statistical information based on the number of data items and the average size of each data item, and the partition statistical information is the amount of data written in each partition. The data lake system 10 calculates the number of buckets for each partition in one or more partitions based on the partition statistical information and the bucket size of each partition.

[0115] Please refer to Figure 6 , Figure 6 is a schematic flow diagram of inferring the number of buckets provided by an embodiment of the present application. In Figure 6In step a of the illustrated example, when there is no pre-bucketing strategy for the partition where the data to be written is to be written, the data lake system 10 performs an aggregation operation on the data to be written and calculates the number of data items in each partition. For example, the data to be written can be data to be written to one or more partitions, and the data lake system 10 calculates the number of data items in each partition. For example, the number of data items to be written in the first partition is r1, the number of data items to be written in the second partition is r2, …, and the number of data items to be written in the nth partition is rn.

[0116] In Figure 6 In steps b to c of the illustrated example, the data lake system 10 can also sample the data to be written and determine the average size of each data item. For example, the data lake system 10 samples the data to be written and determines that the average size of each data item is x. After the data lake system 10 calculates the average data size and the number of data items in each partition, it calculates partition statistics based on the number of data items and the average size of each data item. The partition statistics are the data write volume for each partition.

[0117] For example, the data lake system 10 performs the calculation x * rn to obtain the data volume of the data to be written in the first partition as s1, the data volume of the data to be written in the second partition as s2, …, and the data volume of the number of data items to be written in the nth partition as sn.

[0118] In Figure 6 In step d of the illustrated example, after the data lake system 10 calculates the partition statistics, it calculates the number of buckets for each partition in one or more partitions based on the partition statistics and the bucket size of each partition. For example, if the bucket size of each bucket is 1 GB, the data lake system 10 calculates that the number of buckets for each partition is sn / 1 GB.

[0119] In Figure 6 In steps e to f of the illustrated example, after the data lake system 10 calculates the number of buckets for each partition in one or more partitions, the data lake system 10 saves the calculated number of buckets to the metadata file in the corresponding partition directory and performs bucketed writing according to the number of buckets in the partition directory. An example of the directory structure of the metadata file corresponding to the partition directory is as follows:

[0120] …. / table / 2022-01-05 / bucket.info

[0121] …. / table / 2022-01-06 / bucket.info

[0122] It should be noted that in the above Figure 2In the illustrated embodiment, the data lake system 10 can automatically generate one or more partition policy key-value pairs based on business requirement information. In the embodiments of the present application, the data lake system 10 can be directly configured with policy key-value pairs by the user, and a pre-bucketing policy for one or more partitions in the data lake is generated based on the policy key-value pairs configured by the user. The following will be described in conjunction with Figure 7 the illustrated embodiment:

[0123] Please refer to Figure 7 , Figure 7 , which is a flowchart of another method for bucketing a data lake provided by the embodiments of the present application. In Figure 2 the illustrated example, the method includes the following steps:

[0124] 701. Receive one or more policy key-value pairs input by the user, where the policy key-value pairs are used to indicate the number of buckets for different partitions, and the types of the policy key-value pairs include one or more of: direct matching, regular expression matching, and range matching.

[0125] The data lake system 10 receives one or more policy key-value pairs input by the user, and the policy key-value pairs are used to indicate the number of buckets for different partitions. That is, the user can directly configure the pre-bucketing policy for one or more partitions through the policy key-value pairs. When determining the policy key-value pairs, the user can determine the policy key-value pairs based on business requirement information, and the business requirement information includes the predicted data volume of the data to be written in one or more partitions in the data lake.

[0126] In a possible implementation, the user can write the pre-bucketing policy in the metadata area of the data table, that is, the user can directly write the policy key-value pairs corresponding to the pre-bucketing policy in the metadata area of the data table.

[0127] In a possible implementation, the data lake system 10 can receive the configuration file of the pre-bucketing policy sent by the user and write the loading path of the configuration file in the metadata area of the data table.

[0128] 702. Generate a pre-bucketing policy for one or more partitions in the data lake based on one or more policy key-value pairs, where the pre-bucketing policy is used to indicate the number of buckets for different partitions in the data lake.

[0129] 703. Create a data table in the data lake based on the pre-bucketing policy, where the data table is used to manage the data in the data lake.

[0130] The methods executed by the data lake system 10 in steps 702 and 703 of the embodiments of the present application are similar to the methods executed by the data lake system 10 in steps 202 and 203 in the Figure 2 illustrated embodiment, and will not be described in detail here.

[0131] It can be understood that Figure 7 the illustrated embodiment andFigure 2 The difference between the illustrated embodiments is that Figure 2 in the illustrated embodiment, the data lake system 10 can obtain business requirement information and generate a pre-bucketing policy based on the business requirement information Figure 7 in the illustrated embodiment, the pre-bucketing policy is obtained by directly configuring the policy key-value pairs by the user Figure 7 the illustrated embodiment may also include Figure 2 any possible implementation manner in the illustrated embodiment, which will not be elaborated herein

[0132] As can be seen from the above embodiments, in the embodiments of the present application, the data lake system can configure different pre-bucketing policies for different partitions according to the business requirement information, and the data lake system flexibly sets the number of buckets for different partitions based on the predicted data volume corresponding to the business requirement, thereby improving the configuration accuracy of the number of buckets for different partitions and the flexibility of bucketing, and further improving the data query and writing efficiency of the data lake.

[0133] Based on the above method embodiments, the embodiments of the present application also provide a bucketing device for a data lake. The bucketing device for a data lake provided by the embodiments of the present application will be specifically introduced below.

[0134] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a bucketing device for a data lake provided by an embodiment of the present application. In Figure 8 the illustrated example, the bucketing device 800 for a data lake is used to implement each step executed by the data lake system 10 in the above embodiments. The bucketing device 800 for a data lake includes an acquisition unit 801, a processing unit 802, and a transceiver unit 803.

[0135] In an example of the embodiments of the present application, the bucketing device 800 for a data lake is used to implement each step executed by the data lake system 10 in the above Figure 3 illustrated embodiment.

[0136] Among them, the acquisition unit 801 is used to acquire business requirement information, and the business requirement information is used to indicate the predicted data volume of the data to be written in one or more partitions in the data lake. The data to be written is the data to be written into the data buckets of one or more partitions. The processing unit 802 is used to determine the pre-bucketing policy for one or more partitions according to the business requirement information, and the pre-bucketing policy is used to indicate the number of buckets for different partitions in the data lake. The processing unit 802 is further used to create a data table in the data lake based on the pre-bucketing policy, and the data table is used to manage the data in the data lake.

[0137] In a possible implementation, the obtaining unit 801 is specifically configured to obtain the historical write data volume of one or more historical partitions, predict the service demand information according to the historical write data volume of the one or more historical partitions, and the service demand information includes the predicted data volume of the data to be written in the specified partition. The processing unit 802 is specifically configured to determine the number of buckets in the specified partition according to the predicted data volume of the data to be written in the specified partition and the bucket size of the specified partition.

[0138] In a possible implementation, the transceiver unit 803 is configured to receive the data to be written. The processing unit 802 is further configured to, when there is no pre-bucketing policy for one or more partitions where the data to be written is to be written, automatically determine the number of buckets in the one or more partitions based on the partition statistical information, and the partition statistical information includes the write data volume of the data to be written in the one or more partitions.

[0139] In a possible implementation, the processing unit 802 is specifically configured to perform an aggregation operation on the data to be written, determine the number of data entries in each partition, sample the data to be written, determine the average size of each piece of data, determine the partition statistical information according to the number of data entries and the average size of each piece of data, and calculate the number of buckets in each of the one or more partitions based on the partition statistical information and the bucket size of each partition.

[0140] In a possible implementation, the processing unit 802 is specifically configured to write a pre-bucketing policy in the metadata area of the data table.

[0141] In a possible implementation, the processing unit 802 is further configured to generate a configuration file based on the pre-bucketing policy, the configuration file includes the pre-bucketing policies of one or more partitions, and write the loading path of the configuration file in the metadata area of the data table.

[0142] In a possible implementation, the processing unit 802 is further configured to, when multiple pre-bucketing policies are configured for one partition of the data lake, determine the number of buckets in one partition based on the intersection of the multiple pre-bucketing policies.

[0143] In another example of the embodiments of the present application, the bucketing device 800 of the data lake is used to implement each step performed by the data lake system 10 in the embodiments shown above Figure 7 as shown.

[0144] Among them, the transceiver unit 803 is used to receive one or more policy key-value pairs input by the user. The policy key-value pairs are used to indicate the number of buckets in different partitions. The types of the policy key-value pairs include one or more of the following: direct matching, regular matching, and range matching. The processing unit 802 is used to generate a pre-bucketing policy for one or more partitions in the data lake based on one or more policy key-value pairs. The pre-bucketing policy is used to indicate the number of buckets in different partitions of the data lake. The processing unit 802 is also used to create a data table in the data lake based on the pre-bucketing policy. The data table is used to manage the data in the data lake.

[0145] In a possible implementation manner, the processing unit 802 is specifically used to write the pre-bucketing policy in the metadata area of the data table.

[0146] In a possible implementation manner, the processing unit 802 is specifically used to generate a configuration file based on the pre-bucketing policy. The configuration file includes the pre-bucketing policies for one or more partitions, and write the loading path of the configuration file in the metadata area of the data table.

[0147] In a possible implementation manner, the processing unit 802 is also used to, when multiple pre-bucketing policies are configured for one partition of the data lake, determine the number of buckets for one partition based on the intersection of the multiple pre-bucketing policies.

[0148] It can be understood that the acquisition unit 801, the processing unit 802, and the transceiver unit 803 in the bucketing device 800 of the data lake can be used as functional modules and Figure 1 there is a mapping with each module in the data lake system 10, so as to implement the functions of each module in the data lake system 10.

[0149] It should be understood that the division of units in the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And the units in the device can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by processing elements and partially implemented in the form of hardware. For example, each unit can be a separately established processing element, or can be integrated in a certain chip of the device. In addition, it can also be stored in the memory in the form of a program and called and executed by a certain processing element of the device to perform the functions of the unit. In addition, all or part of these units can be integrated together or can be independently implemented. The processing element mentioned here can also be a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented through the integrated logic circuit in the processor element or in the form of software called by the processing element.

[0150] It should be noted that, for the above method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0151] Other reasonable combinations of steps that can be thought of by those skilled in the art based on the above description also fall within the protection scope of this application. Secondly, those skilled in the art should also be familiar that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0152] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computing device provided by an embodiment of this application. As Figure 9 shown, the computing device 900 includes: a processor 901, a memory 902, a communication interface 903, and a bus 904. The processor 901, the memory 902, and the communication interface 903 are coupled through a bus (not labeled in the figure). The memory 902 stores instructions. When the execution instructions in the memory 902 are executed, the computing device 900 executes the method executed by the computing device in the above method embodiments.

[0153] The computing device 900 can be one or more integrated circuits configured to implement the above method. For example: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. Again, when the units in the device can be implemented in the form of a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. Again, these units can be integrated together to be implemented in the form of a system-on-a-chip (SOC).

[0154] The processor 901 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0155] The memory 902 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0156] The executable program code is stored in the memory 902, and the processor 901 executes the executable program code to implement the functions of the foregoing units or modules respectively, so as to implement the method for bucketing the data lake as described above. That is to say, instructions for executing the method for bucketing the data lake as described above are stored on the memory 902.

[0157] The communication interface 903 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or communication networks.

[0158] In addition to including a data bus, the bus 904 may further include a power bus, a control bus, a status signal bus, etc. The bus may be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0159] Please refer to Figure 10 , Figure 10 for a schematic diagram of a computing device cluster provided by an embodiment of the present application. As Figure 10 shown, the computing device cluster 1000 includes at least one computing device 900.

[0160] As Figure 10 shown, the computing device cluster 1000 includes at least one computing device 900. Instructions for executing the above-mentioned bucketing method of the data lake may be stored in the memories 902 of one or more of the computing devices 900 in the computing device cluster 1000.

[0161] In some possible implementation manners, partial instructions for executing the above-mentioned bucketing method of the data lake may also be stored separately in the memories 902 of one or more of the computing devices 900 in the computing device cluster 1000. In other words, a combination of one or more computing devices 900 may jointly execute the instructions for executing the above-mentioned bucketing method of the data lake.

[0162] It should be noted that the memories 902 in different computing devices 900 in the computing device cluster 1000 may store different instructions, respectively for executing partial functions of the above-mentioned bucketing device of the data lake. That is, the instructions stored in the memories 902 of different computing devices 900 may implement the functions of one or more modules among the obtaining unit, the processing unit, and the transceiver unit.

[0163] In some possible implementation manners, one or more computing devices 900 in the computing device cluster 1000 may be connected through a network. Wherein, the network may be a wide area network or a local area network, etc.

[0164] Please refer to Figure 11, Figure 11 The figure is a schematic diagram showing the network connection of computer devices in a computer cluster provided by an embodiment of the present application. As Figure 11 shown, two computing devices 900A and 900B are connected through a network. Specifically, they are connected to the network through communication interfaces in each computing device.

[0165] In a possible implementation, instructions for implementing the functions of the acquisition unit and the transceiver unit are stored in the memory of computing device 900A. At the same time, instructions for implementing the function of the processing unit are stored in the memory of computing device 900B.

[0166] It should be understood that Figure 11 the functions of computing device 900A shown in

[0167] In another embodiment of the present application, a computer-readable storage medium is further provided. Computer-executable instructions are stored in the computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the method performed by the data lake system in the above method embodiment.

[0168] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer-executable instructions, and the computer-executable instructions are stored in a computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the method performed by the data lake system in the above method embodiment.

[0169] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0170] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0171] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.

Claims

1. A bucketing method for a data lake, characterized in that: include: Acquire business demand information, where the business demand information is used to indicate a predicted amount of data to be written in one or more partitions in the data lake, where the data to be written is used to write into data buckets of the one or more partitions; Determine a pre-bucketing strategy for the one or more partitions according to the business requirement information, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; A data table is created in the data lake based on the pre-bucketing strategy, where the data table is used to manage data in the data lake.

2. The method according to claim 1, characterized in that The obtaining of business demand information includes: Get the historical written data volume of one or more historical partitions; Predicting the business demand information according to the historical written data volume of the one or more historical partitions, the business demand information including the predicted data volume of the data to be written to the specified partition; The determining of the pre-bucketing strategy of the one or more partitions according to the business requirement information includes: The number of buckets of the specified partition is determined according to the predicted data volume of the data to be written to the specified partition and the bucket size of the specified partition.

3. The method according to claim 1, characterized in that The method further comprises: receiving the data to be written; When it is determined that there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, the number of buckets of the one or more partitions is automatically determined based on partition statistical information, and the partition statistical information includes the amount of data written by the data to be written in the one or more partitions.

4. The method according to claim 3, characterized in that The automatically determining the number of buckets of the partition based on the partition statistics information includes: Aggregating the data to be written to determine the number of data entries in each partition; Sampling the data to be written to determine the average size of each piece of data; Determine the partition statistical information according to the number of data pieces and the average size of each data piece; The number of buckets of each of the one or more partitions is calculated based on the partition statistics and the bucket size of each partition.

5. The method according to any one of claims 1 to 4, characterized in that The creating a data table in the data lake based on the pre-bucketing strategy includes: The pre-bucketing strategy is written into the metadata area of ​​the data table.

6. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The creating a data table in the data lake based on the pre-bucketing strategy includes: The loading path of the configuration file is written into the metadata area of ​​the data table.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.

8. A method for bucketing a data lake, characterized in that: include: Receive one or more policy key-value pairs input by a user, where the policy key-value pairs are used to indicate the number of buckets in different partitions, and the type of the policy key-value pairs includes one or more of: direct match, regular match, and range match; Generate a pre-bucketing strategy for one or more partitions in the data lake based on the one or more policy key-value pairs, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; A data table is created in the data lake based on the pre-bucketing strategy, where the data table is used to manage data in the data lake.

9. The method according to claim 8, characterized in that The creating a data table in the data lake based on the pre-bucketing strategy includes: The pre-bucketing strategy is written into the metadata area of ​​the data table.

10. The method according to claim 8, characterized in that: The method further comprises: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The creating a data table in the data lake based on the pre-bucketing strategy includes: The loading path of the configuration file is written into the metadata area of ​​the data table.

11. The method according to any one of claims 8 to 10, characterized in that The method further comprises: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.

12. A bucketing device for a data lake, characterized in that: include: An acquisition unit, used to acquire business demand information, where the business demand information is used to indicate a predicted amount of data to be written in one or more partitions in the data lake, where the data to be written is used to be written into data buckets of the one or more partitions; A processing unit, configured to determine a pre-bucketing strategy for the one or more partitions according to the business requirement information, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; The processing unit is further used to create a data table in the data lake based on the pre-bucketing strategy, where the data table is used to manage the data in the data lake.

13. The device according to claim 12, characterized in that The acquisition unit is specifically used for: Get the historical written data volume of one or more historical partitions; Predicting the business demand information according to the historical written data volume of the one or more historical partitions, the business demand information including the predicted data volume of the data to be written to the specified partition; The processing unit is specifically used for: The number of buckets of the specified partition is determined according to the predicted data volume of the data to be written to the specified partition and the bucket size of the specified partition.

14. The device according to claim 12, characterized in that The device further comprises a transceiver unit, wherein the transceiver unit is used for: receiving the data to be written; The processing unit is also used to automatically determine the number of buckets of the one or more partitions based on partition statistical information when there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, and the partition statistical information includes the amount of data written by the data to be written in the one or more partitions.

15. The device according to claim 14, characterized in that The processing unit is specifically used for: Aggregating the data to be written to determine the number of data entries in each partition; Sampling the data to be written to determine the average size of each piece of data; Determine the partition statistical information according to the number of data pieces and the average size of each data piece; The number of buckets of each of the one or more partitions is calculated based on the partition statistics and the bucket size of each partition.

16. The device according to any one of claims 12 to 15, characterized in that The processing unit is specifically used for: The pre-bucketing strategy is written into the metadata area of ​​the data table.

17. The device according to any one of claims 12 to 15, characterized in that: The processing unit is also used for: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The loading path of the configuration file is written into the metadata area of ​​the data table.

18. The device according to any one of claims 12 to 17, characterized in that The processing unit is also used for: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.

19. A bucketing device for a data lake, characterized in that: include: A transceiver unit, configured to receive one or more policy key-value pairs input by a user, wherein the policy key-value pairs are used to indicate the number of buckets in different partitions, and the types of the policy key-value pairs include one or more of: direct matching, regular matching, and range matching; A processing unit, configured to generate a pre-bucketing strategy for one or more partitions in the data lake based on the one or more policy key-value pairs, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; The processing unit is further used to create a data table in the data lake based on the pre-bucketing strategy, where the data table is used to manage the data in the data lake.

20. The device according to claim 19, characterized in that The processing unit is specifically used for: The pre-bucketing strategy is written into the metadata area of ​​the data table.

21. The device according to claim 19, characterized in that The processing unit is specifically used for: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The loading path of the configuration file is written into the metadata area of ​​the data table.

22. The device according to any one of claims 19 to 21, characterized in that The processing unit is also used for: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.

23. A computing device, characterized in that The invention comprises a processor coupled to a memory, wherein the processor is used to store instructions, and when the instructions are executed by the processor, the computing device executes the method according to any one of claims 1 to 8, or the computing device executes the method according to any one of claims 9 to 11.

24. A computing device cluster, characterized in that: The method comprises at least one computing device, wherein the computing device comprises a processor, the processor is coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method of any one of claims 1 to 8, or the computing device cluster executes the method of any one of claims 9 to 11.

25. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed, the computer is caused to execute the method according to any one of claims 1 to 8, or the computer is caused to execute the method according to any one of claims 9 to 11.

26. A computer program product, comprising instructions, characterized in that: When the instructions are executed, the computer implements the method according to any one of claims 1 to 8, or the computer implements the method according to any one of claims 9 to 11.

Citation Information

Cited By

  • Storage space management method, device and equipment and computer readable storage medium

    CN121541829A

  • Storage space management method, device and equipment and computer readable storage medium

    CN121541829B