Bucketing method and apparatus for data lake
By dynamically configuring the number of buckets in the data lake partition, the problem of uneven data distribution caused by the fixed number of buckets in the existing technology is solved, and the efficiency of data query and writing is improved.
Patent Information
- Application Number
- PCT/CN2024/096958
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2024-06-03
- Publication Date
- 2025-05-22
AI Technical Summary
In the existing data lake bucket writing technology, the number of buckets in partitions is fixed, resulting in uneven distribution of data in different partitions, causing data skew, and reducing query and writing efficiency.
By obtaining business demand information, predicting the amount of data to be written in different partitions, dynamically configure the number of buckets in each partition, generating a pre-packet bucket strategy, and realizing flexible bucketing of the data lake.
Improve the flexibility of bucketing and the efficiency of data query and writing, ensure the uniform distribution of data in the partition, and avoid data skew.
Smart Images

Figure CN2024096958_22052025_PF_FP_ABST
Abstract
Description
A data lake bucketing method and device
[0001] This application claims priority to Chinese patent application No. 202311552902.1, filed with the State Intellectual Property Office of China on November 17, 2023, and entitled “A method and device for writing dynamic buckets”, and No. 202410268575.5, filed with the State Intellectual Property Office of China on March 8, 2024, and entitled “A method and device for bucketing a data lake”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of cloud computing, and in particular to a method and apparatus for bucketing a data lake. Background Art
[0003] With the continuous development of data lake services, users have increasingly higher requirements for the query and update performance of data lakes in some scenarios with large data read and write volumes. Due to the higher performance data query and update capabilities that bucketed writing technology can bring, more and more users tend to choose to use bucketed writing technology to write data in data lakes when building data lakes.
[0004] In current bucketed writing technology, a data lake contains multiple partitions, each of which contains multiple buckets. During data writing, the computing device needs to load the bucketing policy for each partition, which indicates the number of buckets in that partition. Currently, the computing device needs to manually configure the bucketing policy for each partition, and each partition in the current bucketing policy has a fixed number of buckets.
[0005] However, when users write data based on actual business, the written data may be unevenly distributed in different partitions. Therefore, using a fixed number of buckets for each partition will result in some buckets having too much data and some buckets having too little data, making the bucketing inflexible and prone to data skew, which further leads to low data query and writing efficiency.
[0006] Summary of the Invention
[0007] The present invention provides a data lake bucketing method. The data lake system can configure the number of buckets for different partitions according to business needs, thereby improving the flexibility of bucketing and the efficiency of data query and write. The present invention also provides a data lake bucketing device, computing device, computing device cluster, computer-readable storage medium, and computer program product corresponding to the data lake bucketing method.
[0008] In the first aspect, an embodiment of the present application provides a data lake bucketing method, which can be executed by a data lake system, or by a component of the data lake system, such as a processor, chip or chip system of the data lake system, or by a logic module or software that can implement all or part of the data lake system functions. The method provided in the first aspect includes: the data lake system obtains business demand information, the business demand information is used to indicate the predicted amount of data to be written in one or more partitions in the data lake, and the data to be written is used to write into the data buckets of one or more partitions. The data lake system determines a pre-bucketing strategy for one or more partitions based on the business demand information, and the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake, and the number of buckets for different partitions is different. The data lake system creates a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage the data in the data lake. After the data lake system receives the data to be written, it can perform bucketing writing based on the pre-bucketing strategy in the data table.
[0009] In the embodiment of the present application, the data lake system can configure different pre-bucketing strategies for different partitions based on business demand information, and the number of buckets for different partitions can be flexibly set. Compared with the current data lake system that sets a fixed number of buckets, the data lake system in the embodiment of the present application flexibly sets the number of buckets for different partitions based on the predicted data volume corresponding to business needs, thereby improving the configuration accuracy of the number of buckets for different partitions and the flexibility of bucketing, and further improving the data query and write efficiency of the data lake.
[0010] In one possible implementation, in the process of determining the pre-bucketing strategy for each partition based on business requirement information, the data lake system generates a strategy key-value pair for one or more partitions based on the business requirement information. The strategy key-value pair is used to indicate the number of buckets for different partitions. The strategy key-value pair includes a key and a value. The key is used to specify the partition, and the value is used to specify the number of buckets corresponding to the partition. The types of strategy key-value pairs include one or more of: direct matching, regular matching, and range matching. Among them, direct matching refers to directly specifying the number of buckets in a partition, regular matching specifies the number of buckets in a partition that meets the conditions through a string, and range matching specifies the number of buckets in a partition within a range. The data lake system generates a pre-bucketing strategy for one or more partitions based on the strategy key-value pair.
[0011] In the embodiment of the present application, the data lake system can automatically determine the policy key-value pairs according to the business demand information, and generate various types of pre-bucketing strategies based on the policy key-value pairs, thereby improving the configuration efficiency of the number of buckets in different partitions and further improving the data query and writing efficiency of the data lake.
[0012] In one possible implementation, when the data lake system acquires business demand information, it obtains the historical data volume written to one or more historical partitions and predicts business demand information based on the historical data volume written to one or more historical partitions. The business demand information includes the predicted volume of data to be written to a specified partition. When the data lake system determines a pre-bucketing strategy for one or more partitions based on the business demand information, it determines the number of buckets for the specified partition based on the predicted volume of data to be written to the specified partition and the bucket size of the specified partition.
[0013] In the embodiment of the present application, the data lake system can predict business demand information based on the historical write data volume of different partitions, thereby improving the accuracy of the business demand information and further improving the accuracy of the number of buckets in different partitions.
[0014] In one possible implementation, the data lake system calculates the predicted data volume to be written to one or more partitions based on the historical written data volume of one or more historical partitions and weight coefficients corresponding to different historical partitions.
[0015] In the embodiment of the present application, when the data lake system determines the business demand information based on the historical written data volume, it can calculate the predicted data volume of one or more partitions according to the weight coefficients corresponding to different historical partitions, thereby improving the accuracy of the business demand information and further improving the accuracy of the number of buckets in different partitions.
[0016] In one possible implementation, the data lake system calculates the predicted amount of data to be written for one or more partitions based on the historical written data volume of one or more historical partitions and the number of historical partitions, that is, the predicted amount of data to be written is determined based on the average historical written data volume of the historical partitions.
[0017] In the embodiment of the present application, when the data lake system determines the business demand information based on the historical written data volume, it can also determine the predicted data volume of the data to be written according to the average historical written data volume of the historical partition, thereby improving the prediction efficiency of the business demand information and further improving the configuration efficiency of the number of buckets in different partitions.
[0018] In one possible implementation, when a data lake system receives data to be written and determines that no pre-bucketing strategy exists for one or more partitions to which the data to be written is to be written, the data lake system automatically determines the number of buckets for the one or more partitions based on partition statistics, where the partition statistics include the amount of data written for the data to be written in the one or more partitions.
[0019] In an embodiment of the present application, when the data lake system is unable to obtain business demand information to generate a pre-bucketing strategy, the data lake system can determine partition statistics based on the data to be written, and determine the number of buckets for each partition based on the partition statistics, thereby improving the bucketing flexibility and efficiency of the data lake, and further improving the data query and write efficiency of the data lake.
[0020] In one possible implementation, in the process of automatically determining the number of buckets for a partition based on partition statistics, the data lake system aggregates the data to be written to determine the number of data items in each partition, samples the data to be written to determine the average size of each data item, determines partition statistics based on the number of data items and the average size of each data item, and calculates the number of buckets for each partition in one or more partitions based on the partition statistics and the bucket size of each partition.
[0021] In the embodiment of the present application, the data lake system can sample the data to be written to determine the average data size, partition and aggregate the data to be written to determine the number of data items in each partition, and calculate partition statistics based on the average data size and the number of data items in each partition, thereby improving the accuracy of the partition statistics and further improving the configuration accuracy of the number of buckets.
[0022] In one possible implementation, when creating a data table in the data lake based on a pre-bucketing strategy, the data lake system writes the pre-bucketing strategy into the metadata area of the data table. When the data lake system writes data to the data lake, it loads the pre-bucketing strategy from the metadata area of the data table and automatically sets the number of buckets for the partition based on the pre-bucketing strategy.
[0023] In the embodiment of the present application, the data lake system can directly write the pre-bucketing strategy in the metadata area of the data table, thereby improving the feasibility of configuring the pre-bucketing strategy.
[0024] In one possible implementation, when creating a data table in the data lake based on a pre-bucketing strategy, the data lake system generates a configuration file based on the pre-bucketing strategy. The configuration file includes the pre-bucketing strategy for one or more partitions, and the load path of the configuration file is written into the metadata area of the data table. When the data lake system writes data to the data lake, it obtains the load path of the configuration file from the metadata of the data table, loads the configuration file for the pre-bucketing strategy, and automatically sets the number of buckets for the partition based on the pre-bucketing strategy in the configuration file.
[0025] In the embodiment of the present application, the data lake system can also write the loading path of the configuration file corresponding to the pre-bucketing strategy in the metadata area of the data table, thereby improving the feasibility of configuring the pre-bucketing strategy.
[0026] In one possible implementation, when a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets for the partition is determined based on the intersection of the multiple pre-bucketing strategies.
[0027] In an embodiment of the present application, when a partition of the data lake is configured with multiple pre-bucketing strategies, the data lake system can determine the number of buckets for the partition based on the intersection of the multiple pre-bucketing strategies, thereby improving the configuration accuracy of the number of buckets.
[0028] In the second aspect, an embodiment of the present application provides a data lake bucketing method, which can be executed by a data lake system, or by a component of the data lake system, such as a processor, chip or chip system of the data lake system, or by a logic module or software that can implement all or part of the data lake system functions. The method provided in the first aspect includes: the data lake system receives one or more policy key-value pairs input by a user, the policy key-value pairs are used to indicate the number of buckets for different partitions, and the types of policy key-value pairs include one or more of: direct matching, regular matching and range matching. The data lake system generates a pre-bucketing strategy for one or more partitions in the data lake based on the one or more policy key-value pairs, and the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake. The data lake system creates a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage data in the data lake.
[0029] In the embodiment of the present application, users can also directly configure the pre-bucketing strategy through the policy key-value pair. After the data lake system receives the policy key-value pair input by the user, it generates a pre-bucketing strategy based on the policy key-value pair, thereby improving the configuration efficiency of the number of buckets in different partitions and further improving the data query and writing efficiency of the data lake.
[0030] In one possible implementation, when the data lake system creates a data table in the data lake based on the pre-bucketing strategy, the pre-bucketing strategy is written into the metadata area of the data table.
[0031] In one possible implementation, when creating a data table in the data lake based on a pre-bucketing strategy, the data lake system generates a configuration file based on the pre-bucketing strategy. The configuration file includes pre-bucketing strategies for one or more partitions, and the loading path of the configuration file is written into the metadata area of the data table.
[0032] In one possible implementation, when a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets for the partition is determined based on the intersection of the multiple pre-bucketing strategies.
[0033] In a third aspect, an embodiment of the present application provides a data lake bucketing device, which includes an acquisition unit, a processing unit, and a transceiver unit. The acquisition unit is used to acquire business demand information, and the business demand information is used to indicate the predicted amount of data to be written in one or more partitions in the data lake, and the data to be written is used to write into the data buckets of one or more partitions. The processing unit is used to determine the pre-bucketing strategy for one or more partitions based on the business demand information, and the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake. The processing unit is also used to create a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage the data in the data lake.
[0034] In one possible implementation, the acquisition unit is specifically configured to acquire the historical written data volume of one or more historical partitions, and predict business demand information based on the historical written data volume of the one or more historical partitions, where the business demand information includes a predicted data volume of data to be written to a specified partition. The processing unit is specifically configured to determine the number of buckets for the specified partition based on the predicted data volume of data to be written to the specified partition and the bucket size of the specified partition.
[0035] In one possible embodiment, the transceiver unit is configured to receive data to be written. The processing unit is further configured to automatically determine the number of buckets for the one or more partitions based on partition statistics when determining that no pre-bucketing strategy exists for one or more partitions to which the data to be written is to be written, the partition statistics including the amount of data written in the one or more partitions.
[0036] In one possible implementation, the processing unit is specifically used to perform aggregation operations on the data to be written, determine the number of data items in each partition, sample the data to be written, determine the average size of each data item, determine partition statistical information based on the number of data items and the average size of each data item, and calculate the number of buckets for each partition in one or more partitions based on the partition statistical information and the bucket size of each partition.
[0037] In a possible implementation, the processing unit is specifically configured to write a pre-bucketing strategy into a metadata area of the data table.
[0038] In a possible implementation, the processing unit is further configured to generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy for one or more partitions, and write a loading path of the configuration file into the metadata area of the data table.
[0039] In one possible implementation, the processing unit is further configured to, when a partition of the data lake is configured with multiple pre-bucketing strategies, determine the number of buckets for the partition based on the intersection of the multiple pre-bucketing strategies.
[0040] In a fourth aspect, an embodiment of the present application provides a data lake bucketing device, which includes a transceiver unit and a processing unit, wherein the transceiver unit is used to receive one or more policy key-value pairs input by a user, the policy key-value pairs are used to indicate the number of buckets for different partitions, and the types of policy key-value pairs include one or more of: direct matching, regular matching, and range matching. The processing unit is used to generate a pre-bucketing strategy for one or more partitions in the data lake based on one or more policy key-value pairs, and the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake. The processing unit is also used to create a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage data in the data lake.
[0041] In a possible implementation, the processing unit is specifically configured to write a pre-bucketing strategy into a metadata area of the data table.
[0042] In a possible implementation, the processing unit is specifically configured to generate a configuration file based on the pre-bucketing strategy, where the configuration file includes pre-bucketing strategies for one or more partitions, and write a loading path of the configuration file into a metadata area of the data table.
[0043] In one possible implementation, the processing unit is further configured to, when a partition of the data lake is configured with multiple pre-bucketing strategies, determine the number of buckets for the partition based on the intersection of the multiple pre-bucketing strategies.
[0044] In a fifth aspect, an embodiment of the present application provides a computing device, comprising a processor coupled to a memory, the processor being used to store instructions. When the instructions are executed by the processor, the computing device executes the method described in the first aspect or any possible implementation of the first aspect, or the computing device executes the method described in the second aspect or any possible implementation of the second aspect.
[0045] In a sixth aspect, an embodiment of the present application provides a computing device cluster, the computing device cluster including one or more computing devices, the computing device including a processor, the processor coupled to a memory, the processor for storing instructions, and when the instructions are executed by the processor, the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect, or the computing device cluster executes the method described in the second aspect or any possible implementation of the second aspect.
[0046] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed, the computer executes the method described in the first aspect or any possible implementation of the first aspect, or the computer executes the method described in the second aspect or any possible implementation of the second aspect.
[0047] In an eighth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed, the computer implements the method described in the first aspect or any possible implementation of the first aspect, or the computer implements the method described in the second aspect or any possible implementation of the second aspect.
[0048] It can be understood that the beneficial effects that can be achieved by any of the data lake bucketing devices, computing devices, computing device clusters, computer-readable media or computer program products provided above can refer to the beneficial effects in the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG1 is a schematic diagram of the system architecture of a data lake system provided in an embodiment of the present application;
[0050] FIG2 is a flow chart of a data lake bucketing method provided in an embodiment of the present application;
[0051] FIG3 is a schematic diagram of determining business demand information based on historical written data volume according to an embodiment of the present application;
[0052] FIG4 is a flowchart of another data lake bucketing method provided in an embodiment of the present application;
[0053] FIG5 is a flow chart of another data lake bucketing method provided in an embodiment of the present application;
[0054] FIG6 is a flow chart of another data lake bucketing method provided in an embodiment of the present application;
[0055] FIG7 is a flowchart of another data lake bucketing method provided in an embodiment of the present application;
[0056] FIG8 is a schematic diagram of the structure of a data lake bucketing device provided in an embodiment of the present application;
[0057] FIG9 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0058] FIG10 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0059] FIG11 is a schematic diagram of the structure of another computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The embodiments of the present application provide a data lake bucketing method and apparatus for improving the bucketing flexibility of the data lake and the efficiency of data query and writing.
[0061] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0062] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0063] First, some terms involved in the embodiments of the present application are introduced to facilitate those skilled in the art to understand the technical solutions.
[0064] A data lake is a centralized data repository for storing large amounts of structured or unstructured data. Data lakes typically use distributed storage and processing technologies, such as Hadoop and Spark, to support data storage, management, and analysis.
[0065] Partitions are partitions in a data lake that are divided according to certain rules. For example, partition management is performed by time, location, business department, etc. to improve query performance and data management flexibility. A partition can include multiple or multiple buckets.
[0066] Bucketing is a data storage and management technique that distributes data across buckets. These buckets are typically divided according to specific rules, such as hash bucketing based on a column or bucketing based on a range. Bucketing data allows for a more even distribution of data across physical storage, mitigating performance issues caused by excessive data in a single bucket.
[0067] Static bucketing means that data is written using a fixed number of buckets during the bucketing process. Dynamic bucketing means that the number of buckets used during the data bucketing process can change dynamically rather than being completely fixed.
[0068] The number of buckets refers to the number of buckets to which data is written during bucket writing.
[0069] Bucket size refers to the size of the data stored in each bucket.
[0070] In order to make the technical solution of the present application clearer and easier to understand, the system architecture of the present application is introduced below with reference to the accompanying drawings.
[0071] Please refer to Figure 1, which is a schematic diagram of the system architecture of a data lake system provided in an embodiment of the present application. In the example shown in Figure 1, data lake system 10 includes a data lake management subsystem 101, a computing engine subsystem 102, and a centralized storage subsystem 103. Centralized storage subsystem 103 includes one or more partitions, each of which contains multiple data buckets. The following describes in detail the specific functions of each subsystem in data lake system 10.
[0072] The data lake management subsystem 101 implements the data management functions of the data lake system 10. These functions include basic management functions and extended management functions. Basic management functions include data access, access control, task management, and metadata management. Basic management functions also include data governance, quality management, and asset catalog management.
[0073] Among the basic management functions, data access, for example, the data lake management subsystem 101 can connect to various external heterogeneous data sources to extract and migrate data. For example, the data lake management subsystem 101 can set different data access permissions for different users or user groups. For example, the data lake management subsystem 101 can manage and orchestrate tasks in the data lake, enabling automated execution and scheduling. For example, the data lake management subsystem 101 can manage and maintain metadata information in the data lake, including data structure, attributes, and relationships.
[0074] Among the extended management functions, data governance, such as the data lake management subsystem 101, can cleanse, integrate, and standardize data, and define data usage rules. Quality management, such as the data lake management subsystem 101, can assess and manage the quality of data in the data lake, providing data quality reporting and monitoring capabilities to help users understand data quality and identify areas for improvement. Asset catalog management, such as the data lake management subsystem 101, can establish a data asset catalog, categorize and label data, and help users quickly find the data resources they need.
[0075] The computing engine subsystem 102 is used to process and analyze data in the data lake system 10. The computing engine subsystem 102 can provide a variety of computing functions, including batch processing, stream computing, interactive query, and machine learning. Among them, the batch processing function, for example, the computing engine subsystem 102 can perform batch processing and analysis on large amounts of data. Stream computing, for example, the computing engine subsystem 102 can process real-time data streams and perform real-time analysis and processing on the data. Interactive query, for example, the computing engine subsystem 102 can provide interactive query functions, allowing users to retrieve and analyze data through query statements. Machine learning functions, for example, the computing engine subsystem 102 can support the training and prediction tasks of machine learning algorithms, providing data analysis and prediction capabilities.
[0076] The centralized storage subsystem 103 provides data storage within the database system 10. The centralized storage subsystem 103 includes one or more devices, each of which can be logically divided into different partitions. Each partition can store a group of data with similar characteristics or business implications. For example, the centralized storage subsystem 103 can be partitioned based on business attributes such as date, geographic location, and product category.
[0077] The centralized storage subsystem 103 can also set buckets within a partition. That is, the centralized storage subsystem 103 can partition the partitions at a finer granularity according to specified rules, with each bucket of the partition containing a portion of the partition's data. Specified rules include bucketing based on hash values, range values, or sample values. For example, the centralized storage subsystem 103 can hash the data in a data table based on the value of a column and then place data with the same hash value into the same bucket file.
[0078] Based on the data lake system 10 shown in Figure 1, the present application also provides a data lake bucketing method. The data lake bucketing method provided by the present application embodiment is described below in conjunction with an embodiment.
[0079] Please refer to Figure 2, which is a flow chart of a data lake bucketing method provided in an embodiment of the present application. In the example shown in Figure 2, the method includes the following steps:
[0080] 201. Obtain business demand information, where the business demand information is used to indicate a predicted amount of data to be written into one or more partitions in the data lake, where the data to be written is used to write into data buckets of the one or more partitions.
[0081] The data lake system 10 obtains business requirement information, where the business requirement information indicates a predicted amount of data to be written to one or more partitions in the data lake, and the data to be written is used to write data buckets in the one or more partitions. Specifically, the data lake system 10 may receive business requirement information input by a user, or the data lake system 10 may determine the business requirement information based on historical amounts of written data, without limitation.
[0082] In one possible implementation, when the data lake system 10 determines the business demand information based on the historical write data volume, the data lake system 10 first obtains the historical write data volume of one or more historical partitions, and predicts the business demand information based on the historical write data volume of one or more historical partitions. The business demand information includes the predicted data volume of the data to be written in one or more specified partitions.
[0083] Specifically, the data lake system 10 calculates the predicted data volume of the data to be written to the specified partition based on the historical write data volume of one or more historical partitions and the weight coefficients corresponding to different historical partitions, or the data lake system 10 calculates the predicted data volume of the data to be written to the specified partition based on the historical write data volume of one or more historical partitions and the number of historical partitions, that is, the predicted data volume of the data to be written is determined based on the average historical write data volume of the historical partitions.
[0084] Please refer to Figure 3, which is a schematic diagram of a data lake system determining business demand information according to an embodiment of the present application. In the example shown in Figure 3, data lake system 10 obtains the historical write data volume of one or more historical partitions. For example, data lake system 10 obtains the write data volume of historical partition 1, historical partition 2, historical partition 3, ..., historical partition n, respectively, as S1, S2, S3, ..., Sn.
[0085] In the example shown in FIG3 , after obtaining the historical written data volume of one or more historical partitions, the data lake system 10 calculates the predicted data volume to be written to the current partition based on the historical written data volume of the one or more historical partitions and the weight coefficients corresponding to the different historical partitions. For example, if the weight coefficients corresponding to historical partitions 1, 2, 3, ..., and n are w1, w2, w3, ..., wn, where w1, w2, w3, ..., wn satisfy the following formula:
[0086] In the example shown in FIG3 , the data lake system 10 can calculate the predicted amount of data to be written for all specified partitions based on the historical written data volume of historical partitions 1 to n and the weight coefficients w1 to wn corresponding to different historical partitions. The predicted amount of data to be written s satisfies the following formula:
[0087] In the example shown in FIG3 , when the corresponding weight coefficients of different historical partitions are the same, the predicted write data volume of the data to be written in each partition is s / n.
[0088] 202. Determine a pre-bucketing strategy for one or more partitions based on business demand information. The pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake.
[0089] The data lake system 10 determines a pre-bucketing strategy for one or more partitions based on business requirements. The pre-bucketing strategy indicates the number of buckets for each partition in the data lake. Specifically, the data lake system 10 calculates the number of buckets for the partitions based on the predicted amount of data to be written to the partitions and the bucket size of the partitions, where the bucket size is a preset value.
[0090] Please continue to refer to Figure 3. In the example shown in Figure 3, the data lake system 10 can calculate the predicted data volume sn of the data to be written to the specified partition based on the historical written data volume of historical partitions 1 to historical partitions n and the weight coefficients w1 to wn corresponding to different historical partitions. If the bucket size of the partition is k, the data lake system 10 can calculate the pre-bucketing strategy for the partition based on the predicted data volume sn of the data to be written to the partition and the bucket size k. The pre-bucketing strategy means that the number of buckets for the partition is sn / k.
[0091] It can be understood that when the business demand information is the predicted data volume to be written to multiple partitions, the data lake system 10 can determine the corresponding pre-bucketing strategy for each of the multiple partitions, that is, the data lake system 10 can determine the number of buckets corresponding to each partition in the multiple partitions.
[0092] In one possible implementation, when the data lake system 10 determines the pre-bucketing strategy for each partition based on business demand information, the data lake system 10 generates a strategy key-value pair for one or more partitions based on the business demand information. The strategy key-value pair is used to indicate the number of buckets for different partitions. The strategy key-value pair includes a key and a value. The key is used to specify the partition, and the value is used to specify the number of buckets corresponding to the partition. The type of the strategy key-value pair includes one or more of the following: direct matching, regular matching, and range matching. The data lake system 10 generates a pre-bucketing strategy for one or more partitions based on the strategy key-value pair. The following specifically introduces several types of the above-mentioned strategy key-value pairs:
[0093] Direct matching, also known as exact matching, refers to the data lake system 10 directly specifying the number of buckets in a partition. For example, "p=2022-10-11=4" is a configuration statement corresponding to the direct matching type, indicating that the number of buckets in the partition corresponding to "2022-10-11" is 4. Another example is "p=2022-11-11=10," which is also a configuration statement corresponding to the direct matching type and indicates that the number of buckets in the partition corresponding to "2022-11-11" is 11.
[0094] Regular matching means that the Data Lake System 10 specifies the number of buckets in a partition that meets the conditions using a string. For example, "p=****-11-11=9" is a configuration statement corresponding to the regular matching type, indicating that the number of buckets in the partition corresponding to "11-11" every year is 9.
[0095] Range matching refers to the data lake system 10 specifying the number of buckets in a partition within a range. For example, "p>2023-11-12=5" is a configuration statement corresponding to the range matching type, indicating that the number of buckets in the partition corresponding to each day after "2023-11-12" is 5.
[0096] In one possible implementation, if a partition of the data lake is configured with multiple pre-bucketing strategies, the data lake system 10 determines the number of buckets for the partition based on the intersection of the multiple pre-bucketing strategies. For example, if the partition "2023-12-12" corresponds to two pre-bucketing strategies, including "p>2023-11-12=5" and "p=****-12-12=8", the data lake system 10 determines the number of buckets for the partition "2023-12-12" based on the intersection of "p>2023-11-12=5" and "p=****-12-12=8", resulting in a total of 8 buckets for the partition "2023-12-12".
[0097] 203. Create data tables in the data lake based on the pre-bucketing strategy. The data tables are used to manage the data in the data lake.
[0098] After determining the pre-bucketing strategy, the data lake system 10 creates and saves a data table in the data lake based on the pre-bucketing strategy. The data table is used to manage the data in the data lake. When the data lake system 10 receives write data for a partition, it reads the pre-bucketing strategy corresponding to the partition from the data table and writes the data into buckets according to the number of buckets specified by the pre-bucketing strategy.
[0099] Please refer to Figure 4, which is a schematic diagram of a data bucketing method for writing data to a data lake according to an embodiment of the present application. In the example shown in Figure 4, before writing data to one or more partitions received by the data lake system 10, the user needs to configure a pre-bucketing strategy for one or more partitions in the data lake system 10 and create a data table.
[0100] In the example shown in FIG4 , when the data lake system 10 receives write data for a partition, if the partition has been configured with a pre-bucketing strategy, the data lake system 10 loads the pre-bucketing strategy corresponding to the partition from the metadata area and writes the data into buckets according to the number of buckets corresponding to the pre-bucketing strategy. For example, the data lake system 10 receives write data for partition B. Since partition B has been configured with a pre-bucketing strategy, the data lake system 10 loads the pre-bucketing strategy corresponding to partition B from the metadata area and writes the data into buckets according to the number of buckets corresponding to the pre-bucketing strategy. The number of buckets corresponding to the pre-bucketing strategy for partition B is 2. Therefore, the data lake system 10 writes the data into buckets 1 and 2 in partition B respectively.
[0101] Please refer to Figure 5, which is a schematic diagram of a process for writing real-time data into a data lake according to an embodiment of the present application. In steps a to c and steps e to f of the example shown in Figure 5, the data lake system 10 receives write data from one or more partitions. The data lake system 10 then loads and determines whether a pre-bucketing strategy exists for the partition corresponding to the write data. If a pre-bucketing strategy exists for the partition corresponding to the write data, the data lake system 10 saves the pre-bucketing strategy corresponding to the partition to the metadata area of the partition, buckets the partition according to the number of buckets in the pre-bucketing strategy, and performs bucket writing.
[0102] In one possible implementation, when the data lake system 10 creates a data table in the data lake based on the pre-bucketing strategy, the data lake system 10 can write the pre-bucketing strategy into the metadata area of the data table. That is, the data lake system 10 can directly write the strategy key-value pair corresponding to the pre-bucketing strategy into the metadata area of the data table.
[0103] For example, when creating a data table, the data lake system 10 directly writes the pre-bucketing strategy into the "tblproperties" property of the table creation code and saves it to the metadata area of the data table. When writing data to the data lake system 10 later, the pre-bucketing strategy is loaded from the metadata of the data table, and the number of partition buckets is automatically set according to the pre-bucketing strategy. The sample code for directly writing the pre-bucketing strategy into the table creation code is as follows:
[0104] CREATE TABLE LAKETABLE CLUSTED BY(uuid)INTO N BUCKETS
[0105] TBLPROPERTIES(
[0106] 'p=2023-11-12'='4',
[0107] 'p=2022-**-**'='9')
[0108] In one possible implementation, the data lake system 10 generates a configuration file based on a pre-bucketing strategy. The configuration file includes pre-bucketing strategies for one or more partitions. When creating a data table in the data lake based on the pre-bucketing strategy, the data lake system 10 writes the load path of the configuration file into the metadata area of the data table.
[0109] For example, when creating a data table, the data lake system 10 can also write the path to load the pre-bucketing strategy configuration file into the "tblproperties" property of the table creation code. When data is subsequently written to the data lake system 10, the pre-bucketing strategy configuration file is loaded from the metadata area of the data table, and the number of partition buckets is automatically set based on the configuration file. The following is an example of writing the path to load the pre-bucketing strategy configuration file into the table creation code:
[0110] CREATE TABLE LAKETABLE CLUSTED BY(uuid)INTO N BUCKETS
[0111] TBLPROPERTIES(
[0112] bucket.strategy.location=' / tmp / bucketNumStategy')
[0113] In one possible implementation, after the data lake system 10 receives the data to be written, when it is determined that there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, the data lake system 10 estimates the number of partition buckets based on the data to be written, that is, the data lake system 10 automatically determines the number of buckets for one or more partitions based on partition statistical information, and the partition statistical information includes the amount of data written in one or more partitions for the data to be written.
[0114] Continuing with Figure 4, in the example shown in Figure 4, when the data lake system 10 receives write data for a partition, if the partition does not have a pre-bucketing policy configured, the data lake system 10 automatically determines the number of buckets for one or more partitions by inferring the number of partition buckets based on the data to be written. For example, the data lake system 10 receives write data for partition A. Since partition A does not have a pre-bucketing policy configured in advance, the data lake system 10 infers the number of buckets for partition A to be 2 based on the data to be written corresponding to partition A. The data lake system 10 buckets partition A based on the bucket number 2 to obtain buckets 1 and 2, and performs bucket writes in buckets 1 and 2 of partition A.
[0115] Please continue to refer to Figure 5. In step d of the example shown in Figure 5, when the data lake system 10 determines whether there is a pre-bucketing strategy for the partition corresponding to the written data, if there is no pre-bucketing strategy for the partition corresponding to the data to be written, the data lake system 10 estimates the number of buckets based on the amount of data to be written, buckets the partition according to the estimated number of buckets, and performs bucket writing.
[0116] In one possible implementation, when the data lake system 10 automatically determines the number of buckets for a partition based on partition statistics, the data lake system 10 aggregates the data to be written to determine the number of data entries in each partition, and samples the data to be written to determine the average size of each data entry. Based on the number of data entries and the average size of each data entry, the data lake system 10 determines partition statistics, which represent the amount of data written to each partition. Based on the partition statistics and the bucket size of each partition, the data lake system 10 calculates the number of buckets for each of one or more partitions.
[0117] Please refer to Figure 6, which is a flowchart of estimating the number of buckets provided by an embodiment of the present application. In step a of the example shown in Figure 6, when there is no pre-bucketing strategy for the partition to which the data to be written is to be written, the data lake system 10 performs an aggregation operation on the data to be written and calculates the number of data items in each partition. For example, the data to be written may be data to be written to one or more partitions, and the data lake system 10 calculates the number of data items in each partition, for example, the number of data items to be written in the first partition is r1, the number of data items to be written in the second partition is r2, ..., and the number of data items to be written in the nth partition is rn.
[0118] In steps b through c of the example shown in FIG6 , the data lake system 10 may also sample the data to be written and determine the average size of each data entry. For example, the data lake system 10 samples the data to be written and determines the average size of each data entry as x. After calculating the average data size and the number of data entries in each partition, the data lake system 10 calculates partition statistics based on the number of data entries and the average size of each data entry. The partition statistics are the amount of data written to each partition.
[0119] For example, the data lake system 10 performs the calculation x*rn and obtains that the amount of data to be written to the first partition is s1, the amount of data to be written to the second partition is s2, ..., and the amount of data to be written to the nth partition is sn.
[0120] In step d of the example shown in FIG6 , after calculating the partition statistics, the data lake system 10 calculates the number of buckets for each of one or more partitions based on the partition statistics and the bucket size of each partition. For example, if the bucket size of each bucket is 1 GB, the data lake system 10 calculates the number of buckets for each partition as sn / 1 GB.
[0121] In steps e to f of the example shown in FIG6 , after the data lake system 10 calculates the number of buckets for each of one or more partitions, the data lake system 10 saves the calculated number of buckets to the metadata file in the corresponding partition directory and writes the buckets according to the number of buckets in the partition directory. The directory structure of the metadata file corresponding to the partition directory is as follows:
[0122] …. / table / 2022-01-05 / bucket.info
[0123] …. / table / 2022-01-06 / bucket.info
[0124] It should be noted that in the embodiment shown in FIG2 above, the data lake system 10 can automatically generate a policy key-value pair for one or more partitions based on business demand information. In the embodiment of the present application, the data lake system 10 can be directly configured with a policy key-value pair by a user, and a pre-bucketing strategy for one or more partitions in the data lake is generated based on the user-configured policy key-value pair. This will be described below in conjunction with the embodiment shown in FIG7 :
[0125] Please refer to Figure 7, which is a flowchart of another data lake bucketing method provided by an embodiment of the present application. In the example shown in Figure 2, the method includes the following steps:
[0126] 701. Receive one or more policy key-value pairs input by the user, where the policy key-value pairs are used to indicate the number of buckets in different partitions. The types of policy key-value pairs include one or more of: direct matching, regular matching, and range matching.
[0127] The data lake system 10 receives one or more policy key-value pairs input by the user. The policy key-value pairs are used to indicate the number of buckets for different partitions. In other words, the user can directly configure a pre-bucketing strategy for one or more partitions using the policy key-value pairs. When determining the policy key-value pairs, the user can base their decision on business requirements, including the predicted amount of data to be written to one or more partitions in the data lake.
[0128] In one possible implementation, the user may write a pre-bucketing strategy in the metadata area of the data table, that is, the user may directly write a strategy key-value pair corresponding to the pre-bucketing strategy in the metadata area of the data table.
[0129] In one possible implementation, the data lake system 10 may receive a configuration file of a pre-bucketing strategy sent by a user, and write a loading path of the configuration file into a metadata area of a data table.
[0130] 702. Generate a pre-bucketing strategy for one or more partitions in the data lake based on one or more strategy key-value pairs, where the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake.
[0131] 703. Create data tables in the data lake based on the pre-bucketing strategy. The data tables are used to manage data in the data lake.
[0132] The method performed by the data lake system 10 in steps 702 and 703 of the embodiment of the present application is similar to the method performed by the data lake system 10 in steps 202 and 203 of the embodiment shown in Figure 2 above, and will not be repeated here.
[0133] It can be understood that the difference between the embodiment shown in Figure 7 and the embodiment shown in Figure 2 is that the data lake system 10 in the embodiment shown in Figure 2 can obtain business demand information and generate a pre-bucketing strategy based on the business demand information, while in the embodiment shown in Figure 7, the user directly configures the policy key-value pair to obtain the pre-bucketing strategy. The embodiment shown in Figure 7 can also include any possible implementation method of the embodiment shown in Figure 2, and the details will not be repeated.
[0134] It can be seen from the above embodiments that the data lake system in the embodiments of the present application can configure different pre-bucketing strategies for different partitions based on business demand information. The data lake system flexibly sets the number of buckets for different partitions based on the predicted data volume corresponding to the business needs, thereby improving the configuration accuracy of the number of buckets for different partitions and the flexibility of bucketing, and further improving the data query and writing efficiency of the data lake.
[0135] Based on the above method embodiment, the embodiment of the present application also provides a data lake bucketing device. The data lake bucketing device provided by the embodiment of the present application is described in detail below.
[0136] Please refer to Figure 8, which is a schematic diagram of the structure of a data lake bucketing device provided in an embodiment of the present application. In the example shown in Figure 8, the data lake bucketing device 800 is used to implement the various steps performed by the data lake system 10 in the above embodiments. The data lake bucketing device 800 includes an acquisition unit 801, a processing unit 802, and a transceiver unit 803.
[0137] In one example of an embodiment of the present application, the data lake bucketing device 800 is used to implement the various steps performed by the data lake system 10 in the embodiment shown in Figure 3 above.
[0138] The acquisition unit 801 is used to acquire business demand information, where the business demand information indicates the predicted amount of data to be written to one or more partitions in the data lake, where the data to be written is written to the data buckets of the one or more partitions. The processing unit 802 is used to determine a pre-bucketing strategy for the one or more partitions based on the business demand information, where the pre-bucketing strategy indicates the number of buckets for different partitions in the data lake. The processing unit 802 is also used to create a data table in the data lake based on the pre-bucketing strategy, where the data table is used to manage the data in the data lake.
[0139] In one possible implementation, the acquisition unit 801 is specifically configured to acquire the historical written data volume of one or more historical partitions, and predict business demand information based on the historical written data volume of the one or more historical partitions, where the business demand information includes the predicted data volume of data to be written to a specified partition. The processing unit 802 is specifically configured to determine the number of buckets for the specified partition based on the predicted data volume of data to be written to the specified partition and the bucket size of the specified partition.
[0140] In one possible implementation, the transceiver unit 803 is configured to receive data to be written. The processing unit 802 is further configured to automatically determine the number of buckets for the one or more partitions based on partition statistics when determining that no pre-bucketing strategy exists for one or more partitions to which the data to be written is to be written, the partition statistics including the amount of data written in the one or more partitions.
[0141] In one possible implementation, the processing unit 802 is specifically used to perform aggregation operations on the data to be written, determine the number of data items in each partition, sample the data to be written, determine the average size of each data item, determine partition statistical information based on the number of data items and the average size of each data item, and calculate the number of buckets for each partition in one or more partitions based on the partition statistical information and the bucket size of each partition.
[0142] In a possible implementation, the processing unit 802 is specifically configured to write a pre-bucketing strategy into a metadata area of the data table.
[0143] In a possible implementation, the processing unit 802 is further configured to generate a configuration file based on the pre-bucketing strategy, where the configuration file includes pre-bucketing strategies for one or more partitions, and write a loading path of the configuration file into the metadata area of the data table.
[0144] In a possible implementation, the processing unit 802 is further configured to determine the number of buckets for a partition based on the intersection of the multiple pre-bucketing strategies when a partition of the data lake is configured with multiple pre-bucketing strategies.
[0145] In another example of an embodiment of the present application, the data lake bucketing device 800 is used to implement the various steps performed by the data lake system 10 in the embodiment shown in Figure 7 above.
[0146] Among them, the transceiver unit 803 is used to receive one or more policy key-value pairs input by the user, and the policy key-value pairs are used to indicate the number of buckets for different partitions. The types of policy key-value pairs include one or more of: direct match, regular match, and range match. The processing unit 802 is used to generate a pre-bucketing strategy for one or more partitions in the data lake based on the one or more policy key-value pairs, and the pre-bucketing strategy is used to indicate the number of buckets for different partitions in the data lake. The processing unit 802 is also used to create a data table in the data lake based on the pre-bucketing strategy, and the data table is used to manage the data in the data lake.
[0147] In a possible implementation, the processing unit 802 is specifically configured to write a pre-bucketing strategy into a metadata area of the data table.
[0148] In a possible implementation, the processing unit 802 is specifically configured to generate a configuration file based on the pre-bucketing strategy, where the configuration file includes pre-bucketing strategies for one or more partitions, and write a loading path of the configuration file into the metadata area of the data table.
[0149] In a possible implementation, the processing unit 802 is further configured to determine the number of buckets for a partition based on the intersection of the multiple pre-bucketing strategies when a partition of the data lake is configured with multiple pre-bucketing strategies.
[0150] It can be understood that the acquisition unit 801, processing unit 802 and transceiver unit 803 in the data lake bucketing device 800 can be mapped as functional modules to the various modules in the data lake system 10 in Figure 1, thereby realizing the functions of each module in the data lake system 10.
[0151] It should be understood that the division of units in the above device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. Moreover, the units in the device can all be implemented in the form of software called through processing elements; or they can all be implemented in the form of hardware; or some units can be implemented in the form of software called through processing elements, and some units can be implemented in the form of hardware. For example, each unit can be a separately established processing element, or it can be integrated into a certain chip of the device. In addition, it can also be stored in the memory in the form of a program, called by a certain processing element of the device and perform the function of the unit. In addition, all or part of these units can be integrated together, or they can be implemented independently. The processing element described here can also be a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each unit above can be implemented by the integrated logic circuit of the hardware in the processor element or in the form of software called through the processing element.
[0152] It is worth noting that, for the sake of simplicity of description, the above method embodiments are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited to the order of the actions described. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required for this application.
[0153] Other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0154] Please refer to Figure 9, which is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in Figure 9, the computing device 900 includes: a processor 901, a memory 902, a communication interface 903, and a bus 904. The processor 901, memory 902, and communication interface 903 are coupled via a bus (not labeled in the figure). The memory 902 stores instructions. When the execution instructions in the memory 902 are executed, the computing device 900 performs the method performed by the computing device in the above method embodiment.
[0155] The computing device 900 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms. For example, when a unit in the apparatus can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0156] The processor 901 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0157] The memory 902 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0158] Memory 902 stores executable program code, and processor 901 executes the executable program code to implement the functions of the aforementioned units or modules, thereby implementing the aforementioned data lake bucketing method. In other words, memory 902 stores instructions for executing the aforementioned data lake bucketing method.
[0159] The communication interface 903 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or a communication network.
[0160] In addition to the data bus, bus 904 may also include a power bus, a control bus, and a status signal bus. The bus may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), or a Cache Coherent Interconnect for Accelerators (CCIX). Buses can be categorized as address buses, data buses, and control buses.
[0161] Please refer to Figure 10 , which is a schematic diagram of a computing device cluster provided in an embodiment of the present application. As shown in Figure 10 , the computing device cluster 1000 includes at least one computing device 900 .
[0162] As shown in Figure 10, the computing device cluster 1000 includes at least one computing device 900. The memory 902 in one or more computing devices 900 in the computing device cluster 1000 may store the same instructions for executing the above-mentioned data lake bucketing method.
[0163] In some possible implementations, the memory 902 of one or more computing devices 900 in the computing device cluster 1000 may also store partial instructions for executing the aforementioned data lake bucketing method. In other words, the combination of one or more computing devices 900 can jointly execute the instructions for executing the aforementioned data lake bucketing method.
[0164] It should be noted that the memory 902 in different computing devices 900 in the computing device cluster 1000 can store different instructions, each for executing part of the functions of the aforementioned data lake bucketing device. In other words, the instructions stored in the memory 902 in different computing devices 900 can implement the functions of one or more modules in the acquisition unit, processing unit, and transceiver unit.
[0165] In some possible implementations, one or more computing devices 900 in the computing device cluster 1000 may be connected via a network, which may be a wide area network or a local area network.
[0166] Please refer to Figure 11, which is a schematic diagram of computer devices in a computer cluster provided by an embodiment of the present application connected via a network. As shown in Figure 11, two computing devices 900A and 900B are connected via a network. Specifically, the connection to the network is through a communication interface in each computing device.
[0167] In one possible implementation, the memory of the computing device 900A stores instructions for executing the functions of the acquisition unit and the transceiver unit, while the memory of the computing device 900B stores instructions for executing the functions of the processing unit.
[0168] It should be understood that the functions of the computing device 900A shown in Figure 11 may also be completed by multiple computing devices. Similarly, the functions of the computing device 900B may also be completed by multiple computing devices.
[0169] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the method executed by the data lake system in the above method embodiment.
[0170] In another embodiment of the present application, a computer program product is provided, including computer-executable instructions stored in a computer-readable storage medium. When a processor of a device executes the computer-executable instructions, the device performs the method performed by the data lake system in the above method embodiment.
[0171] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0172] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0173] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0174] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
Claims
1. A bucketing method for a data lake, characterized in that: include: Acquire business demand information, where the business demand information is used to indicate a predicted amount of data to be written in one or more partitions in the data lake, where the data to be written is used to write into data buckets of the one or more partitions; Determine a pre-bucketing strategy for the one or more partitions according to the business requirement information, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; A data table is created in the data lake based on the pre-bucketing strategy, where the data table is used to manage data in the data lake.
2. The method according to claim 1, characterized in that The obtaining of business demand information includes: Get the historical written data volume of one or more historical partitions; Predicting the business demand information according to the historical written data volume of the one or more historical partitions, the business demand information including the predicted data volume of the data to be written to the specified partition; The determining of the pre-bucketing strategy of the one or more partitions according to the business requirement information includes: The number of buckets of the specified partition is determined according to the predicted data volume of the data to be written to the specified partition and the bucket size of the specified partition.
3. The method according to claim 1, characterized in that The method further comprises: receiving the data to be written; When it is determined that there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, the number of buckets of the one or more partitions is automatically determined based on partition statistical information, and the partition statistical information includes the amount of data written by the data to be written in the one or more partitions.
4. The method according to claim 3, characterized in that The automatically determining the number of buckets of the partition based on the partition statistics information includes: Aggregating the data to be written to determine the number of data entries in each partition; Sampling the data to be written to determine the average size of each piece of data; Determine the partition statistical information according to the number of data pieces and the average size of each data piece; The number of buckets of each of the one or more partitions is calculated based on the partition statistics and the bucket size of each partition.
5. The method according to any one of claims 1 to 4, characterized in that The creating a data table in the data lake based on the pre-bucketing strategy includes: The pre-bucketing strategy is written into the metadata area of the data table.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The creating a data table in the data lake based on the pre-bucketing strategy includes: The loading path of the configuration file is written into the metadata area of the data table.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.
8. A method for bucketing a data lake, characterized in that: include: Receive one or more policy key-value pairs input by a user, where the policy key-value pairs are used to indicate the number of buckets in different partitions, and the type of the policy key-value pairs includes one or more of: direct match, regular match, and range match; Generate a pre-bucketing strategy for one or more partitions in the data lake based on the one or more policy key-value pairs, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; A data table is created in the data lake based on the pre-bucketing strategy, where the data table is used to manage data in the data lake.
9. The method according to claim 8, characterized in that The creating a data table in the data lake based on the pre-bucketing strategy includes: The pre-bucketing strategy is written into the metadata area of the data table.
10. The method according to claim 8, characterized in that The method further comprises: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The creating a data table in the data lake based on the pre-bucketing strategy includes: The loading path of the configuration file is written into the metadata area of the data table.
11. The method according to any one of claims 8 to 10, characterized in that The method further comprises: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.
12. A bucketing device for a data lake, characterized in that: include: An acquisition unit, used to acquire business demand information, where the business demand information is used to indicate a predicted amount of data to be written in one or more partitions in the data lake, where the data to be written is used to be written into data buckets of the one or more partitions; A processing unit, configured to determine a pre-bucketing strategy for the one or more partitions according to the business requirement information, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; The processing unit is further used to create a data table in the data lake based on the pre-bucketing strategy, where the data table is used to manage the data in the data lake.
13. The device according to claim 12, characterized in that The acquisition unit is specifically used for: Get the historical written data volume of one or more historical partitions; Predicting the business demand information according to the historical written data volume of the one or more historical partitions, the business demand information including the predicted data volume of the data to be written to the specified partition; The processing unit is specifically used for: The number of buckets of the specified partition is determined according to the predicted data volume of the data to be written to the specified partition and the bucket size of the specified partition.
14. The device according to claim 12, characterized in that The device further comprises a transceiver unit, wherein the transceiver unit is used for: receiving the data to be written; The processing unit is also used to automatically determine the number of buckets of the one or more partitions based on partition statistical information when there is no pre-bucketing strategy for one or more partitions to which the data to be written is to be written, and the partition statistical information includes the amount of data written by the data to be written in the one or more partitions.
15. The device according to claim 14, characterized in that The processing unit is specifically used for: Aggregating the data to be written to determine the number of data entries in each partition; Sampling the data to be written to determine the average size of each piece of data; Determine the partition statistical information according to the number of data pieces and the average size of each data piece; The number of buckets of each of the one or more partitions is calculated based on the partition statistics and the bucket size of each partition.
16. The device according to any one of claims 12 to 15, characterized in that The processing unit is specifically used for: The pre-bucketing strategy is written into the metadata area of the data table.
17. The device according to any one of claims 12 to 15, characterized in that The processing unit is also used for: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The loading path of the configuration file is written into the metadata area of the data table.
18. The device according to any one of claims 12 to 17, characterized in that The processing unit is also used for: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.
19. A bucketing device for a data lake, characterized in that: include: A transceiver unit, configured to receive one or more policy key-value pairs input by a user, wherein the policy key-value pairs are used to indicate the number of buckets in different partitions, and the types of the policy key-value pairs include one or more of: direct matching, regular matching, and range matching; A processing unit, configured to generate a pre-bucketing strategy for one or more partitions in the data lake based on the one or more policy key-value pairs, wherein the pre-bucketing strategy is used to indicate the number of buckets of different partitions in the data lake; The processing unit is further used to create a data table in the data lake based on the pre-bucketing strategy, where the data table is used to manage the data in the data lake.
20. The device according to claim 19, characterized in that The processing unit is specifically used for: The pre-bucketing strategy is written into the metadata area of the data table.
21. The device according to claim 19, characterized in that The processing unit is specifically used for: Generate a configuration file based on the pre-bucketing strategy, the configuration file including the pre-bucketing strategy of the one or more partitions; The loading path of the configuration file is written into the metadata area of the data table.
22. The device according to any one of claims 19 to 21, characterized in that The processing unit is also used for: When a partition of the data lake is configured with multiple pre-bucketing strategies, the number of buckets of the partition is determined based on the intersection of the multiple pre-bucketing strategies.
23. A computing device, characterized in that The invention comprises a processor coupled to a memory, wherein the processor is used to store instructions, and when the instructions are executed by the processor, the computing device executes the method according to any one of claims 1 to 7, or the computing device executes the method according to any one of claims 8 to 11.
24. A computing device cluster, characterized in that: The method comprises at least one computing device, wherein the computing device comprises a processor, the processor is coupled to a memory, and the processor is used to store instructions. When the instructions are executed by the processor, the computing device cluster executes the method of any one of claims 1 to 7, or the computing device cluster executes the method of any one of claims 8 to 11.
25. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed, the computer is caused to execute the method according to any one of claims 1 to 7, or the computer is caused to execute the method according to any one of claims 8 to 11.
26. A computer program product, comprising instructions, characterized in that: When the instructions are executed, the computer implements the method according to any one of claims 1 to 7, or the computer implements the method according to any one of claims 8 to 11.
Citation Information
Patent Citations
Data storage method, device and equipment based on distributed system and storage medium
CN115481295A
Data lake data synchronization method and device
CN115840786A
Data processing method and device in Iceberg, storage medium and equipment
CN115952193A
Storage system with bucket contents rebalancer providing adaptive partitioning for database buckets
US10324911B1
Cited By
Front-end processing method, system and equipment for large file and storage medium
CN121056452A
Mass process data presentation method based on Websocket
CN122064754A