Data lake and warehouse management method and equipment
By dynamically adjusting the number of metadata management nodes, the data lake warehouse service has solved the problem of insufficient processing capabilities when processing large amounts of data and redundant performance when the amount of data is small, and the performance and stability of the system are improved.
Patent Information
- Application Number
- CN202510113026.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional data lake warehouse services lack processing capabilities when processing large amounts of data, resulting in crashes, and performance is redundant when the amount of data is small, resulting in waste of resources.
By monitoring the total data volume processed by the operation tasks of the computing cluster, dynamically expand and adjust the number of metadata management nodes in the metadata management cluster, ensuring sufficient load capacity when large-scale data enters the lake warehouse, and avoiding performance redundancy when the data volume is small.
Improve the performance of the data lake warehouse, ensure sufficient processing capabilities during large-scale data processing, avoid resource waste, and improve system stability.
Smart Images

Figure CN120179643A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and in particular, to a method and device for managing a data lake warehouse. Background Art
[0002] In the big data era, the data storage requirements of enterprises and organizations have reached the TB (terabyte) or even PB (petabyte) level. Facing large-scale and diverse data, traditional relational databases such as MySQL and Oracle have many limitations, so the data lake warehouse came into being. The data lake warehouse is a data management architecture that combines the advantages of a data warehouse and a data lake. It widely uses HDFS (Hadoop Distributed File System) technology to implement data storage and uses the form of data lake tables to manage the data files in HDFS.
[0003] Among them, HDFS is inclined to process batch processing tasks of large-scale data sets and is designed specifically for storing large files. Therefore, in the data lake warehouse architecture, some data lake warehouse services adapted to HDFS (such as lake table management components) also need to be configured with corresponding resources. Usually, the resource scale configured for these data lake warehouse services is specified according to experience or is configured after balancing the cost of resources and the amount of data processed by HDFS daily. This makes the processing capacity of the data lake warehouse services with a fixed scale insufficient to handle the large amount of data when a large amount of data enters the data lake warehouse, and even causes crashes. Moreover, the low concurrency rate of the data lake warehouse services slows down the speed of data entering the lake warehouse and increases the probability of data lake warehouse service crashes. However, when a small amount of data enters the lake warehouse, it will cause performance redundancy of the data lake warehouse services and result in resource waste.
[0004] In summary, these drawbacks have imposed great limitations on the performance of the data lake warehouse. Summary of the Invention
[0005] The embodiments of this application provide a method, device, computing device, device cluster, computer storage medium, computer program product, and data lake warehouse system for managing a data lake warehouse, which can improve the performance of the data lake warehouse.
[0006] In a first aspect, the embodiments of this application provide a method for managing a data lake warehouse. The method can be applied to a management node and may include: obtaining the total amount of data processed by an operation task of a computing cluster in the data lake warehouse, where the operation task is used to process target data, and the target data is data in the storage cluster of the data lake warehouse or data to be written into the storage cluster; and scaling the number of metadata management nodes in the metadata management cluster in the data lake warehouse according to the total amount of data, where the metadata management cluster is used to manage the metadata of the data in the storage cluster.
[0007] In this way, in the large-scale data lake warehouse storage scenario, this embodiment can scale the scale (number of nodes) of the metadata management cluster in the data lake warehouse according to the total amount of data processed by the operation tasks of the computing cluster on the storage cluster, ensuring that the metadata management cluster has sufficient load capacity when a large amount of data enters the lake warehouse, or avoiding performance redundancy of the data lake warehouse services provided by the metadata management cluster when the amount of data processed by the computing cluster is small, thus avoiding resource waste.
[0008] In some possible examples, scaling the number of metadata management nodes in the metadata management cluster in the data lake warehouse according to the total amount of data specifically includes: obtaining the first total amount of data processed by the operation tasks in the computing cluster during the current period; determining the change amount between the first total amount of data and the second total amount of data, where the second total amount of data is the total amount of data processed by the operation tasks in the computing cluster during the historical period; predicting the third total amount of data to be processed by the computing cluster in the next period according to the change amount; and determining the number of metadata management nodes required for the third total amount of data from the metadata management cluster according to the scheduling policy to achieve scaling adjustment. Here, the scheduling policy is used to describe the number of metadata management nodes required for different total amounts of data. The metadata management cluster includes multiple metadata management nodes, and each metadata management node is used to manage the metadata of the data in the storage cluster through the data lake table.
[0009] In this example, according to the total amount of data processed by the computing cluster in the current period and the total amount of data processed in the historical period, predict the number of iceberg nodes required in the next period, so as to adaptively adjust the number of metadata management nodes to meet the requirements of the computing cluster. Exemplarily, the total amount of data in the next period can be predicted by comparing the total amount of data of the computing cluster in the current period with the total amounts of data in the previous N (N≥2) historical periods. For example, compare the total amount of data at 8 o'clock today (i.e., the first total amount of data) with the total amounts of data at 6 o'clock and 7 o'clock today (i.e., two second total amounts of data) to obtain the change amount (it can be the comparison between the total amount of data at 6 o'clock and 7 o'clock, and the comparison between the total amount of data at 7 o'clock and 8 o'clock. Determine the corresponding change trend through these two comparison results), and thus predict the total amount of data at 9 o'clock today (i.e., the third total amount of data). Or the total amount of data in the next period can also be predicted according to the difference in the total amount of data between the same time periods within different cycles (such as within 24 hours). For example, compare the total amount of data at 8 o'clock today (i.e., the first total amount of data) with the total amount of data at 8 o'clock yesterday (i.e., the second total amount of data) to obtain a comparison result, and thus predict the total amount of data at 9 o'clock today (i.e., the third total amount of data), and this third total amount of data can be determined according to the total amount of data at 9 o'clock yesterday and this comparison result.
[0010] In some possible examples, after determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling policy, the method includes: comparing the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in the working state during the current period; if the number of metadata management nodes in the working state during the current period is less than the number of metadata management nodes required for the third total data volume, controlling some of the metadata management nodes in the non-working state in the metadata management cluster to start in the next period so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total data volume; and allocating the required resources to the started part of the metadata management nodes.
[0011] In this way, based on the prediction result, if the current metadata management nodes do not meet the processing requirements of the total data volume in the next period, the management nodes will start some metadata management nodes to expand the scale of the metadata management cluster so as to support the task processing in the next period and ensure sufficient load capacity when a large amount of data is input into the lake warehouse.
[0012] In some possible examples, after determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling policy, the method includes: comparing the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in the working state during the current period; if the number of metadata management nodes in the working state during the current period is greater than the number of metadata management nodes required for the third total data volume, determining some of the metadata management nodes in the working state during the current period to be shut down in the next period so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total data volume; and after the shutdown of the part of the metadata management nodes, recycling the resources occupied by the part of the metadata management nodes.
[0013] In this way, based on the prediction result, if the current metadata management nodes exceed the processing requirements of the total data volume in the next period, the management nodes will shut down some metadata management nodes to reduce the scale of the metadata management cluster so as to avoid performance redundancy and resource waste in the next period.
[0014] In some possible examples, after determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling policy, the method includes: comparing the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in the working state during the current period; if the number of metadata management nodes in the working state during the current period is equal to the number of metadata management nodes required for the third total data volume, determining the metadata management nodes in the working state during the current period to be used to support the tasks of the computing cluster in the next period.
[0015] In this way, based on the prediction result, if the current metadata management node can basically meet the processing requirements of the total data volume in the next time period, the management node can keep the scale of the current metadata management cluster unchanged.
[0016] In some possible examples, after determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling policy, the method includes: determining a device list according to the number of metadata management nodes required for the third total data volume, where the device list is used to describe the device information of all metadata management nodes that will be in a working state in the next time period; sending the device list to each computing node of the computing cluster, so that the computing node selects a target metadata management node from the device list for communication connection, and the target metadata management node is a metadata management node of the metadata management cluster.
[0017] In this way, according to the predicted number of metadata management nodes in the next time period, the information of these available metadata management nodes is updated to the device list and sent to the computing node, so that the computing node can complete processing requests such as addition, deletion, modification, and query of data by the service side according to the list.
[0018] In some possible examples, predicting the third total data volume processed by the computing cluster in the next time period according to the change amount includes: if the change amount indicates that the first total data volume is equal to the second total data volume, predicting that the third total data volume is equal to the first total data volume; if the change amount indicates that the first total data volume shows a decreasing trend relative to the second total data volume, predicting that the third total data volume is the difference between the first total data volume and the change amount; if the change amount indicates that the first total data volume shows an increasing trend relative to the second total data volume, predicting that the third total data volume is the sum of the first total data volume and the change amount.
[0019] In some possible examples, before determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling policy, the method further includes: monitoring the status information in the metadata management cluster, where the status information includes the number of metadata management nodes in a working state and / or a non-working state, the currently available resource amount, the used resource amount, and the health status information of each metadata management node. So as to realize the scaling adjustment of the scale of the metadata management cluster according to these status information.
[0020] Second aspect, embodiments of the present application provide a data lake warehouse management device, which may include: an acquisition module, configured to acquire the total amount of data processed by an operation task of a computing cluster in the data lake warehouse, where the operation task is used to process target data, and the target data is data in the storage cluster of the data lake warehouse or data to be written into the storage cluster; a processing module, configured to scale the number of metadata management nodes in the metadata management cluster in the data lake warehouse according to the total amount of data, where the metadata management cluster is used to manage the metadata of the data in the storage cluster.
[0021] In some possible examples, the acquisition module may also be configured to acquire the first total amount of data processed by the operation task in the computing cluster during the current period; the processing module may also be configured to: determine the change amount between the first total amount of data and the second total amount of data, where the second total amount of data is the total amount of data processed by the operation task in the computing cluster during the historical period; predict the third total amount of data processed by the computing cluster in the next period according to the change amount; determine the number of metadata management nodes required for the third total amount of data from the metadata management cluster according to the scheduling policy to implement the scaling adjustment, where the scheduling policy is used to describe the number of metadata management nodes required for different total amounts of data, and the metadata management cluster includes multiple metadata management nodes, and each metadata management node is used to manage the metadata of the data in the storage cluster through a data lake table.
[0022] In some possible examples, the processing module may also be configured to: compare the number of metadata management nodes required for the third total amount of data with the number of metadata management nodes in the working state during the current period; if the number of metadata management nodes in the working state during the current period is less than the number of metadata management nodes required for the third total amount of data, control some of the metadata management nodes in the non-working state in the metadata management cluster to start in the next period, so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total amount of data; allocate the required resources to the started part of the metadata management nodes.
[0023] In some possible examples, the processing module may also be configured to: compare the number of metadata management nodes required for the third total amount of data with the number of metadata management nodes in the working state during the current period; if the number of metadata management nodes in the working state during the current period is greater than the number of metadata management nodes required for the third total amount of data, determine some of the metadata management nodes in the working state during the current period to be closed in the next period, so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total amount of data; after the part of the metadata management nodes is closed, recycle the resources occupied by the part of the metadata management nodes.
[0024] In some possible examples, the processing module may also be used to: compare the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in the working state during the current period; if the number of metadata management nodes in the working state during the current period is equal to the number of metadata management nodes required for the third total data volume, determine that the metadata management nodes in the working state during the current period are used to support the tasks of the computing cluster in the next period.
[0025] In some possible examples, the processing module may also be used to: determine a device list according to the number of metadata management nodes required for the third total data volume, where the device list is used to describe the device information of all metadata management nodes that will be in the working state in the next period; send the device list to each computing node of the computing cluster, so that the computing node selects a target metadata management node from the device list for communication connection, and the target metadata management node is one of the metadata management nodes of the metadata management cluster.
[0026] In some possible examples, the processing module may specifically be used to: if the change amount indicates that the first total data volume is equal to the second total data volume, predict that the third total data volume is equal to the first total data volume; if the change amount indicates that the first total data volume shows a decreasing trend relative to the second total data volume, predict that the third total data volume is the difference between the first total data volume and the change amount; if the change amount indicates that the first total data volume shows an increasing trend relative to the second total data volume, predict that the third total data volume is the sum of the first total data volume and the change amount.
[0027] In some possible examples, the obtaining module may also be used to monitor the status information in the metadata management cluster, where the status information includes the number of metadata management nodes in the working state and / or non-working state, the currently available resource amount, the used resource amount, and the health status information of each metadata management node. So as to realize the scaling adjustment of the metadata management cluster scale according to this status information.
[0028] In a second aspect, an embodiment of the present application provides a computing device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0029] The computing device may be the management node 04 in the embodiment of the present application, and the management node 04 may be a physical server or a virtual machine running on a physical server.
[0030] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when running on a processor, causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0031] In a fourth aspect, an embodiment of the present application provides a computer program product, which, when running on a processor, causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0032] In a fifth aspect, an embodiment of the present application provides a chip, which includes at least one processor and an interface; the at least one processor obtains program instructions or data through the interface; the at least one processor is configured to execute program line instructions to implement the method described in the first aspect or any possible implementation manner of the first aspect.
[0033] In a sixth aspect, an embodiment of the present application provides a data lake warehouse system, which includes: a storage cluster for storing data; a metadata management cluster including a plurality of metadata management nodes, and each metadata management node is configured to manage the metadata of the data in the storage cluster; a computing cluster for receiving an operation task for target data, where the target data is the data in the storage cluster or the data to be written into the storage cluster; the metadata management nodes of the metadata management cluster are further configured to transfer the operation task to the storage cluster for execution; a management node for scaling the number of metadata management nodes in the metadata management cluster according to the total amount of data processed by the operation task in the computing cluster.
[0034] In some possible examples, at least one catalog object is included on the target metadata management node, and the catalog object is configured to manage the metadata of the data lake table, and the metadata of the data in the storage cluster is included in the data lake table; the computing node is further configured to send a connection request to the target metadata management node; the target metadata management node is further configured to establish a connection between the target catalog object and the computing node in response to the connection request, and the metadata of the target data is included in the data lake table of the target catalog object; the computing node is further configured to obtain the metadata of the target data from the catalog object and generate indication information according to the metadata, and the indication information is used to indicate the operation address of the operation task of the target data; the target metadata management node is further configured to call an interface under the catalog object to transfer the executable language of the operation task to the storage cluster for execution and update the data lake table corresponding to the target data when the execution is completed.
[0035] In this way, real-time and high-concurrency data interaction can be achieved between the computing cluster and the storage cluster through the metadata management cluster, which is beneficial to improving the performance of the data lake warehouse system.
[0036] In some possible examples, the storage cluster includes a plurality of storage modules, and the storage module is a database that supports the Simple Storage Service (S3). Compared with HDFS, it has better real-time and low-latency access capabilities, supports efficient storage of small files, avoids waste of storage space and performance in the scenario of small file storage, and can also avoid the limitations in random modification and the disadvantage of insufficient support for concurrent writing existing in HDFS.
[0037] In a seventh aspect, an embodiment of the present application provides a computing device cluster, which is characterized by including at least one computing device, and a program is stored in the memory of the computing device; wherein, when the program stored in the memory of the at least one computing device is executed, the computing device cluster is used to execute the method according to any one of the foregoing first aspect to sixth aspect. The described computing cluster may be the computing cluster 03 in the embodiment of the present application. It can be understood that the beneficial effects of the foregoing second aspect to seventh aspect can be referred to the relevant descriptions in the first aspect, and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the architecture of a data lake warehouse system provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of task processing of the data lake warehouse system in the embodiment of the present application;
[0040] Figure 3 It is a schematic diagram of the architecture of a data lake warehouse system provided by an embodiment of the present application;
[0041] Figure 4 It is a schematic diagram of the management process of the data lake warehouse system in the embodiment of the present application;
[0042] Figure 5 It is a schematic diagram of the process of a data lake warehouse management method provided by an embodiment of the present application;
[0043] Figure 6 It is a schematic diagram of the process of a data lake warehouse management method provided by an embodiment of the present application;
[0044] Figure 7 It is a schematic diagram of the process of a data lake warehouse management method provided by an embodiment of the present application;
[0045] Figure 8 It is a schematic diagram of the structure of a data lake warehouse management device provided by an embodiment of the present application;
[0046] Figure 9 It is a schematic diagram of the structure of a chip provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] As used herein, the term "and / or" describes the associated relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. As used herein, the symbol " / " indicates that the associated objects are in an "or" relationship. For example, A / B means A or B.
[0048] In the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0049] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. For example, a plurality of processing units means two or more processing units, etc.; a plurality of elements means two or more elements, etc.
[0050] To facilitate the understanding of the technical solution of the present application, the terms involved herein are explained below.
[0051] Data Lakehouse: It is a data storage and management architecture that combines the advantages of a data lake and a data warehouse. Among them, a data lake is a system or storage that stores data in its natural / raw format, usually object blocks or files, including copies of raw data generated by the original system and transformed data generated for various tasks, including structured data (rows and columns) from relational databases, semi-structured data (such as CSV, logs, XML, JSON), and unstructured data (such as emails, documents, PDFs, images, audio, videos). A data lake allows various types of data to be directly stored and transformed and analyzed before use.
[0052] Data lake table format: The table format defines the organization method of data and metadata, and provides a unified semantics of "table" upwards.
[0053] HDFS (Hadoop Distributed File System), a distributed file system, is particularly suitable for processing extremely large files and semi-structured and unstructured data. It optimizes the storage and access efficiency of large files through strategies of block storage and parallel reading.
[0054] Hive is a data warehouse infrastructure for querying and managing large-scale datasets in Hadoop. Hive provides a SQL-like query language (HQL), enabling users familiar with SQL to more conveniently process and analyze large-scale data.
[0055] Iceberg is a high-performance and scalable data storage management and analysis tool, especially suitable for data lake scenarios in big data environments. Iceberg aims to provide efficient metadata management for large-scale datasets.
[0056] S3 (Simple Storage Service) is a highly scalable object storage service. It allows users to store and retrieve any amount of data online and access this data anytime, anywhere. S3 provides a highly scalable, durable, secure, fast, and low-cost object storage solution.
[0057] To improve the performance of the data lake warehouse, an embodiment of this application provides a data lake warehouse management method. This method can, in a large-scale data lake warehouse storage scenario, control the scaling of the scale (number of nodes) of the metadata management cluster in the data lake warehouse according to the total amount of data processed by the operation tasks of the computing cluster on the storage cluster in the data lake warehouse, ensuring that the metadata management cluster has sufficient load capacity when a large amount of data enters the lake warehouse, or avoiding performance redundancy of the data lake warehouse services provided by the metadata management cluster when the amount of data processed by the computing cluster is small, thus avoiding resource waste.
[0058] For ease of understanding, the architecture of a data lake warehouse system provided by an embodiment of this application is introduced below.
[0059] Exemplarily, the data lake warehouse system 00 may include a storage cluster 01, a metadata management cluster 02, a computing cluster 03, and a cluster management node 04. Among them:
[0060] The storage cluster 01 includes at least one storage module 1a, 1b, …. The storage modules 1a, 1b, … can be object storage or services that support the Simple Storage Service (S3), or other storage systems suitable for low-latency access and large / small file storage. As an example, the storage modules 1a, 1b, … can be, but are not limited to, the Object Storage Service (OBS), or the object storage system MinIO, etc. These storage modules 1a, 1b, … of the storage cluster 01 form the storage layer of the data lake warehouse, which is used to implement the storage of data in physical storage devices. Among them, the physical storage device can include a hard disk, but is not limited thereto. Compared with the storage capacity provided by HDFS, the storage modules 1a, 1b, … that support S3 have higher advantages in terms of real-time data access and low latency, and are more suitable for application scenarios that require quick responses (such as online services, real-time or interactive applications), achieving a real-time improvement at the millisecond level, and having higher storage efficiency for small files (such as files in the MB or KB level), avoiding waste of storage space and performance. Based on this, it can also achieve high support for concurrent writing and the ability to support random modification.
[0061] The metadata management cluster 02 includes multiple metadata management nodes 2a, 2b, …. These metadata management nodes 2a, 2b, … can be used to provide metadata management services for the storage cluster 01, and are also used to implement data interaction between the storage cluster 01 and the computing cluster 03. Among them, as the service layer of the data lake warehouse, the metadata management nodes 2a, 2b, … can manage the location and status of the data of the storage cluster 01 through metadata files, and provide an interface (Application Programming Interface, API) to support data interaction between the computing cluster 03 and the storage cluster 02, so as to achieve efficient read and write operations.
[0062] Exemplarily, the metadata management nodes 2a, 2b, … can be physical servers or virtualized devices (such as virtual machines or containers), or can also be software deployed on physical servers or virtualized devices. It should be understood that the operation of the metadata management nodes 2a, 2b, … requires certain computing resources, storage resources, and network resources, and these resources come from the physical servers deployed as the metadata management nodes 2a, 2b, …. In this example, the metadata management nodes 2a, 2b, … can be used to provide iceberg (a data lake table format) services. Therefore, in some examples described in this article, the metadata management nodes 2a, 2b, … are also called iceberg nodes.
[0063] Exemplarily, all metadata management nodes 2a, 2b, … of the metadata management cluster 02 can jointly maintain one or more data lake tables (also referred to as Iceberg tables hereinafter). The information recorded in these data lake tables includes, but is not limited to, table definitions and data file lists, etc. Among them, the table definition can be used to describe the organization method of data in the storage layer, and specifically can include information such as columns and partitions. The data file list can be used to describe the access paths and sizes of data files of these partitions (or columns) in the storage layer and other metadata. As a specific example, the metadata management nodes 2a, 2b, … can manage data lake tables through the catalog (file directory) function. The catalog function can implement management operations such as creating (creating a new Iceberg table and storing the metadata of the table), deleting (deleting the Iceberg table and removing all its related metadata), updating (updating the metadata of the table, such as modifying the table structure, adding or deleting columns, etc.), and querying (querying and returning the metadata information of the table so that the system or application can use this information to access and manage the data files corresponding to the table) for these data lake tables. The metadata of the data lake table includes, but is not limited to, the table name, the storage location of the table, etc.
[0064] Optionally, one or more catalog objects can be configured in a metadata management node. The catalog object is an abstracted directory or namespace in the target Iceberg node, has a unique name, and can provide the corresponding catalog function. In this way, the data lake tables on a metadata management node can be managed by different catalog objects, so as to achieve access isolation and permission security control of data files in the storage layer through different catalog objects.
[0065] In this embodiment, the computing cluster 03, as the computing layer of the data lake warehouse, can include at least one computing node 3a, 3b, …. Exemplarily, the computing nodes 3a, 3b, … can be physical servers or virtualized devices (such as virtual machines or containers), or can also be software deployed on these devices. The computing nodes 3a, 3b, … can be connected to the business side, used to receive data from the business side and / or operation tasks on the data, and transfer the data and / or operation tasks to the storage cluster 01 through the metadata management cluster 02 for execution. By way of example and not limitation, the computing nodes 3a, 3b, … can be used to receive data from the business side and execute the operation of writing the data into the specified storage module of the storage cluster 01 through the metadata management cluster 02; or, the computing nodes 3a, 3b, … can be used to receive tasks such as deleting, modifying, and querying data from the business side, and execute the corresponding delete, modify, and query operations on the data of the storage cluster 01 through the metadata management cluster 02, and so on. Among them, the business side can be a user-facing browser, client application or service, etc.
[0066] In some possible implementation manners, each computing node 3a, 3b, … in the computing cluster 03 may maintain a device list, and the device information of the metadata management nodes 2a, 2b, … currently available in the metadata management cluster 02 is recorded in the device list. The device information includes the information required for the computing nodes 3a, 3b, … to connect to the metadata management nodes 2a, 2b, …, specifically including but not limited to the names, IP addresses, and service port numbers of the metadata management nodes 2a, 2b, …, etc.
[0067] In this implementation manner, taking the computing node 3a using the iceberg node to implement data interaction with the storage cluster 01 as an example, the interaction process between the clusters is described. As Figure 2 shown, the computing node 3a receives a processing request for target data sent by the service end in step 1, and the processing request may be an operation request such as deletion, modification, or query of the target data. Then, the computing node 3a can execute step 2 to select a target iceberg node from the currently maintained device list and send a connection request according to the device information of the target iceberg node. It can be understood that the target iceberg node is a node that can manage the target data (or the data set to which the target data belongs), and the selection of the target iceberg node can be achieved through a corresponding preset mechanism. As a possible example rather than a limitation, when the iceberg tables are shared and managed among the iceberg nodes, the computing node can be preset to randomly select a target iceberg table from them, or determine a target iceberg table according to a preset mechanism (round-robin or least connections), etc.
[0068] Then, taking the target iceberg node as the iceberg node 2B as an example, the iceberg node 2B responds to the connection request and establishes a connection between the corresponding catalog object (such as catalog0) and the computing node 3a. Among them, the metadata information such as the location and status of the target data managed in the iceberg table of the catalog object. The "connection" between the catalog object and the computing node 3a can be understood as the association configuration between the catalog object and the computing node 3a.
[0069] Next, the computing node 3a obtains the metadata of the target data (such as information like storage location and status, etc.) from the connected catalog object, and then generates indication information based on the metadata of the target data. This indication information is used to indicate the operation processing corresponding to the processing request for the target data, and the indication information describes on which storage module of the storage cluster 01 the target data for the processing request is to be operated. Then the computing node 3a converts the indication information into an executable SQL language for the target iceberg node 2B and sends it to the target iceberg node 2B through step 3.
[0070] Then the target iceberg node 2B, according to the executable SQL language of the indication information, calls the underlying API interface under the catalog object 0 to communicate with the storage cluster 01 that supports the S3 protocol, so as to instruct the storage cluster 01 to execute the operation tasks described by the SQL language at the corresponding storage location of the storage module where the above target data is located (that is, the storage location of the S3 database therein), such as deleting / modifying / querying the target data at this storage location, etc., and returns the execution result (that is, steps 5 to 7). In addition, after executing this task, step 8 is executed under the catalog object 0 to update the iceberg table corresponding to the target data, realizing the consistent storage of the target data and its metadata.
[0071] It should be understood that the above schematically shows the interaction process between the computing node 3a and an Iceberg node 2B. In actual situations, the computing node 3a needs to determine the number of Iceberg nodes that interact with the computing node 3a and execute tasks according to the total data volume received from the service end. For example, the computing node 3a needs to simultaneously perform data interaction with the Iceberg node 2A and the Iceberg node 2B to implement the processing request of the service end, thereby achieving the performance of high-concurrency business processing. The interaction process between the computing node 3a and the Iceberg node 2A is similar to the interaction process with the Iceberg node 2B and will not be elaborated here.
[0072] Similarly, when the computing node 3b receives a processing request from the service end to write the target data into the storage cluster 01, it also performs an interaction process similar to the above example with the storage cluster 01 through the target ceberg node, thereby writing the target data into the storage cluster 01.
[0073] In this way, real-time and high-concurrency data interaction can be achieved between the computing cluster 03 and the storage cluster 01 through the metadata management cluster 02, which is beneficial to improving the performance of the data lakehouse system 00. In addition, based on the storage module that supports the S3 protocol in the storage cluster 01, the data lakehouse system 00 can support creating, updating, and deleting buckets and storage objects through the API, and is suitable for data storage of various scales. Compared with HDFS, it has better real-time and low-latency access capabilities, supports efficient storage of small files, avoids waste of storage space and performance in the small file storage scenario, and can also avoid the limitations in random modification and the disadvantage of insufficient concurrent write support existing in HDFS.
[0074] In some possible implementation manners, referring again to Figure 1 As shown, the data lakehouse system 00 may further include a management node 04. The management node 04 may be a physical server, a virtualized device, or software deployed on a physical server or a virtualized device. The management node 04 can be used to adaptively adjust the working states of the metadata management nodes 2a, 2b,... in the metadata management cluster 02 according to the task situation processed by the computing cluster 03 in the current period, and allocate corresponding resources (including computing resources, storage resources, network resources, etc.) to the metadata management nodes 2a, 2b,... participating in data processing, so as to achieve high-concurrency data processing and avoid resource waste.
[0075] In a specific example, the data lakehouse system 00 may be deployed in the architecture as Figure 3 shown. Each computing node 3a, 3b,... in its computing layer can be deployed as a big data module supported by Iceberg, such as computing nodes like NiFi, Flink, Spark, etc. Among them, NiFi can process various data sources and data in different formats, can obtain data from one source, transform it, and then push it to another target storage location. Flink is a framework and a distributed processing engine that can be used for stateful computing on unbounded and bounded data streams. Flink can run in a regular cluster environment and can perform calculations at in-memory speed and any scale. Spark is a unified analysis engine for large-scale data processing, with built-in modules for SQL, streaming, machine learning, and graph processing. Each metadata management node 2a, 2b,... in the service layer can be deployed as an Iceberg node to provide Iceberg functions. In addition, each storage module 1a, 2b,... in the storage layer can be a storage product that supports the S3 protocol, such as storage systems like OBS or Minio. And the management node 04 is responsible for dynamically scaling the cluster scale (number of nodes) of the service layer according to the task situation of the computing layer, ensuring a high concurrency rate of the service layer cluster scale while avoiding waste of resources.
[0076] Next, in combination with the attached Figure 4 , the data lake warehouse management principle in the embodiment of this application for the data lake warehouse system 00 will be introduced.
[0077] In some possible implementation manners, as Figure 4 shown, the management node 04 continuously monitors the status information of the metadata management cluster 02, including but not limited to the number of iceberg nodes (including the number of nodes in the working state and the number of nodes in the non-working state), the current available resources (such as CPU, memory, disk space, network bandwidth), the amount of used resources, and the health status of the iceberg nodes (whether they are running normally, whether there are faults, whether they can perform their expected functions, etc.).
[0078] Moreover, the management node 04 also continuously listens to the status information of the computing cluster 03, including but not limited to the number of computing nodes and the amount of data processed by each computing node's current task, etc., and generates corresponding log records.
[0079] In addition, a scheduling policy can be pre-configured on the management node 04. The scheduling policy configures the number of iceberg nodes required by the computing cluster 03 when processing tasks of different data volumes. For example, assuming that the processing capacity of a single iceberg node is not less than 1TB, the scheduling policy can be configured as follows: if the total task data volume ∈ (0, 1TB], the required number of iceberg nodes is 1; if the total task data volume ∈ (1TB, 2TB], the required number of iceberg nodes is 2;.... In this way, after knowing the total data volume of the tasks executed during a certain period, the number of iceberg nodes required for that period can be determined according to the scheduling policy.
[0080] Thus, exemplarily, as Figure 4 shown, after the management node 04 collects the status information of the computing cluster 03 in the current time period through step 11, it can determine the total data volume of the tasks executed by the computing cluster 03 in the current time period according to this status information. Then, the management node 04 searches for the total data volume processed by the computing cluster 03 in the nearest historical time period (which is the sum of the data volumes processed by each computing node) from the historical log records according to a preset time window (which can be at the minute or hour level). Next, the management node 04 can, through step 12, predict the total data volume of the next time period according to the change in the total data volume between the historical time period and the current time, and determine the number of iceberg nodes required for the total data volume of the next time period according to the scheduling policy.
[0081] For example, if in computing cluster 03, the total data volume in the current period is basically equal to that in the previous period, the prediction result can be that the total data volume in the next period is equal to that in the current period. If the total data volume in the current period and the previous period shows a decreasing trend, the reduction amount between the current period and the previous period (i.e., the difference between the total data volume in the previous period and the total data volume in the current period) can be calculated, and it can be predicted that the total data volume in the next period can be the difference between the total data volume in the current period and the reduction amount. If the total data volume in the current period and the previous period shows an increasing trend, the increase amount between the current period and the previous period (i.e., the difference between the total data volume in the current period and the total data volume in the previous period) can be calculated, and it can be predicted that the total data volume in the next period can be the sum of the total data volume in the current period and the increase amount.
[0082] Of course, it is also possible to compare the total data volume of computing cluster 03 in the current period with the total data volume in the previous N (N≥2) periods to predict the total data volume in the next period. Or it is also possible to predict the total data volume in the next period based on the difference in the total data volume between the same periods within different cycles (such as within 24 hours).
[0083] Illustrate with an example. Suppose the total data volume processed by computing cluster 03 within the current hour is 2TB. If the historical log records show that the total data volume processed by computing cluster 03 in the previous period was also 2TB and there is basically no change in the total data volume of these two periods, the prediction result can be that the total data volume in the next period (the next hour) does not exceed 2TB. Similarly, if the total data volume in the previous period was 1.7 and the total data volume shows an increasing trend compared to the current moment, it can be predicted that the total data volume in the next period is 2.3TB, and so on.
[0084] Thus, based on the predicted total data volume in the next period, the possible number of iceberg nodes required in the next period is calculated. Then, continue to refer to Figure 4, compare the number of iceberg nodes currently in working state with the predicted number of nodes required for the next period, and perform step 13 according to the comparison result. For example, if the number of iceberg nodes currently in working state is equal to the predicted number of nodes required for the next period, then ansible node 04 can enable the iceberg nodes currently in working state (such as iceberg nodes 2A and 2B) to continue to provide services for computing cluster 03 in the next period. If after comparison, the number of iceberg nodes currently in working state is lower than the predicted number of nodes required for the next period, then ansible node 04 will start more iceberg nodes (such as iceberg node 2C) to reach the number of nodes required for the next period. On the contrary, if the number of nodes currently in working state is higher than the predicted number of nodes, then ansible node 04 can select a redundant number of iceberg nodes (such as iceberg node 2B) from them, so that after the tasks on these nodes are completed, these redundant iceberg nodes are shut down and the resources allocated to these iceberg nodes are recovered.
[0085] Then, Ansible node 04 allocates corresponding computing resources, storage resources and network resources to the Iceberg nodes required for the next period. In this way, through adaptive dynamic adjustment, the metadata management cluster 02 can achieve high concurrent data processing capabilities in large-scale access task scenarios, while avoiding the waste of resources when there are fewer tasks. It should be noted that having high concurrent data processing capabilities means that the Iceberg nodes currently in working state can share the total amount of data processed by the computing cluster 03.
[0086] In addition, illustratively, the management node 04 may also have a certain fault detection capability. Once an Iceberg node is found to be abnormal, timely measures can be taken, such as rescheduling tasks on the node to other nodes that are working normally) to ensure service continuity and data integrity, and improve the high availability and disaster recovery capabilities of the metadata management cluster.
[0087] In this embodiment, after determining the number of iceberg nodes available in the computing cluster 03 in the next period and the device information of these iceberg nodes, the management node 04 can generate a device list. The device list records the device information of the iceberg nodes currently available in the metadata management cluster, and the device information includes the information required for the computing node to connect to the iceberg node, specifically but not limited to the name, IP address and service port number of the iceberg node, etc.
[0088] Next, the management node 04 sends the device list to each computing node of the computing cluster 03. The computing nodes update the previously received device list with this device list, and select an iceberg node from this device list as the target iceberg node and request a connection. The connection process and the interaction process between clusters 01-03 after the connection can refer to the above description about Figure 2 and will not be elaborated here.
[0089] Next, based on the content described above, a data lakehouse management method provided by an embodiment of the present application will be introduced. It can be understood that this method is proposed based on the content described above, and some or all of the content in this method can refer to the description in the above text.
[0090] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a data lakehouse management method provided by an embodiment of the present application. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. In this example, it is described by taking execution on the above management node 04 as an example. As Figure 5 shown, this method may include:
[0091] S510, obtain the total amount of data processed by the operation tasks of the computing cluster in the data lakehouse.
[0092] In this embodiment, the data lakehouse may be a system including a storage cluster, a metadata management cluster, and a computing cluster, etc. The storage cluster is used to store data, the metadata management cluster is used to manage the metadata of these data, and the computing cluster is used to receive operation tasks from the business side to perform operations such as adding, deleting, modifying, and querying data in the storage cluster.
[0093] Exemplarily, the storage cluster, the metadata management cluster, and the computing cluster may all be clusters including multiple nodes, and the storage cluster and the computing cluster can interact through the metadata management cluster. Among them, each computing node of the computing cluster can execute several operation tasks of the business side, and each operation task is used to process target data, and the target data can be the data already stored in the storage cluster or the data to be written into the storage cluster.
[0094] Moreover, the amount of data processed by each operation task (e.g., the amount of data for deleting / modifying / querying target data or the amount of data when writing target data) can be the same or different. The management node can monitor the amount of data processed by the operation tasks on each computing node within the same time period to obtain the total amount of data processed by the computing cluster during that period. For example, during a certain period, each computing node in the computing cluster receives a write request from the service end to write a large amount of data into the storage cluster, and at the same time, some nodes also receive a modification request to modify data. Then, the management node can monitor the amount of data during the process of the computing cluster processing the data to be written and the amount of data during the process of modifying the data, and the sum of the obtained data amounts is the total amount of data.
[0095] S520. According to the total amount of data, scale up or down the number of metadata management nodes in the metadata management cluster in the data lake warehouse.
[0096] In this step, the management node can dynamically scale up or down the scale of the metadata management cluster in the data lake warehouse, that is, the number of metadata management nodes, according to the obtained total amount of data, so that the metadata management cluster can meet the data processing requirements during the current period and even in the future for a period of time, and avoid wasting resources.
[0097] Next, the principle of the management node scaling up or down the scale of the metadata management cluster will be introduced with reference to the accompanying drawings.
[0098] Exemplarily, as Figure 6 shown, the method may specifically include:
[0099] S610. The computing nodes of the computing cluster receive a processing request for target data.
[0100] In this example, the computing cluster may include one or more computing nodes, and each computing node can receive data and / or processing requests from the service end. Exemplarily, the processing request can be used to request operations such as writing, deleting, modifying, or querying target data, so as to write the target data into the storage cluster, or delete, modify, or query the data in the storage cluster.
[0101] Exemplarily, the storage cluster can include multiple storage modules, and each storage module is a module providing object storage or service with the S3 protocol, supporting the service end to be able to store a large amount of structured or unstructured data (such as pictures, videos, log files, etc.) in real time using the S3 protocol.
[0102] S620. In response to the processing request, the computing node determines a target metadata management node from the metadata management cluster and establishes a connection.
[0103] In this example, each computing node can receive and maintain the device list sent by the management node. This device list records the device information and status information of each metadata management node in the metadata management cluster, etc. The device information may include the information required for the computing node to connect to the metadata management node, specifically including but not limited to the name, IP address, and service port number of the metadata management node. The status information can be used to describe the status of the metadata management node, such as being powered on, connectable, or powered off.
[0104] These metadata management nodes can jointly manage different data sets of the storage cluster through the catalog object. Therefore, for any computing node, when processing the processing request from the service end, the computing node can first randomly select a target metadata management node from the above-mentioned device list that records the metadata management cluster information, or select the metadata management node with the least (or fewer) number of other computing nodes it is connected to as the target metadata management node. It can be understood that for the metadata management node with a smaller number of other computing nodes it is connected to, it undertakes less business volume and has a lighter load. Therefore, it can be preferentially selected as the target metadata management node. It should be understood that for the total data volume in step S510, multiple computing nodes and multiple target metadata management nodes are required for processing, and the device information of multiple target metadata management nodes is included in the device list; step S620 describes the scenario of the interaction between a computing node and a single metadata management node.
[0105] Next, the computing node sends a connection request to the target metadata management node. In response to this request, the target metadata management node causes the corresponding catalog object in the target metadata management node to establish a connection with the computing node. Among them, the catalog object is used to manage metadata such as the location and status of the target data. The "connection" between the catalog object and the computing node can be understood as the associated configuration between the catalog object and the computing node.
[0106] S630. The computing node generates corresponding indication information according to the metadata of the target data provided by the target metadata management node.
[0107] In this example, based on the metadata of the target data provided by the connected catalog object, the computing node can determine which storage module in the storage cluster to perform the corresponding processing operation on the target data. Thus, the computing node generates corresponding indication information. This indication information is used to indicate the operation processing (hereinafter also referred to as the operation task) of the processing request for the target data in corresponding step S610. The indication information describes which storage module in the storage cluster the processing request operates on the target data.
[0108] S640, the computing node converts the operation task into an executable language of the target metadata management node and sends it to the target metadata management node.
[0109] In this example, the computing node can convert the indication information into the SQL language of the target metadata management node and send it to the target metadata management node.
[0110] S650, the target metadata management node instructs the storage cluster to complete the corresponding operation task according to the executable language.
[0111] In this example, according to the received SQL language, the target metadata management node can, under the corresponding catalog object, call the underlying API to communicate with the corresponding storage module of the storage cluster (i.e., the storage module responsible for storing the target data), thereby instructing the storage module to complete the operation task of the target data, such as a write task, or a delete / modify / query task.
[0112] Exemplarily, after completing the operation task, the target metadata management node will receive the execution result of the operation task and update the iceberg table in the catalog object corresponding to the target data according to the execution result, so as to maintain data consistency. In addition, the target metadata management node also returns the current execution result to the computing node. For example, if the operation task is to write the target data, the execution result can be the result of the success or failure of the write operation. If the operation task is to query the target data, the execution result can be the queried target data, and so on.
[0113] Taking a specific application scenario as an example, when a user browses the products on an e-commerce platform and generates browsing data, the computing node can receive this browsing data. The iceberg node connected to it creates a catalog object for the computing node, and the corresponding iceberg table is managed in this catalog object. This iceberg table is used to record metadata information such as the location where these browsing data will be stored. Then, the computing node generates a write-indicating SQL statement. According to this write-indicating SQL statement, the iceberg node calls the API interface to write the browsing data into the specified storage module, completing the ingestion of the browsing data into the lake. Then, the iceberg node updates the corresponding iceberg table in the catalog object to record the metadata of the browsing data (such as the storage location, creation / modification time, etc.).
[0114] Alternatively, in another specific application scenario, a large amount of raw data is stored in each storage module of the storage cluster 01, and the business side needs to use the data of the storage cluster 01 to train a large model. Then, according to the task triggered by the business side, the computing cluster 03 can determine the storage location of the data to be queried through the corresponding catalog object of the connected iceberg node, and then send the sql statement indicating the query to the iceberg node, so as to query this part of the data from the storage cluster through the iceberg node for the business side to train the large model.
[0115] In this way, between the computing cluster and the storage cluster, the high availability and high concurrency rate of the metadata management node cluster can be utilized to achieve real-time and low-latency storage of a large amount of data, thereby improving the data processing efficiency.
[0116] In some possible implementation manners, the management node can monitor the states in the computing cluster and the metadata management cluster, so as to be able to dynamically horizontally match the scale of the metadata management cluster according to the requirements of the computing cluster, improve the concurrency rate of the metadata management node cluster, and avoid wasting resources. Specifically, as Figure 7 shown, this method may further include:
[0117] S701, the management node obtains the first total data volume processed by the operation tasks in the computing cluster during the current period.
[0118] In this example, the management node can continuously monitor the status information of the computing cluster, including but not limited to the number of each computing node and the data volume processed by the current operation tasks of each computing node, etc., and generate corresponding log records.
[0119] Exemplarily, in the computing cluster, each computing node needs to process different operation tasks during the current period, and the sum of the data volumes processed by these operation tasks is the total data volume processed by the cluster during the current period. For the convenience of description, this total data volume is also referred to as the first total data volume.
[0120] S702, determine the change amount between the first total data volume and the second total data volume, where the second total data volume is the total data volume processed by the computing cluster during the historical period.
[0121] In this step, the management node can, according to a preset time window (which can be in minutes or hours), find the total data volume (also referred to as the second total data volume) processed by the computing cluster during the recent historical period from the historical log records, and then determine the change amount between the first total data volume and the second total data volume through comparison. In this way, according to this change amount, the trends and differences such as data volume growth, decrease, or unchanged between the first total data volume and the second total data volume can be determined.
[0122] Exemplarily, it is also possible to compare the total amount of data in the first period of the current period with the total amount of data in the previous N (N≥2) historical periods to predict the total amount of data in the next period. Or it is also possible to predict the total amount of data in the next period based on the difference in the total amount of data between the same periods within different cycles (e.g., within 24 hours). For example, compare the total amount of data in the first period of the computing cluster at 7 am on the current day with the total amount of data in the second period at 7 am on the previous day to determine the change amount.
[0123] S703. Predict the total amount of data in the third period processed by the computing cluster in the next period according to the change amount.
[0124] In this example, the management node can predict the total amount of data (also referred to as the total amount of data in the third period) processed by the computing cluster in the next period according to the obtained change amount through a preset prediction rule.
[0125] Exemplarily, the prediction rule may include:
[0126] If the change amount is 0, that is, the total amount of data in the first and second periods is equal, the prediction result may be that the total amount of data in the third period is equal to the total amount of data in the first period;
[0127] If the total amount of data in the first period is less than the total amount of data in the second period, that is, the data amount shows a decreasing trend from the historical period to the current period, the prediction result may be that the difference between the total amount of data in the first period and the change amount is used as the total amount of data in the third period;
[0128] If the total amount of data in the first period is greater than the total amount of data in the second period, that is, the data amount shows an increasing trend from the historical period to the current period, the prediction result may be that the sum of the total amount of data in the first period and the change amount is used as the total amount of data in the third period.
[0129] S704. Determine the number of metadata management nodes required for the total amount of data in the third period from the metadata management cluster according to the scheduling policy.
[0130] In this example, a scheduling policy is configured in the management cluster, and this scheduling policy is used to describe the number of metadata management nodes required for different total amounts of data. In this way, according to the scheduling policy, the number of metadata management nodes required when the computing cluster processes the task of the total amount of data in the third period in the next period can be predicted.
[0131] S705. Compare the number of metadata management nodes required for the total amount of data in the third period with the number of metadata management nodes in the working state during the current period.
[0132] In this example, the management node continuously monitors the status information of the metadata management cluster, including but not limited to the number of metadata management nodes (including the number of nodes in the working state and the number of nodes in the shutdown state), the health status of the metadata management nodes (whether they are running normally, whether there are faults, whether they can perform their expected functions, etc.), the current available resources (such as CPU, memory, disk space, network bandwidth), and the amount of resources already used, etc.
[0133] Thus, the management node can know how many metadata management nodes support the task of processing the first total data volume during the current period, as well as the resources used by these metadata management nodes and the remaining available resources, and so on.
[0134] Furthermore, the management node can compare the number of metadata management nodes required for the predicted third total data volume with the number of metadata management nodes in the working state during the current period (which can also be the number of metadata management nodes used for the first total data volume). In this way, according to the comparison result, the management node can determine how to adaptively adjust the scale of the data management cluster in the next period.
[0135] S706, if the number of metadata management nodes in the working state during the current period is less than the number of metadata management nodes required for the third total data volume, then control some of the metadata management nodes in the non-working state in the metadata management cluster to start in the next period, so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total data volume.
[0136] In this example, after comparison, the number of metadata management nodes in the working state during the current period is less than the number of metadata management nodes required for the third total data volume, indicating that the number of metadata management nodes in the working state during the current period cannot meet the task requirements in the next period. Then, more metadata management nodes can be started. That is to say, control some of the metadata management nodes in the non-working state in the metadata management cluster to start in the next period (or before), so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total data volume.
[0137] S707, allocate the required resources to the started part of the metadata management nodes.
[0138] In this example, if the metadata management nodes are started, allocate the required computing resources, storage resources, and network resources, etc. to these metadata management nodes, so that this part of the metadata management nodes can participate in the task processing similar to S610 to S650 in the next period.
[0139] S708. If the number of metadata management nodes in the working state during the current period is greater than the number of metadata management nodes required for the third total data volume, then determine some of the metadata management nodes in the working state during the current period to be shut down in the next period, so that the number of metadata management nodes in the working state in the next period reaches the number of metadata management nodes required for the third total data volume.
[0140] In this example, if the number of metadata management nodes in the working state during the current period is greater than the number of metadata management nodes required for the third total data volume, that is, the number of metadata management nodes in the working state during the current period may be redundant in the next period. Therefore, the management node can select some of the metadata management nodes in the working state during the current period to be shut down after completing the current operation tasks. In this way, these metadata management nodes do not participate in the task processing similar to S610 to S650 in the next period, and the number of remaining metadata management nodes in the working state is equal to the number of metadata management nodes required for the third total data volume, avoiding resource waste.
[0141] S709. After these metadata management nodes are shut down, recycle the resources occupied by these metadata management nodes.
[0142] In this example, the management node recycles the resources occupied by these shut-down metadata management nodes for reallocation.
[0143] S710. If the number of metadata management nodes in the working state during the current period is equal to the number of metadata management nodes required for the third total data volume, then determine that the metadata management nodes in the working state during the current period are used to support the tasks of the computing cluster in the next period and complete the task processing similar to S610 to S650.
[0144] In this embodiment, if the number of metadata management nodes in the working state during the current period is equal to the number of metadata management nodes required for the third total data volume, then the management node can make the currently working metadata management nodes continue to provide services for the computing cluster in the next period.
[0145] S711. Determine the device list according to the number of metadata management nodes required for the third total data volume.
[0146] In this example, after the management node determines the number of metadata management nodes and which specific metadata management nodes in the next period, it obtains the device list of these metadata management nodes to describe the device information of all metadata management nodes that will be in the working state in the next period.
[0147] S712. Send the device list to the computing nodes of the computing cluster so that the computing nodes can select a target metadata management node from the device list for communication connection.
[0148] In this embodiment, the management node sends the device list to each computing node of the computing cluster. The computing nodes use the device list to update and overwrite the previously received device list, and select a metadata management node from the device list as the target metadata management node when processing operation tasks to implement the processes of S610 to S650 above, which will not be elaborated here.
[0149] Based on the method in the above embodiment, an embodiment of the present application provides a data lake warehouse management device. Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a data lake warehouse device provided by an embodiment of the present application.
[0150] As Figure 8 shown, the device 800 may include an acquisition module 801 and a processing module 802. Among them, the acquisition module 801 can be used to acquire the total amount of data processed by the operation tasks of the computing cluster in the data lake warehouse, and the operation tasks are used to process target data, and the target data is the data in the storage cluster of the data lake warehouse or the data to be written into the storage cluster. The processing module 802 can be used to scale the scale of the metadata management cluster in the data lake warehouse according to the total amount of data, and the metadata management cluster is used to manage the metadata of the data in the storage cluster.
[0151] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program modules in the device, their implementation principles and technical effects are similar to the descriptions in the above method. The working process of the device can refer to the corresponding process in the above method, which will not be elaborated here.
[0152] Based on the method in the above embodiment, an embodiment of the present application provides a computing device. The computing device may include: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the method in the above embodiment.
[0153] Based on the method in the above embodiment, an embodiment of the present application provides a computing device cluster. The computing device cluster may include at least one computing device, and a program is stored in the memory of the computing device; wherein, when the program stored in the memory of at least one computing device is executed, the computing device cluster is used to execute the method in the above embodiment.
[0154] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0155] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product, characterized in that when the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0156] Based on the method in the above embodiments, an embodiment of the present application further provides a chip. Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a chip provided by an embodiment of the present application. As Figure 9 shown, the chip 900 includes one or more processors 901 and an interface circuit 902. Optionally, the chip 900 may further include a bus 903. Among them:
[0157] The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method may be completed by the integrated logic circuit in the processor 901 or instructions in software form. The above-mentioned processor 901 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0158] The interface circuit 902 may be used for sending or receiving data, instructions, or information. The processor 901 may utilize the data, instructions, or other information received by the interface circuit 902 for processing, and may send the processed information through the interface circuit 902.
[0159] Optionally, the chip 900 further includes a memory. The memory may include a read-only memory and a random access memory, and provide operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM).
[0160] Optionally, the memory stores an executable software module or data structure. The processor may execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions may be stored in the operating system).
[0161] Optionally, the interface circuit 902 may be used to output the execution result of the processor 901.
[0162] It should be noted that the functions corresponding to the processor 901 and the interface circuit 902 can be implemented through hardware design, can also be implemented through software design, or can be implemented through a combination of software and hardware, and there is no limitation here.
[0163] It should be understood that the steps of the above method embodiments can be completed by the logic circuit in the form of hardware or instructions in the form of software in the processor.
[0164] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to the actual situation, can be partially executed, or can be fully executed, and there is no limitation here.
[0165] It can be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.
[0166] The method steps in the embodiments of the present application can be implemented in a hardware manner or can be implemented by the processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in the ASIC.
[0167] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0168] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
Claims
1. A data lake warehouse management method, characterized in that: The method is applied to a management node, and the method comprises: Obtain the total amount of data processed by the operation tasks of the computing cluster in the data lake warehouse, where the operation tasks are used to process target data, and the target data is data in the storage cluster of the data lake warehouse or data to be written to the storage cluster; According to the total data volume, the number of metadata management nodes in the metadata management cluster in the data lake warehouse is expanded or reduced, and the metadata management cluster is used to manage the metadata of the data in the storage cluster.
2. The method according to claim 1, characterized in that The step of scaling the number of metadata management nodes in the metadata management cluster in the data lake warehouse according to the total data volume specifically includes: Obtaining a first total amount of data processed by the operation task in the computing cluster in a current period; Determine a change between the first total data volume and a second total data volume, where the second total data volume is a total data volume processed by operation tasks in the computing cluster within a historical period; Predicting, according to the change amount, a third total amount of data to be processed by the computing cluster in the next period; According to the scheduling strategy, the number of metadata management nodes required for the third total data volume is determined from the metadata management cluster to achieve the scaling adjustment, wherein the scheduling strategy is used to describe the number of metadata management nodes required for different total data volumes, and the metadata management cluster includes multiple metadata management nodes, each of which is used to manage metadata of data in the storage cluster through a data lake table.
3. The method according to claim 2, characterized in that After determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling strategy, the method includes: Compare the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in working state during the current period; If the number of metadata management nodes in working state in the current time period is less than the number of metadata management nodes required for the third total data volume, control some metadata management nodes in non-working state in the metadata management cluster to start in the next time period, so that the number of metadata management nodes in working state in the next time period reaches the number of metadata management nodes required for the third total data volume; Allocate the required resources to the started metadata management node.
4. The method according to claim 2 or 3, characterized in that: After determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling strategy, the method includes: Compare the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in working state during the current period; If the number of metadata management nodes in working state in the current time period is greater than the number of metadata management nodes required for the third total data volume, some metadata management nodes are determined to be shut down in the next time period from the metadata management nodes in working state in the current time period, so that the number of metadata management nodes in working state in the next time period reaches the number of metadata management nodes required for the third total data volume; After the part of metadata management nodes is closed, the resources occupied by the part of metadata management nodes are recovered.
5. The method according to any one of claims 2 to 4, characterized in that: After determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling strategy, the method includes: Compare the number of metadata management nodes required for the third total data volume with the number of metadata management nodes in working state during the current period; If the number of metadata management nodes in working state in the current time period is equal to the number of metadata management nodes required for the third total data volume, the metadata management nodes in working state in the current time period are determined to be used to support the tasks of the computing cluster in the next time period.
6. The method according to any one of claims 2 to 5, characterized in that: After determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling strategy, the method includes: Determine a device list according to the number of metadata management nodes required for the third total data volume, wherein the device list is used to describe device information of all metadata management nodes that will be in working state in the next period; The device list is sent to each computing node of the computing cluster, so that the computing node selects a target metadata management node from the device list for communication connection, and the target metadata management node is a metadata management node of the metadata management cluster.
7. The method according to any one of claims 2 to 6, characterized in that: The predicting, according to the change amount, a third total amount of data to be processed by the computing cluster in the next period of time includes: If the variation indicates that the first total data amount is equal to the second total data amount, then predicting that the third total data amount is equal to the first total data amount; If the variation indicates that the first total data amount is decreasing relative to the second total data amount, then predicting the third total data amount to be the difference between the first total data amount and the variation; If the change amount indicates that the first total data amount shows an increasing trend relative to the second total data amount, then the third total data amount is predicted to be the sum of the first total data amount and the change amount.
8. The method according to any one of claims 2 to 7, characterized in that: Before determining the number of metadata management nodes required for the third total data volume from the metadata management cluster according to the scheduling strategy, the method further includes: Monitor status information in the metadata management cluster, the status information including the number of metadata management nodes in working state and / or non-working state, currently available resources, used resources and health status information of each metadata management node.
9. A computing device, characterized in that include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; Wherein, when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-7.
10. A computing device cluster, characterized in that: comprising at least one computing device, wherein a memory of the computing device Programs are stored; Wherein, when the program stored in the memory of the at least one computing device is executed, the computing device cluster is used to execute the method according to any one of claims 1-7.