Resource Allocation Method, Device, Computer Equipment, and Storage Medium
By real-time monitoring of resource parameters of distributed database clusters and load type determination, rematching cluster users in resource consumption queues, solving the problem of low resource allocation efficiency in the existing technology, and achieving more efficient resource allocation and stability of data processing platform.
Patent Information
- Application Number
- CN202111367947.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-11-18
AI Technical Summary
The existing distributed database resource allocation efficiency is low, especially when high concurrent query requests, concurrent tasks are prone to preemption of resources, resulting in slower response speed of query statements and even unavailable services of the entire data processing platform.
By monitoring the resource parameters of the distributed database cluster based on preloaded task management configuration information, determining the resource pool load type, and rematching the cluster users in the resource consumption queue according to the load type to achieve more reasonable resource allocation.
Through real-time monitoring and dynamic adjustment of resource allocation, the resource allocation efficiency of distributed database clusters is improved, the problems of resource preemption and reduced response speed are avoided, and the stability and high availability of the data processing platform are ensured.
Smart Images

Figure CN114035962B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a resource allocation method, apparatus, computer device, and storage medium. Background Art
[0002] With the continuous development of business and the continuous iterative upgrade of IT technology, the database management system with an all-in-one architecture can no longer meet the needs of massive data storage and high-concurrency data access.
[0003] In related technologies, a database based on a distributed architecture is usually adopted to support a data warehouse system; although the distributed architecture reduces the hardware cost of building a database cluster and improves the processing performance of the database for massive data, it also brings challenges to the load management of the database cluster. If the resources of the distributed database are not reasonably controlled, when performing resource-intensive jobs and high-concurrency query requests, the phenomenon of concurrent tasks preempting resources will occur, resulting in a slow response speed of query statements, and even causing the operating system to exceed the designed load, and even making the entire data processing platform service unavailable. Therefore, the efficiency of the existing resource allocation is still relatively low. Summary of the Invention
[0004] Based on this, it is necessary to provide a resource allocation method, apparatus, computer device, and storage medium for the above technical problems.
[0005] A resource allocation method, characterized by comprising:
[0006] Monitoring resource parameters of a distributed database cluster based on pre-loaded task management configuration information;
[0007] When it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determining the resource pool load type of the distributed database cluster according to the resource parameters;
[0008] Determining cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
[0009] In one embodiment, the resource parameters include the number of first database statements and the number of second database statements;
[0010] The determining the resource pool load type of the corresponding distributed database cluster according to the resource parameters includes:
[0011] Obtaining the ratio of the number of first database statements to the number of second database statements as the resource pool load value;
[0012] If the resource pool load value is greater than or equal to the preset load threshold, determine that the corresponding distributed database cluster is of the first resource pool load type; if the resource pool load value is less than the preset load threshold, determine that the resource pool load type of the corresponding distributed database cluster is of the second resource pool load type.
[0013] In one embodiment, before determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type, it further includes:
[0014] Obtain the resource consumption parameters and weight parameters of the cluster users;
[0015] Determine the resource consumption value of the cluster users according to the resource consumption parameters and the weight parameters;
[0016] If the resource consumption value is greater than or equal to the preset resource consumption threshold, identify the resource consumption type of the corresponding cluster user as the heavy resource consumption type; if the resource consumption value is less than the preset resource consumption threshold, identify the resource consumption type of the corresponding cluster user as the light resource consumption type.
[0017] In one embodiment, the determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type includes:
[0018] Allocate the cluster users of the heavy resource consumption type to the distributed database cluster of the second resource pool load type, and allocate the cluster users of the light resource consumption type to the distributed database cluster of the first resource pool load type.
[0019] In one embodiment, after monitoring the resource parameters of the distributed database cluster, it further includes:
[0020] Obtain the historical resource parameters of the distributed database cluster; the historical resource parameters carry a parameter source identifier;
[0021] Obtain the weight information matching the parameter source identifier, and perform calculation processing on the historical resource parameters according to the weight information to obtain the statistical result of the historical resource parameters;
[0022] Send the statistical result to the preset terminal.
[0023] In one embodiment, after monitoring the resource parameters of the distributed database cluster, it further includes:
[0024] Generate alarm message information according to the resource parameters of the distributed database cluster;
[0025] Filter the alarm message information according to the preset warning configuration information to obtain the target alarm message information;
[0026] Send the target alarm message information to a preset terminal.
[0027] A resource allocation device, the device includes:
[0028] A resource parameter monitoring module, configured to monitor the resource parameters of a distributed database cluster based on pre-loaded task management configuration information;
[0029] A load type determination module, configured to determine the resource pool load type of the distributed database cluster according to the resource parameters when it is detected that the resource parameters of the distributed database cluster meet the preset resource allocation conditions;
[0030] A cluster user matching module, configured to determine the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
[0031] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0032] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0033] A computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0034] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0035] Monitor the resource parameters of the distributed database cluster based on pre-loaded task management configuration information;
[0036] When it is detected that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determine the resource pool load type of the distributed database cluster according to the resource parameters;
[0037] Determine the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
[0038] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0039] Monitor the resource parameters of the distributed database cluster based on pre-loaded task management configuration information;
[0040] When it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determine the resource pool load type of the distributed database cluster according to the resource parameters;
[0041] Determine the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
[0042] The above resource allocation method, device, computer device and storage medium, the method includes: monitoring the resource parameters of the distributed database cluster based on pre-loaded task management configuration information; when it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determining the resource pool load type of the distributed database cluster according to the resource parameters; determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type. This application monitors the resource parameters of the distributed database cluster by pre-loading task management configuration, determines the load type of the resource pool of the distributed database cluster, and then re-determines the cluster users in the queue, realizing the allocation of different resources to suitable cluster users according to the real-time situation of the distributed database cluster, improving the resource allocation efficiency. Description of the Drawings
[0043] Figure 1 It is a schematic flowchart of the resource allocation method in an embodiment;
[0044] Figure 2 It is a structural block diagram of the resource allocation device in an embodiment;
[0045] Figure 3 It is a schematic structural diagram of the resource allocation system in an embodiment;
[0046] Figure 4 It is a schematic diagram of the resource allocation system communicating with multiple distributed databases in an embodiment;
[0047] Figure 5 It is a schematic flowchart of a resource allocation process of the resource allocation system in an embodiment;
[0048] Figure 6 It is a schematic flowchart of a process control module executing in the resource allocation system in an embodiment;
[0049] Figure 7 It is a schematic flowchart of another task management module executing in the resource allocation system in an embodiment;
[0050] Figure 8It is a schematic flow diagram of another task warning module executed in a resource allocation system in an embodiment;
[0051] Figure 9 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0053] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties; correspondingly, the present disclosure also provides a corresponding user authorization entry for the user to choose to authorize or choose to refuse.
[0054] It should be noted that the resource allocation determination method and device of the present disclosure can be used in the cloud computing field, and can also be used in any field other than the cloud computing field. The application field of the resource allocation determination method and device of the present disclosure is not limited.
[0055] In one embodiment, as Figure 1 shown, a resource allocation method is provided, including the following steps:
[0056] Step S110, monitor the resource parameters of the distributed database cluster based on the pre-loaded task management configuration information.
[0057] Among them, the task management configuration information is the configuration information preset and loaded in the server to monitor the resource parameters of the distributed database cluster, including information such as monitoring time, monitoring frequency, and monitoring data type.
[0058] Step S120, when it is detected that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determine the resource pool load type of the distributed database cluster according to the resource parameters.
[0059] Among them, the preset resource allocation conditions are the index conditions of the resource parameters determined according to the task management configuration information; when the index of a certain resource parameter exceeds the corresponding index condition, it means that the resource parameter needs to be allocated so that the resource parameter can be adjusted after allocation to achieve the effect of efficient utilization of resources.
[0060] Among them, during the process of using the resource pool, there are various different load types according to the resource utilization situation. For example, a high load means that more resources in the resource pool are currently being utilized and it cannot undertake the resource consumption of new cluster users. A medium load means that the resources currently being utilized in the resource pool are moderate and it can still undertake the resource consumption requests of some new cluster users. A low load means that fewer resources in the resource pool are currently being utilized and there are more idle resources, which can undertake the resource consumption requests of new cluster users. It should be noted that the load type of the resource pool can be dynamically adjusted according to the actual scenario of the application of the present disclosure, not limited to high, medium, low and other load types.
[0061] Step S130, determine the cluster users in the resource consumption queue of the distributed database cluster according to the load type of the resource pool.
[0062] Among them, the resource pool is a component for the distributed database cluster to provide resources externally. The resource pool can provide various resources including computing resources, data resources, information resources, etc. The present disclosure does not limit the types of resources stored in the resource pool.
[0063] Among them, the resource consumption queue is the way for cluster users to obtain resources in the resource pool of the distributed database cluster.
[0064] In the above resource allocation method based on the distributed database, by preloading the task management configuration to monitor the resource parameters of the distributed database cluster, determining the load type of the resource pool of the distributed database cluster, and then re-determining the cluster users in the queue, it realizes the allocation of different resources to suitable cluster users according to the real-time situation of the distributed database cluster, improving the resource allocation efficiency.
[0065] In one embodiment, the resource parameters include the number of the first database statements and the number of the second database statements;
[0066] Determining the load type of the resource pool of the corresponding distributed database cluster according to the resource parameters includes:
[0067] Obtain the ratio of the number of the first database statements to the number of the second database statements as the resource pool load value; if the resource pool load value is greater than or equal to the preset load threshold, determine that the corresponding distributed database cluster is of the first resource pool load type; if the resource pool load value is less than the preset load threshold, determine that the load type of the resource pool of the corresponding distributed database cluster is the second resource pool load type.
[0068] Among them, the first database statement and the second database statement refer to two different types of database statements, such as fast statement concurrent resources and slow statement concurrent resources. Fast statements refer to SQL statements with short running time and small occupation of computing resources; slow statements refer to SQL statements with long running time and large occupation of computing resources.
[0069] Among them, the resource pool load value is a judgment value for the load situation of the resource pool; the current load type of the resource pool can be determined according to the size of the resource pool load value; the resource pool load value can be updated in real time and dynamically changed according to the utilization situation of the resource pool.
[0070] Among them, the preset load threshold is the basis for judging the load type corresponding to the resource pool load value. For example, if the preset load threshold is set to be above 10 for high load, that is, the first resource pool load type, and below 10 for medium load, that is, the second resource pool load type, then if the resource pool load value is 8, it is determined as the second resource pool load type, and after reaching 15, it is determined as the first resource pool load type.
[0071] In this embodiment, the ratio of the number of the first database statements to the number of the second database statements is used as the resource pool load value to describe the load situation of the resource pool, realizing the accurate acquisition of the real-time load situation of the resource pool; at the same time, the load type of the resource pool can be intuitively determined through the preset load threshold, improving the resource allocation efficiency.
[0072] In one embodiment, before determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type, it further includes:
[0073] Obtain the resource consumption parameters and weight parameters of the cluster users; determine the resource consumption value of the cluster users according to the resource consumption parameters and weight parameters; if the resource consumption value is greater than or equal to the preset resource consumption threshold, then identify the resource consumption type of the corresponding cluster user as the heavy resource consumption type; if the resource consumption value is less than the preset resource consumption threshold, then identify the resource consumption type of the corresponding cluster user as the light resource consumption type.
[0074] Among them, the resource consumption parameter is the evaluation basis for the speed of resource consumption of the cluster users, and the weight parameter is the adjustment value for adjusting the resource consumption parameter of the cluster users. The resource consumption value refers to the size, speed, etc. of the resources consumed by the cluster users.
[0075] Among them, the resource consumption threshold is the boundary for distinguishing the resource consumption types of the cluster users. If the resource consumption value of the cluster users is less than this resource consumption threshold, then the resource consumption type of the cluster users is identified as the light resource consumption type. If the resource consumption value of the cluster users is greater than or equal to this resource consumption threshold, then the resource consumption type of the cluster users is identified as the heavy resource consumption type.
[0076] In this embodiment, the resource consumption value of the cluster users is determined through the resource consumption parameters and weight parameters of the cluster users, and then the consumption types of the cluster users are classified and judged according to the resource consumption threshold, improving the efficiency of judging the consumption types of the cluster users.
[0077] In one embodiment, determining cluster users in the resource consumption queue of a distributed database cluster according to the resource pool load type includes:
[0078] Assigning cluster users of the heavy resource consumption type to the distributed database cluster of the second resource pool load type, and assigning cluster users of the light resource consumption type to the distributed database cluster of the first resource pool load type.
[0079] Specifically, the heavy resource consumption type refers to cluster users that consume resources more and faster; therefore, matching cluster users of the heavy resource consumption type with the distributed database cluster of the second resource pool load type enables the distributed database cluster of the second resource pool load type to adapt to the resource consumption capacity of cluster users of the heavy resource consumption type, so as to achieve a rapid response and processing of the database processing requirements of cluster users of the heavy resource consumption type.
[0080] Similarly, the light resource consumption type refers to cluster users that consume fewer resources and have a slower resource consumption speed; therefore, matching cluster users of the light resource consumption type with the distributed database cluster of the first resource pool load type enables the distributed database cluster of the first resource pool load type to adapt to the resource consumption capacity of cluster users of the light resource consumption type, so as to achieve a rapid response and processing of the database processing requirements of cluster users of the light resource consumption type.
[0081] In this embodiment, by assigning cluster users of the heavy resource consumption type to the distributed database cluster of the second resource pool load type, and assigning cluster users of the light resource consumption type to the distributed database cluster of the first resource pool load type, it is possible to adjust the cluster users received according to the real-time resource pool load type of the distributed database cluster, that is, to achieve the allocation of resources of the distributed database cluster, and improve the resource allocation efficiency.
[0082] In one embodiment, after monitoring the resource parameters of a distributed database cluster, it further includes:
[0083] Obtaining the historical resource parameters of the distributed database cluster; the historical resource parameters carry a parameter source identifier; obtaining the weight information matching the parameter source identifier, calculating and processing the historical resource parameters according to the weight information to obtain the statistical result of the historical resource parameters; and sending the statistical result to a preset terminal.
[0084] Among them, the historical resource parameters refer to the resource parameter situation of the distributed database cluster before the current moment, and the historical operation situation of the distributed database cluster can be determined according to the historical resource parameters.
[0085] Among them, the parameter source identifier refers to the mark of the source of resource parameters. Since the importance of data generated from different data sources is different, by setting different weight information, it is possible to identify resource parameters that have a greater impact on system operation and resource parameters that have a smaller impact on system operation in the resource parameters.
[0086] Among them, the statistical result refers to the analysis result of the historical operation of the distributed database cluster using historical resource parameters. Through the statistical result, problems in the operation process of the distributed database cluster can be discovered for timely adjustment.
[0087] In this embodiment, by obtaining the statistical result of historical resource parameters, the historical operation of the distributed database cluster is mastered.
[0088] In one embodiment, after monitoring the resource parameters of the distributed database cluster, it further includes:
[0089] Generating alarm message information according to the resource parameters of the distributed database cluster; screening the alarm message information according to the preset early warning configuration information to obtain target alarm message information; and sending the target alarm message information to a preset terminal.
[0090] Among them, the alarm message information is a warning that there may be problems with the distributed database cluster judged and analyzed according to the resource parameters of the distributed database cluster; however, the data of the distributed database cluster is extensive and the running time is long, so the quality of the generated alarm message information may be inconsistent. Therefore, it is necessary to screen the alarm message information, and the screened target alarm message information is sent to the preset terminal to reduce the amount of information obtained by the terminal personnel and improve the efficiency of the terminal personnel in judging the problems of the distributed database cluster.
[0091] In one embodiment, as Figure 2 shown, a resource allocation device is provided, including: a resource parameter monitoring module 21, a load type determination module 22, and a cluster user matching module 23, where:
[0092] The resource parameter monitoring module 21 is used to monitor the resource parameters of the distributed database cluster based on the pre-loaded task management configuration information;
[0093] The load type determination module 22 is used to determine the resource pool load type of the distributed database cluster according to the resource parameters when it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions;
[0094] The cluster user matching module 23 is used to determine the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
[0095] To further illustrate the present application, a resource allocation system applying the resource allocation method of the present disclosure is described. The structure of the resource allocation system is as shown in Figure 3 shown, and at least includes: a process control module 31 and a task management module 32; the process control module 31 and the task management module 32 are communicatively connected;
[0096] Among them, the task management configuration is equivalent to the task management configuration information, and the preset task condition is equivalent to the preset resource allocation condition.
[0097] The process control module 31 is configured to load the task management configuration in response to a start instruction; monitor the status of at least one distributed database cluster based on the task management configuration; and generate a task execution instruction corresponding to at least one preset task condition when it is detected that the status of at least one distributed database cluster meets at least one preset task condition;
[0098] The task management module 32 is configured to obtain resource parameters of at least one distributed database cluster when the task type of the task execution instruction is a resource allocation task; determine the load type of the resource pool of the distributed database cluster according to the resource parameters, and re-determine the cluster users matching the resource pool according to the load type.
[0099] Among them, the process control module may at least include a general control and self-check unit; the self-check unit at least includes a parameter parsing unit and an environment inspection module; the general control provides functions such as overall process control and retry of failed tasks, that is, it can globally control the overall process and obtain the running status of the system in real time. The parameter parsing unit and the environment inspection unit in the self-check unit can perform legality verification on input parameters, load necessary configuration parameters, and pre-check the availability of the environment. If the pre-check process passes and all configuration parameters are loaded, a signal is sent to the general control indicating that the resource allocation system has entered the running state.
[0100] Among them, the task management module communicates directly with each distributed database cluster and at least includes a computing resource information collection module unit, a computing resource allocation strategy management unit, and a computing resource regulator unit. The computing resource information collection unit is used to collect the core configuration and metrics of the database computing resources, which is equivalent to the "eyes" of the present resource allocation system; the computing resource allocation strategy management unit is used to execute the intelligent generation and management of the database cluster resource strategy relying on statistical data stored in time series, historical operation and maintenance experience, expert rules, etc., which is equivalent to the "brain" of the present resource allocation system; the computing resource regulator unit is used to obtain and submit change commands for the resource management strategies of each set of distributed database clusters, so that the resource allocation strategies in the computing resource allocation strategy management unit take effect in the distributed database cluster, which is equivalent to the "hand" of the present resource allocation system.
[0101] Among them, the distributed database cluster can be a data processing cluster composed of distributed parallel processing databases; a distributed parallel processing database (Massive Parallel Processing Database, MPPDB) is a database that adopts a large-scale distributed, share-nothing processing architecture; it consists of multiple nodes that have independent and non-shared system resources such as CPUs, memory, and storage; in such a system architecture, business data is scattered and stored on multiple physical nodes, and data analysis tasks are pushed to be executed near the location where the data is located. Through the coordination of the control module, large-scale data processing work is completed in parallel, achieving a fast response to data processing.
[0102] Among them, resource allocation refers to the reasonable allocation of CPU resources, memory resources, IO resources, etc. of the distributed database cluster to avoid the inefficient operation of the system or system operation problems caused by unreasonable resource occupation.
[0103] Among them, a resource pool (Resource Pool) is a distributed database system resource configuration mechanism used to divide host resources (such as memory, IO, etc.) and provide concurrent control capabilities for Structured Query Language. The resource pool manages resources by binding control groups (Control Groups, Cgroups); cluster users can achieve resource load management of the jobs under them by binding to the resource pool. A control group is a mechanism provided by the Linux kernel that can limit, record, and isolate the physical resources (such as CPUs, memory, IO, etc.) used by a process group. If a process joins a certain control group, the control group has strict restrictions on the system resources of Linux, and the process cannot exceed its maximum limit when using these resources.
[0104] Among them, the start instruction can be generated and sent by the terminal or automatically generated by the process control module according to a preset time period to achieve periodic processing of resource allocation.
[0105] Such as Figure 4 As shown in the figure, it is a schematic diagram of the communication connection between the resource allocation system and multiple distributed databases; it includes: big data platform 12, big data platform ETL scheduling and management subsystem 13, resource allocation system 14, big data platform operation and maintenance center 15, distributed database cluster MPP-DBMS1 16, distributed database coordination CN node 17, distributed database data DN node 18, distributed database cluster management service module 19, distributed database system management node 20, distributed database security management node 21;
[0106] Specifically, the big data platform 12 is a system that realizes complex data processing and calculation based on massive data and business requirements. It consists of a scheduling and management subsystem 13, a resource allocation and management device 14, an operation and maintenance center 15, and a computing engine with several distributed database clusters 16 at the bottom layer. As the core computing engine of the big data platform, the distributed database cluster 16 mainly consists of the following parts: a coordination CN node 17 responsible for coordinating computing tasks, a DN data node 18 for storing distributed data, and a cluster management module 19 responsible for managing and monitoring each functional unit and physical resource of the database.
[0107] Illustrated with a typical application scenario: The ETL scheduling and management module 13 submits data processing instructions to the designated distributed database cluster CN node 17 through JDBC or a dedicated client. The distributed database cluster CN node 17 distributes the execution plan to each data DN node and aggregates the processing execution results of all DN nodes. The distributed database cluster management service module 19 coordinates the system management node 20 and the security management node 21, and is responsible for managing and monitoring the operation of each functional unit and physical resource in the distributed system to ensure the stable operation of the entire distributed database cluster. During the operation of the data processing job, the big data platform users and operation and maintenance management personnel can directly monitor the data processing process through the platform operation and maintenance management center 15.
[0108] It should be noted that the resource allocation system 14 mentioned in the present invention can be used as a functional component of the big data platform and communicate directly with the core subsystems in the big data platform (such as the scheduling system 13 that initiates jobs, the distributed database cluster 16 at the bottom layer computing engine, the operation and maintenance management center 15, etc.) through interfaces or clients. It can perceive the busy degree of the bottom layer computing engine in real time. By analyzing the priority weights, job quantities, and operation peaks and valleys of different job groups in the scheduling system 13, and the load conditions of the distributed database cluster 16 at the bottom layer computing engine, and combining with the operation and maintenance monitoring experience data accumulated by the operation and maintenance management center 15, the resource elastic allocation management device 14 can generate a reasonable computing resource allocation strategy to automatically allocate the computing resource parameters of the distributed database cluster and maximize the utilization rate of computing resources.
[0109] Specifically, the process control module loads corresponding task management configuration information according to the start instruction; the task management configuration information may include connection status information of each connected distributed database cluster, corresponding communication interfaces, operation history records, monitoring trigger conditions, etc., that is, the task management configuration information is used for the process control module to more accurately monitor the status of the distributed database cluster. When the process control module monitors that the status of at least one distributed database cluster reaches the preset task conditions corresponding to the monitoring trigger conditions, it generates a task execution instruction corresponding to the preset task conditions and sends it to the task management module, so that the task management module performs resource allocation and other processing on the monitored distributed database cluster according to the task execution instruction. After receiving the task execution instruction, the task management module first determines the task type corresponding to the task execution instruction. If the task type is a resource allocation task, it obtains the resource parameters of at least one distributed database cluster; the resource parameters can reflect the real-time operation status of the distributed database cluster. The task management module performs calculation processing based on the collected resource parameters to obtain a numerical or level information that can reflect a certain load type, and determines the load type corresponding to the resource pool of the distributed database cluster according to the numerical or level information; different cluster users have different resource consumption capabilities, so they can be redistributed to the resource pool corresponding to the appropriate load type for mounting according to the resource consumption capabilities of the cluster users, realizing the reasonable allocation of resources.
[0110] Specifically, the process control module is further configured to respond to the start instruction, load the task warning configuration; monitor the status of at least one distributed database cluster based on the task warning configuration; and generate a warning execution instruction corresponding to at least one preset warning condition when it is monitored that the status of at least one distributed database cluster meets at least one preset warning condition.
[0111] Specifically, in response to the start instruction, the process control module loads the task warning configuration in addition to the task management configuration; the task warning configuration includes multiple preset warning conditions, and each preset warning condition includes various warning parameters, which can perform real-time calculation according to the status of the distributed database cluster to determine whether one of the preset warning conditions is reached. If a certain preset warning condition is reached, a corresponding warning execution instruction is generated according to the triggered preset warning condition.
[0112] The resource allocation system further includes: a task warning module; the task warning module is communicatively connected to the process control module; the task warning module is configured to respond to the warning execution instruction and determine the warning task type of the warning execution instruction; in the case where the warning task type is a statistical analysis task, obtain the operation and maintenance data of the distributed database cluster within a preset time range; and input the operation and maintenance data into the statistical analysis model to obtain a statistical analysis result.
[0113] Among them, the task warning module is an important operation and maintenance assistance part of the resource allocation system, which can serve the operation and maintenance team in the production environment. Its function is to accumulate operation and maintenance monitoring indicators in the time dimension, conduct intelligent analysis based on such data, so as to provide a reliable decision-making basis for the task management module and the maintenance team of the resource allocation system. The task warning module can at least include a statistical analysis unit and an alarm unit, and can further integrate monitoring data analysis components and corresponding alarm platforms, such as influxDB (an open-source (MIT) distributed time series database developed in Go language) for monitoring data recording or analysis, ClickHouse (a columnar database management system for online analytical processing (OLAP)), Grafana (a cross-platform open-source metric analysis and visualization tool), etc., and Prometheus (an open-source monitoring and alarm system and time series database (TSDB) developed by SoundCloud) platform for event monitoring and alarm. The task warning module can be set in a form that does not communicate directly with the distributed database, so that processes such as monitoring analysis and alarm do not occupy the computing resources of business queries and data processing batches, improving the operation efficiency of the resource allocation system.
[0114] Specifically, after the task warning module responds to the warning execution instruction, it first judges the warning task type of the warning execution instruction; in the case where the warning task type is a statistical analysis task, it obtains the operation and maintenance data of the distributed database cluster within a preset time range; inputs the operation and maintenance data into the statistical analysis model to obtain the statistical analysis result. The statistical analysis result can be stored as historical data on the operation and maintenance of the distributed database cluster for subsequent statistical analysis; it can also be further used to generate corresponding messages for warning notifications. The task warning module is also used to obtain the message to be sent corresponding to the message sending task in the case where the warning task type is a message sending task; screen the message to be sent according to the preset warning strategy to obtain the target message to be sent; send the target message to be sent to the target terminal.
[0115] Specifically, if the warning task type specifically refers to a message sending task, the task warning module sorts out all messages to be sent according to the preset warning strategy, determines whether each message to be sent is sent, as well as the specific time, object, cycle, etc. of sending. Finally, the message to be sent is sent to the target terminal; the target terminal performs corresponding processing according to the received message information to ensure normal operation.
[0116] The task management module is also used to obtain the computing resource information and the task information to be processed of the distributed database cluster in the case where the task type is an information query task; obtain the resource parameters of the distributed database cluster according to the computing resource information and the task information to be processed.
[0117] Among them, the information query task is used to obtain the computing resource information of the distributed database cluster and the information of tasks to be processed; the task management module can determine the resource parameters of the distributed database cluster according to the computing resource information and the information of tasks to be processed.
[0118] Furthermore, the task management module is also used to determine the resource consumption information and weight information of each cluster user according to the resource parameters; determine the resource consumption type of each cluster user according to the resource consumption information and weight information, and the resource consumption type corresponds to the load type of the resource pool; re-determine the cluster users matching the load type of the resource according to the resource consumption type.
[0119] Among them, different cluster users correspond to different weights and resource consumption capabilities. Therefore, the resource consumption type of the cluster user can be determined according to the resource consumption information and weight information of the cluster user; and different resource pools correspond to different load types; therefore, the load efficiency of the resource pool and the resource consumption type of the cluster user can be matched, so that the resources of the distributed database cluster can be reasonably allocated.
[0120] For example, as Figure 5 shown is a schematic diagram of a resource allocation process of a resource allocation system; specifically, it is a schematic diagram of the resource allocation situation before and after adjustment of a single distributed database cluster under the action of the resource allocation system; among them, the cylinder (i.e., the resource pool resource pool) represents a resource pool of a single CN control node of a distributed database. Each resource pool contains 53 concurrent resources for slow statements and 54 concurrent resources for fast statements respectively. Each distributed database cluster user 55 can only mount one resource pool, and the users mounting the same resource pool belong to the same queue; fast statements refer to SQL statements with short running time and small occupancy of computing resources; slow statements refer to SQL statements with long running time and large occupancy of computing resources.
[0121] Specifically, in the initial state, each resource pool has the same proportion of fast and slow statements; all cluster users are grouped into queues based on business volume and business weight. Each queue contains users with high weight and high business volume (i.e., strong resource consumption ability) and users with low weight and low business volume (i.e., low resource consumption ability). Different types of cluster users in each queue are distributed as evenly as possible, and try to avoid allocating users with heavy resource consumption to the same queue. After the distributed database cluster provides services externally, the resource allocation system starts to intervene in the management of computing resources, accumulates various monitoring index data in the running state of the cluster, automatically judges the time point of resource adjustment, and executes automatic adjustment of resource parameters.
[0122] Based on the resource parameters obtained from each distributed database cluster, the resource allocation system makes adaptive adjustments to the ratio of fast and slow statements in each distributed database and the queues where cluster users are located respectively, and realizes the re-determination of cluster users matching the resource pool. On the one hand, it identifies the resource pool with heavy consumption of cluster computing resources and the resource pool with low consumption of resources (that is, determines the resource pools of two load types); on the other hand, it re-groups the cluster users, marks the cluster users with heavy resource consumption but high weights, re-assigns them to the heavy resource consumption queue, and appropriately reduces the number of cluster users in this queue. After the resource parameter allocation is completed, these cluster users will obtain the tilt of cluster computing resources, which can ensure that the jobs submitted by such cluster users are preferentially executed in the cluster on the premise of stable operation of the cluster. For cluster users with light resource consumption and low business weights, they are re-assigned to the light resource consumption queue, and the number of users in this queue can be appropriately increased. After the resource allocation is completed, these users can maximize the utilization rate of the allocated cluster computing resources and still complete queries or batch jobs according to the scheduled time without occupying too much cluster resources.
[0123] In the above resource allocation system, the process control module loads the corresponding task management configuration information according to the start instruction to generate a task execution instruction; the task management module obtains the resource parameters of the distributed database cluster according to the task execution instruction, determines the load type of the resource pool, and then re-determines the mounted cluster users, realizing the allocation of different resources to suitable cluster users according to the real-time situation of the distributed database cluster, and improving the resource allocation efficiency.
[0124] In one embodiment, as Figure 6 shown, a process executed by the process control module in the resource allocation system is provided, including:
[0125] Step S601: After the resource allocation system based on the distributed database is deployed, the main control program starts.
[0126] Step S602: Start the self-check unit and execute the pre-check process, including the legality check of parameters, environmental inspection, and availability check of external systems or services connected.
[0127] Step S603: Load the task management configuration or task warning configuration. The task management configuration is a task for the configuration and allocation of computing resources that directly interacts with the backend distributed database cluster; the task warning configuration does not directly interact with the backend distributed data cluster, but performs intelligent analysis based on the monitoring data with a time dimension in the production environment, and sends a message to the operation and maintenance management platform based on the analysis data, providing accurate decision-making basis for the operation and maintenance management of the big data platform.
[0128] Step S604: After the main control program obtains all task management configurations or task warning configurations, it uses an asynchronous parallel method to check whether the execution conditions of relevant tasks are met at the running time. Specifically, in an asynchronous thread, if situations such as the distributed database cluster being temporarily unavailable, the time-series database for monitoring data storage being unavailable, or the pre-configuration information being incomplete occur, it is determined that the task execution conditions are not met, the thread status is reset, and the task management configuration or task warning configuration is not triggered; if all task execution conditions are met, step S605 is invoked asynchronously.
[0129] Step S605: The main control program determines that the current environment does not meet the running conditions of the task management configuration or task warning configuration, and asynchronously invokes the task management module or task warning module according to the type of specific task configuration to further execute the corresponding process.
[0130] Step S606: After the relevant tasks in step S605 are completed, the main control program records the successful status or failure error message of the corresponding task execution in the result table.
[0131] Step S607: The main control program enters the sleep state and returns to step S63 after the sleep ends to enter the next loop.
[0132] In one embodiment, as Figure 7 shown, another process for the task management module to execute in the resource allocation system is provided, including:
[0133] After the process control module invokes the task management module asynchronously in step S604, after the task management module is invoked, it enters the execution process corresponding to the task type, including:
[0134] Step S701: The main control unit sends a task execution instruction and passes all necessary parameters to the task management module, such as cluster connection information, resource management policy parameters, etc.
[0135] Step S702: The task management module determines whether the current task is only an information query task. If so, it enters the computing resource information collection branch of step S703, will not change the resource configuration parameters of the distributed database cluster, and will not affect the existing business operations; if not, it enters the computing resource allocation branch of step S707, changes the resource configuration ratio of the distributed database cluster, and the computing resources occupied by the cluster users on the cluster are elastically allocated to improve the data processing batch or query job execution efficiency on the distributed database cluster.
[0136] Steps S703 - S706 are an independent information query task, and the details are as follows:
[0137] Step S703: This is the first step of the information collection branch for the distributed database cluster. The task management module collects computing resource information from the coordination CN node or the cluster management module of the distributed database cluster through JDBC or HTTP Client respectively, and obtains the scale of the jobs to be submitted from the scheduling management subsystem. Specifically, the information collected from the distributed database includes pre-configured load parameters on the distributed database cluster (such as the number of statements that can be executed concurrently at the same time for each cluster user queue), real-time load metrics of each cluster user queue on the distributed database cluster (such as the number of running SQLs for each cluster user queue in the current database), real-time and historical resource usage of job runs for each cluster user queue in the distributed database cluster, etc.; the information collected from the scheduling management subsystem includes the number of jobs to be submitted, the priority of the scheduling job group, etc. After all the information is collected, the task management module formats and assembles the relevant data and proceeds to the next step S704 for further processing.
[0138] Step S704: The computing resource monitoring data summary processing model classifies the data obtained in step S703 into two categories according to the source: real-time computing resource parameters of the distributed database cluster and job parameters of the scheduling system, and encapsulates the unprocessed information into a query result object in a fixed format.
[0139] Step S705: Obtain the query result object in the previous step, call the operation data update interface in step S706, and then return an instruction indicating that the query task is completed to the request initiator.
[0140] Step S706: Obtain the query result object from the previous step, cache the result in the time series database asynchronously, and archive the results with too long time limits, record the details of the computing resource monitoring information for other modules to call.
[0141] Steps S707 - S710 are an independent resource allocation task, and the details are as follows:
[0142] Step S707: Load the resource allocation policy pool, select the optimal resource allocation method from the statistical data relying on time series storage, historical operation and maintenance experience, expert rules, combined with the real-time operation status of the distributed database cluster.
[0143] Step S708: Map the resource allocation policy selected in step S707 into two types of change parameters, namely cluster-side parameter adjustment and reallocation of cluster user queues. Among them, the resource parameters of the distributed database cluster, including the number of concurrency, CPU quota, I / O quota, etc., reconfigure the resource pool parameters of the distributed database cluster; specifically, the reallocation of cluster user queues is to group and mark the distributed database cluster users according to the policy.
[0144] Step S709: Submit the parameters output in step S708 to the distributed database cluster to make the resource allocation strategy formulated in step S707 effective, asynchronously call the operation data update interface of S710, and the allocation task is completed.
[0145] Step S710: Obtain the specific parameter results of the resource allocation strategy from the previous step, cache the results in the time series database asynchronously, archive the results with too long time limit, and record the details of the effective calculation resource strategy allocation for other modules to analyze and audit.
[0146] In one embodiment, as Figure 8 shown, another execution process of the task warning module in the resource allocation system is provided, including:
[0147] After the process control module starts the task warning module asynchronously in step S604, after the task warning module is started, it enters the execution process corresponding to the task type, including:
[0148] Step S801: The main control unit sends a warning execution instruction to the task warning module and passes all necessary parameters to the task warning module, such as cluster connection information, resource management strategy parameters, alarm task queue information, etc.
[0149] Step S802: The task warning module determines whether the current task is only a statistical analysis task. If so, it enters the monitoring information statistical analysis branch of step S803 and performs intelligent analysis based on various monitoring metrics accumulated in the operation database; if not, it enters the alarm branch of step S807 and conveys the exception message to the operation and maintenance team through the alarm platform.
[0150] Steps S803 - S806 are an independent statistical analysis task, and the details are as follows:
[0151] In step S803, the task warning module loads the key operation and maintenance data within a certain time span from the monitoring database, specifically including the throughput of the distributed database cluster, the resource consumption of the operating system, the running status of the scheduling jobs, etc., and transfers the loaded data as input parameters to step S804.
[0152] In step S804, the task warning module loads the statistical analysis strategy. Different statistical analysis strategies assign different weights to the monitoring data according to the source. After being calculated by the statistical analysis model, the monitoring data analysis results are obtained, and the analysis results are encapsulated and transmitted to step S805.
[0153] Step S805 obtains the monitoring data analysis results and distributes them to different channels, such as the resource allocation center, the alarm platform, other subscription channels, etc.
[0154] Step S806 records the relevant information of the current statistical task, such as category, distribution channel, execution time, etc., to the operation center as an audit record.
[0155] Steps S807 - S810 are an independent alarm task, and the details are as follows:
[0156] Step S807 loads and parses all the messages to be sent from the detailed list of alarm information.
[0157] Step S808 loads the pre-configured alarm policies, filters the message information that does not conform to the rules, obtains the final alarm message object to be distributed, and transmits it to Step S809.
[0158] Step S809 sends the alarm message to various channels, such as emails, text messages, or a unified alarm platform, etc.
[0159] Step S810 records the relevant information of the current alarm task, such as category, distribution channel, execution time, etc., to the operation center as an audit record.
[0160] It should be understood that although Figure 1 、 6 each step in the flowchart of Figure 1 、 6 -8 is shown in sequence according to the arrow indication, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0161] For the specific limitations of the resource allocation device, reference can be made to the limitations on the resource allocation method in the above text, which will not be elaborated here. Each module in the above resource allocation device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0162] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 9As shown. The computer device includes a processor, a memory, and a network interface connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store resource allocation data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a resource allocation method.
[0163] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0164] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0165] Based on the pre-loaded task management configuration information, monitor the resource parameters of the distributed database cluster;
[0166] When it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determine the resource pool load type of the distributed database cluster according to the resource parameters;
[0167] According to the resource pool load type, determine the cluster users in the resource consumption queue of the distributed database cluster.
[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented:
[0169] Based on the pre-loaded task management configuration information, monitor the resource parameters of the distributed database cluster;
[0170] When it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determine the resource pool load type of the distributed database cluster according to the resource parameters;
[0171] According to the resource pool load type, determine the cluster users in the resource consumption queue of the distributed database cluster.
[0172] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0173] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0174] The above various embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A resource allocation method, characterized in that, Including: Monitoring the resource parameters of a distributed database cluster based on pre-loaded task management configuration information; the resource parameters include the number of first database statements and the number of second database statements; When it is monitored that the resource parameters of the distributed database cluster meet the preset resource allocation conditions, determining the resource pool load type of the distributed database cluster according to the resource parameters, including: obtaining the ratio of the number of first database statements to the number of second database statements as the resource pool load value; If the resource pool load value is greater than or equal to the preset load threshold, determining that the corresponding distributed database cluster is of the first resource pool load type; if the resource pool load value is less than the preset load threshold, determining that the resource pool load type of the corresponding distributed database cluster is the second resource pool load type; Determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
2. The method according to claim 1, characterized in that, Before determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type, it further includes: Obtaining the resource consumption parameters and weight parameters of the cluster users; Determining the resource consumption value of the cluster users according to the resource consumption parameters and the weight parameters; If the resource consumption value is greater than or equal to the preset resource consumption threshold, identifying the resource consumption type of the corresponding cluster user as a heavy resource consumption type; if the resource consumption value is less than the preset resource consumption threshold, identifying the resource consumption type of the corresponding cluster user as a light resource consumption type.
3. The method according to claim 2, characterized in that, The determining the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type includes: Allocating the cluster users of the heavy resource consumption type to the distributed database cluster of the second resource pool load type, and allocating the cluster users of the light resource consumption type to the distributed database cluster of the first resource pool load type.
4. The method according to claim 1, characterized in that, After monitoring the resource parameters of the distributed database cluster, it further includes: Obtaining the historical resource parameters of the distributed database cluster; the historical resource parameters carry a parameter source identifier; Obtaining the weight information matching the parameter source identifier, and performing calculation processing on the historical resource parameters according to the weight information to obtain the statistical result of the historical resource parameters; Sending the statistical result to a preset terminal.
5. The method according to claim 1, characterized in that, After monitoring the resource parameters of the distributed database cluster, it further includes: Generating alarm message information according to the resource parameters of the distributed database cluster; Filtering the alarm message information according to the preset early warning configuration information to obtain the target alarm message information; Sending the target alarm message information to a preset terminal.
6. A resource allocation device, characterized in that, The device includes: A resource parameter monitoring module, configured to monitor the resource parameters of a distributed database cluster based on pre-loaded task management configuration information; the resource parameters include the number of first database statements and the number of second database statements; A load type determination module, configured to determine the resource pool load type of the distributed database cluster according to the resource parameters when it is detected that the resource parameters of the distributed database cluster meet the preset resource allocation conditions; The load type determination module is specifically configured to obtain the ratio of the number of the first database statements to the number of the second database statements as the resource pool load value; if the resource pool load value is greater than or equal to a preset load threshold, it is determined that the corresponding distributed database cluster is of the first resource pool load type; if the resource pool load value is less than the preset load threshold, it is determined that the resource pool load type of the corresponding distributed database cluster is the second resource pool load type; A cluster user matching module, configured to determine the cluster users in the resource consumption queue of the distributed database cluster according to the resource pool load type.
7. A computer device, including a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer program product, including a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Cloud database system and cloud database resource dynamic adjustment method
CN107085539A
Multi-user task scheduling method and device for calculation cluster
CN107291545A