Database cluster management optimization method and device, equipment and storage medium

By intelligently dynamic allocation of cluster resources and real-time detection of node health status, and generating optimization suggestions strategies, the problem of lack of intelligence and automation of existing cluster management methods is solved, and the automation and optimization of cluster management is realized, and the system operation efficiency and stability is improved.

CN120029848APending Publication Date: 2025-05-23SHANGHAI DONGPU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510064658.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The lack of intelligent and automated support for existing cluster management methods leads to waste of resources, performance bottlenecks, and increased management complexity.

Method used

By identifying each cluster node in the database cluster, collecting performance data, determining resource requirements, monitoring resource usage, analyzing load changes, dynamically adjusting cluster size, realizing intelligent and dynamic allocation of resources, and real-time detection of node health status, and generating optimization suggestions strategies.

Benefits of technology

It realizes automation and optimization of cluster management, improves the overall operating efficiency and operation stability of the system, and reduces resource waste and management complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029848A_ABST
    Figure CN120029848A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information processing, in particular to an optimization method, device and equipment for database cluster management and a storage medium, and aims at optimizing the database cluster management by identifying each cluster node in a database cluster, collecting performance data of each cluster node, determining resource requirements according to the performance data and monitoring resource use conditions in each cluster node. The method comprises the steps of analyzing a resource use condition, adjusting a cluster scale of each cluster node according to load change in a load change condition analysis result, dynamically allocating cluster resources according to resource requirements and the cluster scale, and realizing automation and optimization of cluster management, so that the overall operation efficiency of a system is improved; the node health state of each cluster node is detected in real time after resource allocation, the optimization suggestion strategy is generated according to the node health state, and the optimization suggestion strategy is sent to the management end, so that the cluster management strategy is adjusted in time, and the overall operation stability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to an optimization method, device, equipment and storage medium for database cluster management. Background Art

[0002] Cluster management refers to the management process of configuring, monitoring, scheduling, maintaining and optimizing a group of computers or servers (i.e., a cluster). A cluster is usually composed of multiple computers that work together and share tasks to provide higher computing power, reliability and scalability. The purpose of cluster management is to ensure that the resources in the cluster can run efficiently and stably and can cope with load changes or hardware failures. At present, cluster management is a key challenge in large-scale computing and data processing environments. Cluster resource allocation, load balancing, troubleshooting and other issues require efficient management and optimization solutions. Traditional cluster management methods often lack intelligent and automated support, resulting in resource waste, performance bottlenecks and increased management complexity. Summary of the invention

[0003] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide an optimization method, device, equipment and storage medium for database cluster management that realizes automation and optimization of cluster management through intelligent dynamic allocation of cluster resources, thereby improving the overall operating efficiency of the system, and generates optimization recommendation strategies according to the health status of nodes after resource allocation, thereby improving the overall operating stability of the system.

[0004] The first aspect of the present invention provides an optimization method for database cluster management, comprising: identifying each cluster node in the database cluster and collecting performance data of each of the cluster nodes; determining resource requirements based on the performance data and monitoring resource usage in each of the cluster nodes; analyzing the resource usage to obtain load change analysis results, and adjusting the cluster scale of each of the cluster nodes based on the load changes in the load change analysis results; dynamically allocating cluster resources based on the resource requirements and the cluster scale, and detecting the node health status of each of the cluster nodes in real time; generating an optimization recommendation strategy based on the node health status, and sending the optimization recommendation strategy to a management end.

[0005] Optionally, in a first implementation method of the first aspect of the present invention, the identifying each cluster node in the database cluster and collecting the performance data of each cluster node include: calling the cluster configuration file of the database cluster and extracting all node IDs listed in the cluster configuration file; identifying all corresponding cluster nodes in the database cluster according to the node ID; and collecting the performance data of each cluster node, the performance data including CPU usage, memory usage, disk I / O and network bandwidth.

[0006] Optionally, in a second implementation method of the first aspect of the present invention, determining resource requirements based on the performance data and monitoring resource usage in each of the cluster nodes include: calculating resource utilization of the cluster nodes based on the performance data; determining resource requirements based on the resource utilization; and using the Prometheus monitoring tool to regularly collect resource usage in each of the cluster nodes.

[0007] Optionally, in a third implementation method of the first aspect of the present invention, the resource usage is analyzed to obtain a load change analysis result, and the cluster size of each cluster node is adjusted according to the load change in the load change analysis result, including: obtaining historical resource usage data, and performing data cleaning on the historical resource usage data; constructing a resource usage trend prediction model based on the historical resource usage data; analyzing the resource usage using the resource usage trend prediction model to obtain a load change analysis result; and adjusting the cluster size of each cluster node according to the load change in the load change analysis result.

[0008] Optionally, in a fourth implementation method of the first aspect of the present invention, the cluster resources are dynamically allocated according to the resource requirements and the cluster size, and the node health status of each of the cluster nodes is detected in real time, including: generating a node task requirement table according to the resource requirements and the cluster size, and configuring the priority of the node task requirement table according to the urgency and importance of the task; retrieving the cluster resources corresponding to the node task requirement table based on the priority; allocating the cluster resources to each of the cluster nodes using a resource scheduling algorithm; and detecting the node health status of each of the cluster nodes in real time.

[0009] Optionally, in a fifth implementation method of the first aspect of the present invention, generating an optimization recommendation strategy based on the node health status and sending the optimization recommendation strategy to a management end includes: determining whether the node health status is an unhealthy state; if so, backing up the data of the cluster node corresponding to the unhealthy state, and troubleshooting the cluster node; restarting the cluster node after troubleshooting, and obtaining a troubleshooting log; generating an optimization recommendation strategy based on the fault type and severity in the troubleshooting log; evaluating and verifying the optimization recommendation strategy to obtain an evaluation and verification result; when the evaluation and verification result is passed, generating a visualization strategy report for the optimization recommendation strategy, and sending the visualization strategy report to a management end.

[0010] Optionally, in a sixth implementation method of the first aspect of the present invention, after generating an optimization recommendation strategy based on the node health status and sending the optimization recommendation strategy to the management end, it also includes: obtaining the policy application result feedback from the management end, and obtaining a feedback timestamp from a timestamp service platform based on the feedback time; merging the feedback timestamp, the optimization recommendation strategy and the strategy application result to obtain merged information; encrypting the merged information to obtain encrypted information; and uploading the encrypted information to the blockchain.

[0011] The second aspect of the present invention provides an optimization device for database cluster management, including: an identification and collection module, used to identify each cluster node in the database cluster and collect performance data of each cluster node; a determination monitoring module, used to determine resource requirements based on the performance data and monitor resource usage in each cluster node; an analysis and adjustment module, used to analyze the resource usage, obtain load change analysis results, and adjust the cluster size of each cluster node according to the load changes in the load change analysis results; an allocation detection module, used to dynamically allocate cluster resources according to the resource requirements and the cluster size, and detect the node health status of each cluster node in real time; a generation and sending module, used to generate an optimization recommendation strategy based on the node health status, and send the optimization recommendation strategy to the management end.

[0012] Optionally, in a first implementation manner of the second aspect of the present invention, the identification and collection module includes: a calling extraction unit, used to call the cluster configuration file of the database cluster, and extract all node IDs listed in the cluster configuration file; an identification unit, used to identify all corresponding cluster nodes in the database cluster according to the node ID; and a first collection unit, used to collect performance data of each of the cluster nodes, the performance data including CPU usage, memory usage, disk I / O and network bandwidth.

[0013] Optionally, in a second implementation of the second aspect of the present invention, the monitoring module includes: a calculation unit, used to calculate the resource utilization of the cluster node based on the performance data; a determination unit, used to determine the resource requirement based on the resource utilization; and a second collection unit, used to use the Prometheus monitoring tool to periodically collect resource usage in each of the cluster nodes.

[0014] Optionally, in a third implementation method of the second aspect of the present invention, the analysis and adjustment module includes: an acquisition and cleaning unit, used to acquire historical resource usage data and perform data cleaning on the historical resource usage data; a construction unit, used to construct a resource usage trend prediction model based on the historical resource usage data; an analysis unit, used to use the resource usage trend prediction model to analyze the resource usage and obtain a load change analysis result; and an adjustment unit, used to adjust the cluster size of each cluster node according to the load changes in the load change analysis result.

[0015] Optionally, in a fourth implementation of the second aspect of the present invention, the allocation detection module includes: a generation configuration unit, used to generate a node task requirement table according to the resource requirements and the cluster size, and configure the priority of the node task requirement table according to the urgency and importance of the task; a calling unit, used to call the cluster resources corresponding to the node task requirement table based on the priority; an allocation unit, used to use a resource scheduling algorithm to allocate the cluster resources to each of the cluster nodes; and a detection unit, used to detect the node health status of each of the cluster nodes in real time.

[0016] Optionally, in a fifth implementation of the second aspect of the present invention, the generation and sending module includes: a judgment unit, used to judge whether the health status of the node is an unhealthy state; a backup and troubleshooting unit, used to back up the data of the cluster node corresponding to the unhealthy state if so, and perform troubleshooting on the cluster node; a restart acquisition unit, used to restart the cluster node after troubleshooting, and obtain the troubleshooting log; a generation unit, used to generate an optimization recommendation strategy according to the fault type and severity in the troubleshooting log; an evaluation and verification unit, used to evaluate and verify the optimization recommendation strategy to obtain an evaluation and verification result; a generation and sending unit, used to generate a visualization strategy report for the optimization recommendation strategy when the evaluation and verification result is passed, and send the visualization strategy report to the management end.

[0017] Optionally, in a sixth implementation method of the second aspect of the present invention, it also includes: an acquisition module, used to obtain the policy application results fed back by the management end, and obtain a feedback timestamp from a timestamp service platform based on the feedback time; a merging module, used to merge the feedback timestamp, the optimization suggestion strategy and the policy application result to obtain merged information; an encryption module, used to encrypt the merged information to obtain encrypted information; and an uploading module, used to upload the encrypted information to the blockchain.

[0018] A third aspect of the present invention provides an optimization device for database cluster management, the optimization device for database cluster management comprising: a memory and at least one processor, the memory storing instructions; at least one of the processors calling the instructions in the memory so that the optimization device for database cluster management executes each step of any one of the above-mentioned optimization methods for database cluster management.

[0019] A fourth aspect of the present invention provides a computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the steps of any of the above-mentioned methods for optimizing database cluster management are implemented.

[0020] In the technical solution of the present invention, by identifying each cluster node in the database cluster, collecting the performance data of each cluster node, determining the resource demand based on the performance data, monitoring the resource usage in each cluster node, analyzing the resource usage, adjusting the cluster scale of each cluster node based on the load change in the load change analysis result, dynamically allocating cluster resources based on resource demand and cluster scale, realizing automation and optimization of cluster management, thereby improving the overall operation efficiency of the system, detecting the node health status of each cluster node in real time after resource allocation, generating an optimization recommendation strategy based on the node health status, and sending the optimization recommendation strategy to the management end so as to adjust the cluster management strategy in time, thereby improving the overall operation stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A first flow chart of the optimization method for database cluster management provided by an embodiment of the present invention;

[0022] Figure 2 A second flow chart of the optimization method for database cluster management provided by an embodiment of the present invention;

[0023] Figure 3 A third flow chart of the optimization method for database cluster management provided by an embodiment of the present invention;

[0024] Figure 4 A fourth flow chart of the optimization method for database cluster management provided by an embodiment of the present invention;

[0025] Figure 5 A schematic diagram of the structure of an optimization device for database cluster management provided by an embodiment of the present invention;

[0026] Figure 6 Another structural schematic diagram of the optimization device for database cluster management provided by an embodiment of the present invention;

[0027] Figure 7A schematic diagram of the structure of an optimization device for database cluster management provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The present invention provides a method, device, equipment and storage medium for optimizing database cluster management, which realizes automation and optimization of cluster management by intelligently and dynamically allocating cluster resources, thereby improving the overall operating efficiency of the system, and generates optimization suggestion strategies according to the health status of nodes after resource allocation, thereby improving the overall operating stability of the system.

[0029] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , an embodiment of the optimization method of database cluster management in an embodiment of the present invention includes:

[0031] 101. Identify each cluster node in the database cluster and collect performance data of each cluster node;

[0032] In this embodiment, the cluster configuration file is located, and a programming language (such as Python, Go, or Bash) is used to parse the configuration file. After parsing, detailed information (such as node ID, node IP, role, etc.) of each node in the cluster can be extracted. By parsing the node list in the cluster configuration file, the unique identifier of each node is extracted, the role type of the node is confirmed, and Prometheus, Zabbix, or Nagios is configured to collect performance data from each cluster node. After the monitoring tool is configured, it is necessary to ensure that real-time data is obtained from each cluster node and the data is stored in a suitable storage system. The performance data of the node, such as CPU, memory, disk, and network, is automatically collected, and monitoring agents are deployed on each cluster node to ensure that they can regularly push performance data to the Prometheus server.

[0033] 102. Determine resource requirements based on performance data and monitor resource usage in each cluster node;

[0034] In this embodiment, the peak value, average value and fluctuation of the CPU usage of each node are calculated, the high-load period is determined, and the number of CPU cores required to ensure normal operation is estimated. The memory usage of each node is analyzed, especially the maximum memory usage, memory swap and memory leakage are checked. The disk read and write volume and network traffic of the node are analyzed to evaluate whether there is a bottleneck. In particular, high-frequency disk I / O operations and large-scale data transmission may affect the cluster performance. If the CPU usage is often close to or reaches 100%, it means that the CPU is a bottleneck resource, and it is necessary to increase the number of CPU cores or adjust the load balancing. If the system frequently swaps memory or the memory usage is too high, it means that the memory resources are insufficient. At this time, it is necessary to increase the memory or optimize the query and data storage strategy. If the disk read and write speed is very slow, or the disk I / O queue length is very long, it means that the disk may be a bottleneck, and it is necessary to upgrade the hard disk or optimize the storage architecture (such as using SSD, adding RAID configuration). If the network bandwidth is too high and causes an increase in the communication delay between nodes, it means that the network may be a bottleneck. It is necessary to evaluate whether it is necessary to increase the bandwidth or optimize the network topology structure, and monitor the resource usage in each cluster node.

[0035] 103. Analyze resource usage to obtain load change analysis results, and adjust the cluster size of each cluster node according to the load change in the load change analysis results;

[0036] In this embodiment, the resource usage is analyzed to obtain the load change analysis results. The load change refers to the change of resource consumption of nodes in the cluster over time. The trend of system load in the past period of time (for example, hourly, daily or weekly) is checked, and the time period and duration of load peak are analyzed. For example, the load of some clusters may increase abnormally during a specific period of time. At this time, the scale of the cluster may need to be increased. If the load peak reaches the resource limit, it may affect the performance. At this time, the cluster needs to be expanded, the load distribution between the nodes is checked, and whether there is any node overload or idle situation. When the system load approaches or exceeds the processing capacity of the node, the number of cluster nodes can be increased by automatic expansion. For example, during the peak load period, the system can dynamically add more nodes to share the load. When the load is low, the system can reduce the size of the cluster to reduce resource waste. For example, when the CPU and memory usage of the node are low for a long time, the number of nodes can be reduced to save costs. The automatic expansion tool is used to realize automatic cluster scale adjustment based on load changes.

[0037] 104. Dynamically allocate cluster resources according to resource requirements and cluster size, and detect the node health status of each cluster node in real time;

[0038] In this embodiment, during peak load periods, the size of the cluster is automatically expanded according to the increase in resource demand, and more nodes are added. Conversely, when the load decreases, costs are saved by reducing unnecessary nodes. Horizontal expansion specifically increases or decreases the number of nodes, and the load is shared by adding more computing or storage nodes. Vertical expansion specifically increases the resources of a single node (such as CPU, memory, etc.) to cope with the increased load, but this is usually limited by the capabilities of the hardware. When the cluster size increases, resources can be flexibly allocated by configuring elastic resource pools in the cloud environment. Check whether the node is still alive. If the node is detected to be in an unrecoverable state (such as a service crash), it will be removed from the cluster to avoid affecting other nodes. Check whether the node is ready to accept traffic. Even if the node survives, if its internal service is not ready to process requests (for example, at startup), it needs to be temporarily removed from the load balancing pool to prevent request errors. Perform health checks on the node's resources (such as CPU, memory, disk I / O, etc.) to ensure that the node resources are not overloaded.

[0039] 105. Generate an optimization suggestion strategy based on the node health status, and send the optimization suggestion strategy to the management end;

[0040] In this embodiment, once the health status data is collected, an in-depth analysis is performed to discover potential bottlenecks or anomalies. According to the set thresholds and health check rules, the system will automatically determine whether each node has an abnormality. For example, the utilization rate of resources such as CPU, memory, disk or network is higher than the set threshold, the node has a resource bottleneck, a service (such as database, Web service, etc.) fails to start as expected, responds to frequent errors, or has a long delay, the node continuously times out and does not respond to the heartbeat signal, or the node has a hardware failure. Through comparative analysis of real-time indicators and historical data, any behavior that deviates from the normal operating state is discovered in a timely manner. Once a health status abnormality is detected, optimization suggestions are automatically generated according to the type of abnormality. The optimization suggestions include resource expansion or reduction, service optimization, node health status optimization, hardware repair suggestions, and load balancing and scheduling optimization. According to the severity of the health status and the impact on cluster performance, the optimization strategy can be prioritized. When the optimization suggestions are generated, the cluster management system will encapsulate these suggestions into an optimization suggestion report and send it to the management end or administrator in an appropriate manner.

[0041] In an embodiment of the present invention, by identifying each cluster node in the database cluster, collecting performance data of each cluster node, determining resource requirements based on the performance data, monitoring resource usage in each cluster node, analyzing resource usage, adjusting the cluster scale of each cluster node based on load changes in the load change analysis results, dynamically allocating cluster resources based on resource requirements and cluster scale, realizing automation and optimization of cluster management, thereby improving the overall operating efficiency of the system, detecting the node health status of each cluster node in real time after resource allocation, generating optimization recommendation strategies based on the node health status, and sending the optimization recommendation strategies to the management end so as to adjust the cluster management strategies in a timely manner, thereby improving the overall operating stability of the system.

[0042] See also Figure 2 The second embodiment of the optimization method for database cluster management in the embodiment of the present invention includes:

[0043] 201. Call the cluster configuration file of the database cluster and extract all node IDs listed in the cluster configuration file;

[0044] In this embodiment, the cluster configuration file of the database cluster is called, the cluster configuration file is read, the file is parsed and node information is extracted according to different configuration file formats, and the node information is saved as a list, which includes all node IDs listed in the cluster configuration file.

[0045] 202. Identify all corresponding cluster nodes in the database cluster according to the node ID;

[0046] In this embodiment, the node ID is a unique identifier for each cluster node. Based on the node ID provided by the user, the corresponding node is searched in the node information of the cluster. If the cluster configuration file contains multiple nodes and there may be multiple node ID matches, further deduplication or other processing can be performed. Once a matching node ID is found, detailed information of the node can be returned, including IP address, role, health status, etc.

[0047] 203. Collect performance data of each cluster node, including CPU usage, memory usage, disk I / O, and network bandwidth;

[0048] In this embodiment, the performance data collection task is executed by combining Python scripts with system commands to collect performance data of each cluster node, including CPU usage, memory usage, disk I / O, and network bandwidth.

[0049] 204. Calculate resource utilization of cluster nodes based on performance data;

[0050] In this embodiment, based on the performance data, the resource utilization of the cluster nodes is calculated. The resource utilization includes CPU utilization, memory utilization, disk I / O utilization and network bandwidth utilization. The CPU utilization indicates the current usage percentage of the CPU, the memory utilization indicates the ratio of the current used memory to the total memory, the disk I / O utilization indicates the ratio of the current disk read and write operations to the maximum read and write capacity of the disk, and the network bandwidth utilization indicates the percentage of the current network transmission bandwidth to the total available bandwidth.

[0051] 205. Determine resource requirements based on resource utilization;

[0052] In this embodiment, the current resource usage status of the cluster nodes can be evaluated according to the resource utilization rate, and the future resource demand can be inferred based on this.

[0053] 206. Use the Prometheus monitoring tool to regularly collect resource usage in each cluster node;

[0054] In this embodiment, Prometheus is installed and deployed, and the Prometheus monitoring tool is used to regularly collect resource usage in each cluster node. Prometheus stores the captured time series data in the local disk.

[0055] 207. Obtain historical resource usage data and perform data cleaning on the historical resource usage data;

[0056] In this embodiment, historical resource usage data is obtained through the Prometheus API. The return data of Prometheus is usually in JSON format, including a timestamp and a corresponding indicator value. Outliers are filtered out by setting a threshold, and missing values ​​are filled by linear interpolation using valid data points before and after. Duplicates are removed by timestamp, and only the last record or average value at each time point is retained. The data is standardized so that the values ​​of different indicators are on the same scale. A common method is to subtract the mean and divide by the standard deviation to convert the data to a fixed range. The data is smoothed using methods such as rolling window averaging, weighted averaging, or moving average to reduce high-frequency short-term fluctuations.

[0057] 208. Construct a resource usage trend prediction model based on historical resource usage data;

[0058] In this embodiment, historical resource usage data is converted into features suitable for a machine learning model, a suitable prediction algorithm model is selected, the prediction algorithm model is trained using the features to obtain a resource usage trend prediction model, and the trained resource usage trend prediction model is deployed in a production environment, and data is updated regularly and real-time predictions are performed.

[0059] 209. Analyze resource usage using the resource usage trend prediction model to obtain load change analysis results;

[0060] In this embodiment, a trained model is used to predict future load conditions. The trend of resource usage is identified through the results of model prediction. For example, the CPU load may increase significantly during certain time periods, or memory usage may increase as user visits increase. By analyzing the load changes, the peak moments of resource usage can be identified. The predicted load conditions can be used to find potential bottlenecks in the system, and the load change analysis results can be obtained based on the load change trend.

[0061] 210. Adjust the cluster size of each cluster node according to the load change in the load change analysis result;

[0062] In this embodiment, a cluster size adjustment strategy is formulated according to the load change in the load change analysis result. When the load forecast shows that the demand is low, some nodes can be reduced to save resource costs. When the load forecast shows that the demand is high, increasing the number of nodes can effectively improve the computing power and processing power, and avoid performance degradation or system crash caused by insufficient resources. If the CPU load is high and it is predicted that it will be in a high load state for a long time, more computing nodes can be added or the computing power of existing nodes can be improved. If it is predicted that the memory usage will reach a bottleneck, more memory or storage nodes can be added to ensure that the cluster can smoothly handle the load. If it is predicted that the network bandwidth is insufficient to support high-traffic requests, this problem can be alleviated by adding network nodes or adjusting the network architecture. The cluster size is automatically adjusted using container orchestration tools (such as Kubernetes, DockerSwarm, etc.). As the cluster size is adjusted, the load balancer (such as Nginx, HAProxy, cloud load balancing service) needs to make corresponding configuration adjustments to ensure that requests can be evenly distributed to each node to avoid the situation where some nodes are overloaded and other nodes are idle.

[0063] In the embodiment of the present invention, the process realizes accurate monitoring of cluster node resource utilization by automatically collecting and analyzing performance data, reduces manual intervention and ensures the accuracy of resource configuration, and combines real-time data and historical trend predictions to intelligently adjust the cluster scale according to load changes, thereby improving the elasticity and scalability of the system. Through detailed resource utilization analysis and load prediction, it can optimize resource allocation to avoid overload or waste, thereby improving the overall performance and stability of the cluster. Relying on historical data and prediction models, it ensures the scientificity and efficiency of resource management decisions and maximizes the resource utilization of the cluster.

[0064] See also Figure 3The third embodiment of the optimization method for database cluster management in the embodiment of the present invention includes:

[0065] 301. Generate a node task requirement table according to resource requirements and cluster size, and configure the priority of the node task requirement table according to the urgency and importance of the task;

[0066] In this embodiment, a node task requirement table is generated based on resource requirements and cluster size. The node task requirement table is a detailed resource allocation plan that lists the resources required for each task in the cluster, including computing resources, storage resources, memory resources, network bandwidth, etc. The priority of the node task requirement table is configured according to the urgency and importance of the task.

[0067] 302. Retrieve cluster resources corresponding to the node task requirement table based on the priority;

[0068] In this embodiment, the relationship between priority and the node task requirement table is understood, and the cluster resources required for each task are retrieved according to the node task requirement table. The priority of the task is divided into three levels: high, medium, and low. The combination of task priority and resource requirement table determines the resource allocation method. For high-priority tasks, the scheduling system will give priority to allocating computing resources, memory, etc. to them, while for low-priority tasks, the system may allocate them when resources are idle, or delay execution when the cluster load is low.

[0069] 303. Use a resource scheduling algorithm to allocate cluster resources to each cluster node;

[0070] In this embodiment, tasks are sorted according to their priorities to ensure that tasks with high priorities are allocated resources first. In the queue, tasks are arranged according to their priorities, with tasks with high priorities at the front of the queue and tasks with low priorities at the back. Once the task sorting is completed, the scheduling system will select a suitable node to allocate resources to the task based on the resource conditions of each node in the cluster. For example, if a high-priority computing task requires a large amount of CPU resources, and the CPU resources of a node are idle, the system will schedule the task to be executed on this node.

[0071] 304. Detect the node health status of each cluster node in real time;

[0072] In this embodiment, the online status of the node is confirmed by periodically sending heartbeat signals, and by analyzing the log files of the node to check whether there are abnormal errors or warning messages, it can be found whether there are potential problems with the node.

[0073] 305. Determine whether the node health state is an unhealthy state;

[0074] In this embodiment, the definition of an unhealthy state is clearly defined. For example, hardware resources such as the CPU, memory, and disk fail or have abnormal usage rates; key services on the node (such as databases, Web servers, etc.) fail to start, respond, or crash; the node's network connection is unstable and cannot communicate normally with other nodes in the cluster; the CPU or memory usage is in a high-load state for a long time, resulting in resource exhaustion; or the load of a node is much higher than that of other nodes. The above indicators are used to determine whether the node health state is an unhealthy state.

[0075] 306. If yes, back up the data of the cluster nodes corresponding to the unhealthy state and perform troubleshooting on the cluster nodes;

[0076] In this embodiment, if the node health status is unhealthy, the data of the cluster node corresponding to the unhealthy status is backed up, and the cluster node is troubleshooted. During the troubleshooting process, the faulty node can be automatically eliminated, the node can be automatically restored, or the cluster can be rebuilt or migrated.

[0077] 307. After troubleshooting, restart the cluster node and obtain the troubleshooting log;

[0078] In this embodiment, after troubleshooting, the health status of the node is confirmed before the restart operation is performed, and then the cluster node is restarted. If it is a physical server, the operating system restart command can be used or restarted through the hardware console. If it is a virtual machine or containerized environment, the node can be restarted through the virtualization management platform. After restarting the node, log information related to troubleshooting and restarting needs to be collected.

[0079] 308. Generate optimization suggestion strategies based on the fault type and severity in the troubleshooting log;

[0080] In this embodiment, the logs generated during the troubleshooting process are comprehensively collected and reviewed. By analyzing these logs, the specific type of fault, the root cause, and the scope of impact can be extracted. The severity of each fault is evaluated based on the scope of impact and the difficulty of recovery. Based on the type and severity of the fault, a specific optimization recommendation strategy is generated.

[0081] 309. Evaluate and verify the optimization suggestion strategy to obtain evaluation and verification results;

[0082] In this embodiment, it is evaluated whether the optimization strategy has successfully solved the problems that have occurred previously, for example, whether problems such as hardware failure, resource overload, and service crash are alleviated or completely eliminated, and whether the optimization strategy has improved the performance of the system, including response time, throughput, resource utilization, etc. It is evaluated whether the optimized system has improved availability and stability and reduced the frequency or impact range of failures. The key performance indicators of the system before and after optimization are obtained through monitoring tools, and the failure records of the optimized system are tracked for a period of time to check whether the number of failures has been reduced or the time for failure handling has been shortened. The performance and responsiveness of the system are tested under high load conditions to verify whether the optimization measures can maintain the stability of the system under pressure, and to ensure that the optimization strategy does not introduce new problems, especially the impact on other functions of the system. Some failures or pressures (such as network interruptions, hardware failures, etc.) are deliberately introduced to test the performance of the system in these situations to ensure that the redundancy and fault-tolerant mechanisms are effective.

[0083] 310. When the evaluation verification result is passed, the visualization tool (such as Tableau, Power BI, Grafana or other data visualization platforms) will generate a visualization strategy report based on the optimization recommendation strategy, including various charts (such as line charts, bar charts, pie charts, heat maps, etc.) to display the optimization results, create an interactive dashboard for the report, allow the management end to view detailed data by selecting different time periods, indicator categories and other parameters, use report generation tools (such as Google Docs, Microsoft Word, LaTeX, Markdown, etc.) or dedicated reporting tools, write and format the report according to the standard template, and send the visualization strategy report to the management end via email or web platform;

[0084] In this embodiment, when the evaluation verification result is passed, the optimization suggestion strategy is generated into a visualization strategy report, and the visualization strategy report is sent to the management end.

[0085] In the embodiment of the present invention, cluster resources are accurately allocated through priority configuration and resource scheduling algorithm, resource utilization and task execution efficiency are improved, node health status is detected in real time, faulty nodes are quickly identified and processed, system stability is ensured and downtime is reduced, cluster performance is continuously optimized through log analysis and optimization strategy evaluation after troubleshooting, continuous improvement of the system is promoted, and visualization reports are generated for optimization strategies to provide clear decision support and help managers make subsequent adjustments and management.

[0086] See also Figure 4 The fourth embodiment of the optimization method for database cluster management in the embodiment of the present invention includes:

[0087] 401. Obtain the policy application result fed back by the management end, and obtain the feedback timestamp from the timestamp service platform based on the feedback time;

[0088] In this embodiment, the management personnel will provide an evaluation on the effectiveness of the optimization strategy execution, whether the expected goals have been achieved, for example, whether the system performance has been improved, whether the failures have been reduced, whether the user experience has been improved, etc. The feedback content includes text descriptions, ratings or questionnaire results. The feedback from the management side is usually collected through e-mail, internal reports, team collaboration tools or through questionnaires, and then the strategy application results fed back by the management side are obtained. Once the feedback is collected from the management side, the "feedback time" information is passed to the timestamp service platform through the API or data interface to obtain the corresponding timestamp.

[0089] 402. Merge the feedback timestamp, the optimization suggestion strategy, and the strategy application result to obtain merged information;

[0090] In this embodiment, each feedback timestamp is associated with its related optimization suggestion strategy and strategy application result to create a unified data structure, which includes a feedback timestamp field, an optimization suggestion strategy field and a strategy application result field. Unique identifiers (such as feedback ID, strategy ID) are used to connect different information together. The three items of information, feedback timestamp, optimization suggestion strategy and strategy application result, need to be associated through these identifiers to ensure that each feedback can follow up the corresponding strategy suggestion and application result, fill the collected data into a unified data structure, and perform data verification.

[0091] 403. Encrypt the combined information to obtain encrypted information;

[0092] In this embodiment, an asymmetric encryption algorithm is selected, and a pair of keys, namely a public key and a private key, is used. The public key is used to encrypt the combined information to obtain the encrypted information, and the private key is used to decrypt it.

[0093] 404. Upload the encrypted information to the blockchain;

[0094] In this embodiment, a blockchain platform is selected, and before uploading the encrypted information to the blockchain, the encrypted information is converted into a format suitable for storage in the blockchain, a hash value of a fixed length is generated from the encrypted data through a hash algorithm (such as SHA-256), the encrypted information and related metadata are packaged into a transaction, and prepared to be submitted to the blockchain, and a blockchain transaction containing the encrypted information is created. The transaction includes a summary or hash value of the encrypted information, related metadata (such as the creation time of the encrypted information, the source of the data, the type of encryption algorithm, etc.), and the identity information of the sender (usually an account address on the blockchain), and the transaction is signed using the sender's private key to ensure the integrity and identity of the transaction, and the signed transaction is broadcast to the blockchain network. When the transaction is added to the block, it will be confirmed by multiple nodes and eventually become part of the blockchain. Once the encrypted information is uploaded to the blockchain, it will be permanently recorded in the block, and anyone can view the data in the block, but only the person holding the private key can decrypt the data content.

[0095] In the embodiment of the present invention, encryption measures are taken to ensure that policy feedback and related data are protected during transmission to prevent information leakage and tampering. Blockchain technology is used to upload data to ensure that the information is tamper-proof and open and transparent, and to enhance trust in and audit capabilities for policy execution. By combining timestamps and blockchain technology, a clear record of the policy application process is provided to facilitate later tracing and verification, ensuring the integrity of historical data. The efficiency of the policy execution process is improved through automated merging, encryption, and uploading processes, ensuring that policy execution and feedback can be carried out efficiently and reliably.

[0096] The above describes the optimization method for database cluster management in the embodiment of the present invention. The following describes the optimization device for database cluster management in the embodiment of the present invention. Figure 5 In one embodiment of the present invention, an optimization device for database cluster management includes:

[0097] An identification and collection module 501 is used to identify each cluster node in the database cluster and collect performance data of each cluster node;

[0098] Determine monitoring module 502, used to determine resource requirements based on performance data and monitor resource usage in each cluster node;

[0099] The analysis and adjustment module 503 is used to analyze the resource usage, obtain the load change analysis result, and adjust the cluster size of each cluster node according to the load change in the load change analysis result;

[0100] Allocation detection module 504, used to dynamically allocate cluster resources according to resource requirements and cluster size, and detect the node health status of each cluster node in real time;

[0101] The generating and sending module 505 is used to generate an optimization suggestion strategy according to the node health status and send the optimization suggestion strategy to the management end.

[0102] In this embodiment, by identifying each cluster node in the database cluster, collecting performance data of each cluster node, determining resource requirements based on the performance data, monitoring resource usage in each cluster node, analyzing resource usage, adjusting the cluster scale of each cluster node based on load changes in the load change analysis results, dynamically allocating cluster resources based on resource requirements and cluster scale, and realizing automation and optimization of cluster management, thereby improving the overall operation efficiency of the system, and detecting the node health status of each cluster node in real time after resource allocation, generating optimization recommendation strategies based on the node health status, and sending the optimization recommendation strategies to the management end so as to adjust the cluster management strategies in time, thereby improving the overall operation stability of the system.

[0103] See also Figure 6 Another embodiment of the optimization device for database cluster management in the embodiment of the present invention includes:

[0104] An identification and collection module 501 is used to identify each cluster node in the database cluster and collect performance data of each cluster node;

[0105] Determine monitoring module 502, used to determine resource requirements based on performance data and monitor resource usage in each cluster node;

[0106] The analysis and adjustment module 503 is used to analyze the resource usage, obtain the load change analysis result, and adjust the cluster size of each cluster node according to the load change in the load change analysis result;

[0107] Allocation detection module 504, used to dynamically allocate cluster resources according to resource requirements and cluster size, and detect the node health status of each cluster node in real time;

[0108] A generating and sending module 505 is used to generate an optimization suggestion strategy according to the node health status and send the optimization suggestion strategy to the management end;

[0109] In this embodiment, the identification and collection module 501 includes: a calling extraction unit 5011, which is used to call the cluster configuration file of the database cluster and extract all node IDs listed in the cluster configuration file; an identification unit 5012, which is used to identify all corresponding cluster nodes in the database cluster according to the node ID; a first collection unit 5013, which is used to collect performance data of each cluster node, and the performance data includes CPU usage, memory usage, disk I / O and network bandwidth.

[0110] In this embodiment, it is determined that the monitoring module 502 includes: a calculation unit 5021, which is used to calculate the resource utilization of the cluster nodes based on the performance data; a determination unit 5022, which is used to determine the resource demand according to the resource utilization; and a second collection unit 5023, which is used to use the Prometheus monitoring tool to regularly collect the resource usage in each cluster node.

[0111] In this embodiment, the analysis and adjustment module 503 includes: an acquisition and cleaning unit 5031, which is used to acquire historical resource usage data and perform data cleaning on the historical resource usage data; a construction unit 5032, which is used to construct a resource usage trend prediction model based on the historical resource usage data; an analysis unit 5033, which is used to use the resource usage trend prediction model to analyze the resource usage and obtain a load change analysis result; an adjustment unit 5034, which is used to adjust the cluster size of each cluster node according to the load change in the load change analysis result.

[0112] In this embodiment, the allocation detection module 504 includes: a generation configuration unit 5041, which is used to generate a node task requirement table according to resource requirements and cluster size, and configure the priority of the node task requirement table according to the urgency and importance of the task; a calling unit 5042, which is used to call the cluster resources corresponding to the node task requirement table based on the priority; an allocation unit 5043, which is used to use a resource scheduling algorithm to allocate cluster resources to each cluster node; and a detection unit 5044, which is used to detect the node health status of each cluster node in real time.

[0113] In this embodiment, the generation and sending module 505 includes: a judgment unit 5051, which is used to judge whether the health status of the node is an unhealthy state; a backup and troubleshooting unit 5052, which is used to back up the data of the cluster node corresponding to the unhealthy state if so, and perform troubleshooting on the cluster node; a restart acquisition unit 5053, which is used to restart the cluster node after troubleshooting and obtain the troubleshooting log; a generation unit 5054, which is used to generate an optimization recommendation strategy according to the fault type and severity in the troubleshooting log; an evaluation and verification unit 5055, which is used to evaluate and verify the optimization recommendation strategy and obtain an evaluation and verification result; a generation and sending unit 5056, which is used to generate a visualization strategy report for the optimization recommendation strategy when the evaluation and verification result is passed, and send the visualization strategy report to the management end.

[0114] In this embodiment, it also includes: an acquisition module 506, which is used to obtain the policy application results fed back by the management end, and obtain the feedback timestamp from the timestamp service platform based on the feedback time; a merging module 507, which is used to merge the feedback timestamp, the optimization suggestion policy and the policy application results to obtain merged information; an encryption module 508, which is used to encrypt the merged information to obtain encrypted information; and an uploading module 509, which is used to upload the encrypted information to the blockchain.

[0115] above Figure 5 and Figure 6 The optimization device for database cluster management in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The optimization device for database cluster management in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0116] Figure 7 6 is a schematic diagram of the structure of an optimization device for database cluster management provided by an embodiment of the present invention. The optimization device 600 for database cluster management may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 610 (for example, one or more processors) and a memory 620, and one or more storage media 630 (for example, one or more mass storage devices) storing application programs 633 or data 632. Among them, the memory 620 and the storage medium 630 may be temporary storage or permanent storage. The program stored in the storage medium 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the optimization device 600 for database cluster management. Furthermore, the processor 610 may be configured to communicate with the storage medium 630, and execute a series of instruction operations in the storage medium 630 on the optimization device 600 for database cluster management to implement the steps of the optimization method for database cluster management provided by the above-mentioned various method embodiments.

[0117] The optimization device 600 for database cluster management may also include one or more power supplies 640, one or more wired or wireless network interfaces 650, one or more input and output interfaces 660, and / or one or more operating systems 631, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 7 The illustrated structure of the optimization device for database cluster management does not constitute a limitation on the optimization device based on database cluster management, and may include more or less components than those illustrated, or a combination of certain components, or a different arrangement of components.

[0118] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the optimization method for database cluster management.

[0119] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0120] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0121] Finally, it should be noted that the above description is only a preferred example of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing database cluster management, characterized in that: include: Identify each cluster node in the database cluster and collect performance data of each cluster node; Determine resource requirements based on the performance data, and monitor resource usage in each of the cluster nodes; Analyze the resource usage to obtain a load change analysis result, and adjust the cluster size of each cluster node according to the load change in the load change analysis result; Dynamically allocate cluster resources according to the resource demand and the cluster size, and detect the node health status of each cluster node in real time; Generate an optimization suggestion strategy according to the node health status, and send the optimization suggestion strategy to the management end.

2. The optimization method for database cluster management according to claim 1, characterized in that: The identifying each cluster node in the database cluster and collecting performance data of each cluster node includes: Calling a cluster configuration file of the database cluster, and extracting all node IDs listed in the cluster configuration file; Identify all corresponding cluster nodes in the database cluster according to the node ID; The performance data of each of the cluster nodes is collected, wherein the performance data includes CPU usage, memory usage, disk I / O, and network bandwidth.

3. The optimization method for database cluster management according to claim 1, characterized in that: Determining resource requirements according to the performance data and monitoring resource usage in each of the cluster nodes includes: Based on the performance data, calculating resource utilization of the cluster nodes; determining resource demand based on the resource utilization; Use the Prometheus monitoring tool to regularly collect resource usage in each cluster node.

4. The optimization method for database cluster management according to claim 1, characterized in that: The analyzing the resource usage to obtain a load change analysis result, and adjusting the cluster size of each cluster node according to the load change in the load change analysis result, includes: Acquire historical resource usage data, and perform data cleaning on the historical resource usage data; constructing a resource usage trend prediction model based on the historical resource usage data; Analyzing the resource usage using the resource usage trend prediction model to obtain a load change analysis result; The cluster size of each of the cluster nodes is adjusted according to the load change in the load change analysis result.

5. The optimization method for database cluster management according to claim 1, characterized in that: The dynamically allocating cluster resources according to the resource demand and the cluster size, and detecting the node health status of each cluster node in real time, includes: Generate a node task requirement table according to the resource requirements and the cluster size, and configure the priority of the node task requirement table according to the urgency and importance of the task; Based on the priority, cluster resources corresponding to the node task requirement table are retrieved; Allocate the cluster resources to each of the cluster nodes using a resource scheduling algorithm; The node health status of each of the cluster nodes is detected in real time.

6. The optimization method for database cluster management according to claim 1, characterized in that: Generating an optimization suggestion strategy according to the node health status and sending the optimization suggestion strategy to a management terminal includes: Determine whether the node health state is an unhealthy state; If so, backing up the data of the cluster node corresponding to the unhealthy state and performing troubleshooting on the cluster node; After troubleshooting, restart the cluster node and obtain troubleshooting logs; Generate an optimization suggestion strategy based on the fault type and severity in the troubleshooting log; Evaluate and verify the optimization suggestion strategy to obtain evaluation and verification results; When the evaluation verification result is passed, the optimization suggestion strategy is used to generate a visualization strategy report, and the visualization strategy report is sent to the management end.

7. The optimization method for database cluster management according to claim 1, characterized in that: After generating the optimization suggestion strategy according to the node health status and sending the optimization suggestion strategy to the management end, the method further includes: Obtaining the policy application result fed back by the management end, and obtaining a feedback timestamp from a timestamp service platform based on the feedback time; Merging the feedback timestamp, the optimization suggestion strategy and the strategy application result to obtain merged information; Encrypting the combined information to obtain encrypted information; The encrypted information is uploaded to the blockchain.

8. An optimization device for database cluster management, characterized in that: include: An identification and collection module, used to identify each cluster node in the database cluster and collect performance data of each cluster node; Determine a monitoring module, which is used to determine resource requirements based on the performance data and monitor resource usage in each of the cluster nodes; An analysis and adjustment module, configured to analyze the resource usage to obtain a load change analysis result, and adjust the cluster size of each cluster node according to the load change in the load change analysis result; An allocation detection module, used to dynamically allocate cluster resources according to the resource demand and the cluster size, and detect the node health status of each cluster node in real time; A generating and sending module is used to generate an optimization suggestion strategy according to the health status of the node and send the optimization suggestion strategy to the management end.

9. An optimization device for database cluster management, characterized in that: The optimization device for database cluster management includes: a memory and at least one processor, wherein instructions are stored in the memory; At least one of the processors calls the instructions in the memory to enable the database cluster management optimization device to execute each step of the database cluster management optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, the various steps of the optimization method for database cluster management as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Content delivery network scheduling method, system and equipment

    CN120321201A

  • Node management method and device, equipment, storage medium and program product

    CN120729865A