Database node management method and device, electronic equipment and storage medium

By predicting the load change trend of database nodes and adding alternate read-only nodes, the problem of delay and expansion of database nodes in the prior art is solved, and the performance and scalability of database clusters are improved.

CN119938294APending Publication Date: 2025-05-06ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311468002.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing database node management solution has delays, and it takes a long time to expand the read-only node operation, making it difficult to deal with the sudden increase in database cluster load in a timely manner, resulting in performance problems.

Method used

By obtaining the load timing data of each database node, predicting future load change trends and determining whether to add alternate read-only nodes to optimize the load management of the database cluster.

Benefits of technology

It greatly shortens the node expansion operation time, improves the concurrency, reliability and scalability of the database cluster, and promptly deals with performance losses at high loads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938294A_ABST
    Figure CN119938294A_ABST
Patent Text Reader

Abstract

The invention provides a database node management method and apparatus, an electronic device and a storage medium. The database node management method comprises the steps of obtaining load time sequence data of each database node in a first preset duration before a current moment and closest to the current moment; based on the load time sequence data of each database node, respectively predicting load change trend information of each database node in a second preset duration in the future; and based on the load change trend information of each database node, determining whether to add a standby read-only node of a cluster to which each database node belongs. According to the method and the device, the time consumed by node expansion operation can be greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of database technology, and specifically relates to a database node management method, device, electronic device and storage medium. Background Art

[0002] With the continuous development of cloud computing technology, various cloud applications have emerged. Among them, cloud database is a relatively mature and widely used cloud application technology. In recent years, Serverless (Serverless computing) database, as a cloud database solution based on serverless architecture, has been known and applied by more and more people. Serverless is a cloud computing model based on Platform as a Service (PaaS) and provides a micro architecture. End customers do not need to deploy, configure or manage server services. The server services required for code operation are all provided by the cloud platform.

[0003] The Serverless database can automatically allocate and expand resources on demand to adapt to the application load. This feature can help users reduce the waste of resources caused by over-reserving server resources, or the cost waste and time delay caused by having to perform secondary expansion due to under-reserving server resources. Therefore, it is necessary to reasonably manage and control the database nodes of the Serverless database.

[0004] In the related art, database node management solutions usually have a certain delay, and the operation of expanding read-only nodes often takes minutes. Therefore, the existing methods are difficult to deal with the sudden increase of database cluster load in time, which will cause certain performance problems. Summary of the invention

[0005] The present application proposes a database node management method, device, electronic device and storage medium, which can significantly shorten the time consumed by node expansion operations.

[0006] The first embodiment of the present application proposes a database node management method, including:

[0007] Obtaining load time series data of each database node before the current moment and within a first preset time period closest to the current moment;

[0008] Based on the load time series data of each database node, respectively predict the load change trend information of each database node within a second preset time period in the future;

[0009] Based on the load change trend information of each database node, determine whether to add a standby read-only node to the cluster to which each database node belongs.

[0010] In some embodiments of the present application, the step of obtaining the load time series data of each database node within a first preset time period before the current moment and closest to the current moment includes:

[0011] Obtaining performance indicator data of each database node at each moment within a first preset time period before the current moment and closest to the current moment;

[0012] Based on the performance indicator data of each database node, the load time series data of each database node is determined.

[0013] In some embodiments of the present application, the load change trend information of each database node within a second preset time period in the future is predicted based on the load time series data of each database node, including:

[0014] Perform noise smoothing on the load time series data of each database node to obtain new load time series data of each database node;

[0015] Based on the new load time series data of each database node, the load change trend information of each database node within a second preset time period in the future is predicted respectively.

[0016] In some embodiments of the present application, the noise smoothing of the load time series data of each database node to obtain new load time series data of each database node includes:

[0017] Determine the neighboring points of each data point in the load time series data of each database node;

[0018] Based on the weights of each neighboring point of the data point, the smoothing value of the corresponding data point is calculated;

[0019] Based on the smoothed value of each data point in each database node, new load time series data corresponding to each database node is formed.

[0020] In some embodiments of the present application, before calculating the smoothing value of the corresponding data point based on the weights of each neighbor point of the data point, the method further includes:

[0021] Determine the time windows corresponding to each data point in the load time series data of each database node;

[0022] The weight of each neighbor point is determined according to the distance between the neighbor point of each data point in the corresponding time window and the corresponding data point.

[0023] In some embodiments of the present application, the load change trend information of each database node within a second preset time period in the future is predicted based on the new load time series data of each database node, including:

[0024] Predicting the change trend information of the new load time series data of each database node within the second preset time period based on the autoregressive model;

[0025] Based on the change trend information of the new load time series data, the load change trend information of the corresponding database node within the second preset time length is determined.

[0026] In some embodiments of the present application, the determining whether to add a standby read-only node to the cluster to which each database node belongs based on the load change trend information of each database node includes:

[0027] Based on the load change trend information of each database node, respectively determine the predicted load level of each database node after the second preset time period;

[0028] Filter out high-load nodes whose predicted load levels reach a first preset threshold from each database node;

[0029] Based on the high-load node, it is determined whether to add a standby read-only node to the cluster to which the high-load node belongs.

[0030] In some embodiments of the present application, determining whether to add a standby read-only node to the cluster to which the high-load node belongs based on the high-load node includes:

[0031] Based on the high-load nodes, determining the number of high-load nodes belonging to the same cluster;

[0032] Based on the number of high-load nodes belonging to the same cluster, it is determined whether to increase a standby read-only node of the cluster.

[0033] In some embodiments of the present application, the determining whether to add a standby read-only node of the cluster based on the number of high-load nodes belonging to the same cluster includes:

[0034] When the number of high-load nodes belonging to the same cluster is greater than or equal to a preset number, determining whether the total number of database nodes in the cluster reaches a preset upper limit threshold, and whether there are standby read-only nodes in the cluster;

[0035] When the total number of database nodes in the cluster does not reach the preset upper limit threshold and there is no standby read-only node in the cluster, it is determined to increase the standby read-only node of the cluster.

[0036] In some embodiments of the present application, after determining whether to add a standby read-only node to the cluster to which the high-load node belongs, the method further includes:

[0037] Determining the total load level of the cluster and the activation of standby read-only nodes in the cluster;

[0038] When the total load level of the cluster and the activation status of the standby read-only node meet preset conditions, determine to delete the standby read-only node; the preset conditions include: the total load in the cluster does not reach a second preset threshold within a third preset time period, and the standby read-only node is never activated.

[0039] The second aspect of the present application provides a database node management device, including:

[0040] A data acquisition module, used to acquire the load time series data of each database node before the current moment and within a first preset time period closest to the current moment;

[0041] A load prediction module, used to predict the load change trend information of each database node within a second preset time period in the future based on the load time series data of each database node;

[0042] The node management module is used to determine whether to add a standby read-only node to the cluster to which each database node belongs based on the load change trend information of each database node.

[0043] The third aspect of the present application provides a database node management system, including a computing center end, an application service end and a node control end, wherein the computing center end is used to execute the database node management method described in the first aspect;

[0044] The application server is used to monitor the load level of the database node and send an instruction to add a node to the node control terminal when the load level is greater than a preset first threshold;

[0045] The node control end adds a spare read-only node to the cluster when the computing center determines to add a spare read-only node; and brings the spare read-only node online when receiving the instruction to add the node.

[0046] An embodiment of the fourth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.

[0047] An embodiment of the fifth aspect of the present application provides a computer-readable storage medium on which a computer program is stored, and the program is executed by a processor to implement the method described in the first aspect above.

[0048] The technical solution provided in the embodiments of the present application has at least the following technical effects or advantages:

[0049] The database node management method provided in the embodiment of the present application can predict the load change trend information of each database node in the future period based on the load time series data of each database node in the recent period, and determine whether to increase the backup read-only node of the cluster to which each database node belongs based on the load change trend information. In this way, the embodiment of the present application is based on the load time series data of each database node in the recent period of time. According to the correspondence between load and time in the load time series data, a mathematical statistical method can be used to predict the load change trend information of each database node in the future period of time. The load change trend information of each database node can be used to predict the load level of each database node in the future period of time, that is, it can be predicted in time before the database node reaches a high load, so that the backup read-only node of the cluster to which the database node belongs can be added in advance, so that when the load of the database node where the client is located reaches a high load, the added backup read-only node can be directly put online and expanded for use, thereby reducing the node expansion operation at the minute level to the second level, greatly shortening the expansion operation time of the database node, and promptly responding to the performance loss problem when the database node is under high load, which can improve the concurrency, reliability, scalability and overall performance of the entire database cluster, and at the same time can achieve more accurate database node expansion operations to avoid waste of resources and costs.

[0050] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] By reading the detailed description of the preferred embodiment below, various other advantages and benefits will become clear to those of ordinary skill in the art. The accompanying drawings are only used for the purpose of illustrating the preferred embodiment and are not considered to be limitations of the present application. In addition, the same reference symbols are used to represent the same components throughout the accompanying drawings.

[0052] In the attached picture:

[0053] Figure 1 An application system and principle diagram of a database node management method provided by some embodiments of the present application are shown;

[0054] Figure 2 A schematic diagram showing a flow chart of a database node management method provided in some embodiments of the present application;

[0055] Figure 3 A specific flow chart of step S2 in some embodiments of the present application is shown;

[0056] Figure 4 A specific flow chart of step S21 in some embodiments of the present application is shown;

[0057] Figure 5 A specific flow chart of step S22 in some embodiments of the present application is shown;

[0058] Figure 6 A specific flow chart of step S3 in some embodiments of the present application is shown;

[0059] Figure 7 A specific flow chart of step S33 in some embodiments of the present application is shown;

[0060] Figure 8 A schematic diagram showing the structure and working principle of a database node management system provided in some embodiments of the present application is shown;

[0061] Fig. 9 A schematic diagram showing the structure of a database node management device provided in some embodiments of the present application is shown;

[0062] Fig.10 A schematic diagram of an electronic device provided by an embodiment of the present application is shown;

[0063] Fig.11 A schematic diagram of a storage medium provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0064] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0065] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by technicians in the field to which this application belongs.

[0066] In traditional database node management solutions, it is usually necessary to manually adjust the configuration of database resources (such as CPU, memory, disk capacity, read-only nodes in the cluster, etc.) to adapt to the database load. This method requires manual intervention, is costly and inefficient. The elastic scaling of database Serverless is the ability to automatically increase or decrease database resources according to the application load. When the application load increases, the database can automatically allocate more resources to meet the demand, and when the load decreases, the database can automatically release excess resources. Therefore, database Serverless can reduce the waste of resources caused by too many reserved server resources, or the cost waste and time delay caused by having to perform secondary expansion due to too few reserved server resources. However, the database nodes of database Serverless need to be reasonably managed and controlled to better realize the elastic scaling function of database Serverless.

[0067] The elastic scaling of database read-only nodes can automatically increase or decrease the number of read-only nodes according to the cluster load to meet the reading requirements under different loads, thereby improving the performance, reliability and scalability of the entire database system, while avoiding resource waste and cost waste. In related technologies, the elastic scaling of database read-only nodes can usually adaptively adjust the number of nodes and resource configuration based on predefined rules or threshold algorithms, such as adjusting the number of read-only nodes and resource configuration based on indicators such as CPU usage or number of connections to adapt to different load requirements. However, this type of method based on the current load for rule or threshold matching usually has a certain delay, and the operation of expanding read-only nodes takes minutes, making it difficult for existing database node expansion solutions to respond to sudden increases in database cluster load in a timely manner, which will cause certain performance problems.

[0068] In order to solve the above-mentioned problems, the embodiments of the present application have conducted a lot of research on the existing database management methods, and found that the elastic scaling technology of the database is mainly divided into two categories: rule-based elastic scaling technology and automatic learning-based elastic scaling technology. Among them, the rule-based elastic scaling technology automatically adjusts the configuration of database resources through pre-defined rules or thresholds. For example, it automatically allocates and releases the computing and storage resources of the database instance according to indicators such as CPU usage or I / O operation number. Most of the existing technologies also use this solution to achieve elastic scaling of the database, but this technology is difficult to adapt to different load conditions, and there may be certain delays and performance issues.

[0069] The elastic scaling technology based on automatic learning automatically learns the database resource usage and load pattern, predicts the future load situation, and automatically adjusts the configuration of database resources. However, the implementation of this technology requires a large amount of historical data and feedback information to train the database's automatic scaling model, and the model needs to be updated regularly to maintain accuracy. In this way, not only is the training cost very high, but the design, development and maintenance of the model also require a considerable amount of resources, technology, manpower and time, resulting in a very high implementation cost. Therefore, it has basically not been effectively implemented.

[0070] In view of the above problems, an embodiment of the present application provides a database node management method, which can predict the load change trend information of each database node in the future period based on the load time series data of each database node in the recent period, and determine whether to add a backup read-only node to the cluster to which each database node belongs based on the load change trend information. In this way, the embodiment of the present application is based on the load time series data of each database node in the recent period of time. According to the correspondence between load and time in the load time series data, a mathematical statistical method can be used to predict the load change trend information of each database node in the future period of time. The load change trend information of each database node can be used to predict the load level of each database node in the future period of time, that is, it can be predicted in time before the database node reaches a high load, so that the backup read-only node of the cluster to which the database node belongs can be added in advance, so that when the load of the database node where the client is located reaches a high load, the added backup read-only node can be directly put online and expanded for use, thereby reducing the node expansion operation at the minute level to the second level, greatly shortening the expansion operation time of the database node, and promptly responding to the performance loss problem when the database node is under high load, which can improve the concurrency, reliability, scalability and overall performance of the entire database cluster, and at the same time can achieve more accurate database node expansion operations to avoid waste of resources and costs.

[0071] Cluster load prediction can be achieved through distributed stream processing and batch processing open source platforms, such as Flink. By streaming and reading second-level performance indicator data, cluster load conditions can be analyzed in real time, and the load of database nodes that are on an upward trend can be located in real time to predict whether they are about to break through the high level line. Low-latency load prediction can be achieved to ensure the stability and efficiency of the entire system.

[0072] Specifically, this embodiment can be applied to Figure 1 The database node management system shown in Figure 1As shown, the database node management system may include a computing center end, an application server end and a node control end provided by a cloud service provider, wherein the application server end is connected to the client of the application cloud database, and can collect performance indicator data of the database node where the client is located in real time, and judge whether the load of the database node has reached a high level based on the collected performance indicator data, such as CPU usage or number of connections, etc., and send an instruction to go online to the node to the node control end when the load reaches a high level.

[0073] The application server will also report the collected performance indicator data to the computing center in real time. The computing center can store the received performance indicator data in real time, for example, in a log (the log can store the received data and the time when the data was received). Then the computing center can execute the database node management method provided in this embodiment, obtain the performance indicator data of each database node from the log in real time, and form the load time series data of each database node in the recent period of time, and then predict the load change trend information of each database node in the future period of time based on the obtained load time series data, and then determine whether to add a standby read-only node to the cluster to which each database node belongs based on the load change trend information. In the case of determining to add a standby read-only node, an instruction to add a standby read-only node is sent to the node control end. The computing center can also execute a recovery mechanism for standby read-only nodes, and send an instruction to delete standby read-only nodes to the node control end.

[0074] The node control end is connected to the computing center end and the application server end respectively, and is used to actually perform operations such as adding database nodes, deleting database nodes, and bringing standby read-only nodes online. After receiving the instruction to add standby read-only nodes sent by the computing center end, it adds standby read-only nodes to the cluster; after receiving the instruction to add nodes sent by the application server end, it brings standby read-only nodes online; after receiving the instruction to delete standby read-only nodes sent by the computing center end, it deletes standby read-only nodes in the cluster; and when there are no standby read-only nodes in the cluster, it adds read-only nodes.

[0075] It can be understood that the database node management method provided in this embodiment is applied to the above-mentioned database node management system, which is only one implementation method of this embodiment, and this embodiment is not limited to this. For example, the above-mentioned database node management system can also be integrated into a database node management instance for implementation, and the database node management method can also be applied to the database node management instance.

[0076] The following is a detailed description of the database node management method provided in the embodiment of the present application, taking the application in the above-mentioned computing center as an example and combining with the accompanying drawings.

[0077] See also Figure 2, is a flow chart of the database node management method provided in the embodiment of the present application, such as Figure 1 As shown, the database node management method may include the following steps:

[0078] Step S1, obtaining the load time series data of each database node within a first preset time period before the current moment and closest to the current moment.

[0079] Among them, the database node can be understood as a database instance. In a cluster environment, one or more database instances can run on a server. The load time series data of the database node can be understood as a data column in which the load of the database node is recorded in chronological order, that is, it includes the load of the database node at different times and the corresponding time. The first preset time length can be any time period that can accurately predict the future change trend of the database node, that is, by obtaining the load time series data in the time period before the current moment and closest to the current moment, the future change trend of the database node can be accurately predicted. The specific value can be obtained through a limited number of tests. This embodiment does not make specific limitations on it. For example, it can be a few minutes, more than ten minutes, dozens of minutes, etc.

[0080] In view of the real-time variability of database nodes, historical data closer to the current moment is of decisive significance for predicting future load change trends. Therefore, this embodiment can only obtain the load time series data of each database node before the current moment and within the first preset time period closest to the current moment. In this way, a small window is used to save the latest load data in a short period of time, and data that is too old and has no relevance is discarded. There is no need to conduct large-scale data learning, which can effectively reduce the amount of data processing and increase the speed of data processing on the basis of improving prediction accuracy.

[0081] Specifically, the above step S1 may include the following processing: obtaining the performance indicator data of each database node at each moment before the current moment and within the first preset time length closest to the current moment; determining the load time series data of each database node based on the performance indicator data of each database node.

[0082] In actual applications, after receiving the performance indicator data reported by the application server, the computing center can store the received performance indicator data in chronological order to quickly obtain the historical load data of the database node. Then the computing center can obtain the performance indicator data stored in the first preset time period before the current moment and closest to the current moment in real time, and form the load time series data of each database node in the recent period of time, so as to accurately predict the load change trend information of each database node in the future period of time based on the load time series data.

[0083] Step S2: Based on the load time series data of each database node, predict the load change trend information of each database node within a second preset time period in the future.

[0084] The second preset duration can be any duration, as long as the node control end can complete the operation of adding the standby read-only node within the duration, when the application server sends a node adding instruction after the second preset duration, the node control end can bring the added standby read-only node online in time, so as to shorten the node expansion operation time. The specific value of the second preset duration is similar to that of the first preset duration, and can also be obtained through a limited number of tests. This embodiment does not specifically limit it, for example, it can also be a few minutes, more than ten minutes, etc.

[0085] In some embodiments, Figure 3 As shown, step S2 may include the following processing: step S21, performing noise smoothing on the load time series data of each database node to obtain new load time series data of each database node; step S22, based on the new load time series data of each database node, predicting the load change trend information of each database node within a second preset time period in the future.

[0086] The new load time series data may be understood as data obtained after the original load time series data has been transformed by one or more smoothing algorithms.

[0087] In this embodiment, the noise of the load time series data of each database node is smoothed, and the smoothing algorithm can be used to perform multiple transformations on the original load time series data to extract the trend characteristics of the overall sequence and retain the detailed characteristics of each interval, so as to obtain new load time series data that can accurately reflect the load change trend of each database node. Then, the current load change trend information is determined based on the new load time series data to improve the accuracy of the current load change trend information prediction. Based on the current load change trend information, the load change trend information of the database node within the second preset time length in the future can be predicted.

[0088] Furthermore, if Figure 4 As shown, the above step S21 may include the following processing: step S211, determining the neighbor points of each data point in the load time series data of each database node; step S212, calculating the smoothing value of the corresponding data point based on the weight of each neighbor point of the data point; step S213, forming new load time series data corresponding to each database node based on the smoothing value of each data point in each database node.

[0089] The data points in the load time series data can be understood as the load value at a specific moment. The neighbor points of a data point can be understood as the load value in a period close to the corresponding moment of the data point. The weight of the neighbor point can be understood as the degree of influence of the neighbor points around a data point on the data point, and its value can be 0-1. The larger the value, the greater the influence on the data point.

[0090] This embodiment adopts the idea of ​​statistics and calculates the smoothing value of the corresponding data point based on the weight of each neighbor point of the data point. It does not need to make assumptions about the distribution of the data points and can be applied to more data types. And based on the weight of the neighbor point, the data point is locally smoothed, which retains the overall trend characteristics of the load time series data and the detailed characteristics of each data point, making the prediction results based on the new load time series data more accurate. In addition, the calculation time of the smoothing value is independent of the size of the data set, which is more suitable for real-time data processing and online analysis.

[0091] In some embodiments, before executing the above step S212, the time windows corresponding to each data point in the load time series data of each database node can be determined; then the weight of each neighbor point is determined according to the distance between the neighbor point of each data point in the corresponding time window and the corresponding data point. In this way, the local smoothing effect of the data point can be adjusted by adjusting the size of the time window and the weight of the neighbor point. On the basis of retaining the overall trend characteristics of the load time series data, the detailed characteristics of each data point can be better reflected. In this way, the data point can be locally smoothed by adjusting the size of the time window and the weight of the neighbor point, so that the detailed characteristics are more prominent. It is also possible to adapt to different data sets by adjusting the degree of smoothing and the number of neighbor points to further improve the versatility of the database node management method.

[0092] Furthermore, for each data point, a time window can be selected, which includes some data points around the data point as the neighbor points of the data point. The weight of each neighbor point is calculated according to the distance between the neighbor point and the data point. The closer the neighbor point is to the data point, the greater the weight. Then, for each data point, the least squares algorithm is used to perform weighting according to the weight of the neighbor point to obtain the smoothing value of the point.

[0093] It is understandable that this embodiment does not specifically limit the size of the time window. For example, it can be several seconds, tens of seconds, etc. The smaller the window, the more detailed features of the data will be retained. The larger the window, the more the overall features of the data will be highlighted, and the resulting curve will be smoother.

[0094] The specific calculation method of the weight is not specifically limited in this embodiment, as long as the calculated weight can accurately reflect the distance relationship between the neighbor point and the data point. For example, the weight calculation formula can be shown as the following formula (1):

[0095] w(x)=(1-|d| 3 ) 3 (1)

[0096] Among them, w(x) is the weight of the neighbor point x, and d is the distance between the neighbor point x and the data point.

[0097] Specifically, a non-parametric regression smoothing method, such as but not limited to the LOWESS sequence smoothing algorithm, can be used to smooth the noise of the load time series data of the database node, and curve fitting can also be performed on the smoothed new load time series data to better observe the trends and laws of the data.

[0098] In order to select a better smoothing algorithm, this embodiment uses a simple moving average algorithm, a central moving average algorithm, a weighted moving average algorithm, an exponentially weighted moving average algorithm, and a non-parametric regression smoothing algorithm to perform a smoothing test on the load time series data. The test shows that when the simple moving average algorithm is used to smooth the load time series data, the result has a lag and the smoothing effect is not very good; when the central moving average algorithm is used to smooth the load time series data, not only the smoothing effect is not good, but also there is no prediction value at the head and tail nodes, which may reduce the prediction accuracy; when the weighted moving average algorithm or the exponentially weighted moving average is used to smooth the load time series data, a good smoothing effect cannot be achieved. When the non-parametric regression smoothing algorithm is used to smooth the load time series data, a better smoothing effect can be achieved, and the processing time of the sequence change is short, there is basically no lag, and the prediction of future load change trend information is more accurate.

[0099] Among them, the LOWESS sequence smoothing algorithm is a non-parametric local weighted regression smoothing algorithm, which is used to perform noise smoothing and curve fitting on load time series data, so as to better observe the change trend and law of load time series data. The algorithm is based on the weighted least squares method, which converts it into a weighted regression problem by assigning a weight to each data point. Then, curve fitting can be performed based on the current load change trend information to obtain the current load change curve of the database node.

[0100] In other embodiments, Figure 5 As shown, the above step S22 may include the following processing: step 221, predicting the change trend information of the new load time series data of each database node within the second preset time length based on the autoregressive model; step 222, determining the load change trend information of the corresponding database node within the second preset time length based on the change trend information of the new load time series data.

[0101] The autoregressive model (AR model for short) is a statistical method for processing time series. It uses the data of the same variable, such as x, before time t, that is, x1 to x t-1 To predict the current time x t performance and assume that they have a linear relationship.

[0102] In this embodiment, during the normal service of the database node, the load change of the database node can be regarded as a continuous change, the load values ​​at similar times are highly correlated, and the autocorrelation coefficient of the load time series data is relatively large. Regression modeling is performed based on the relationship between the variable values ​​of the load in different periods of the sequence in the load time series data, so as to predict the future load value or load change trend. The change trend of each database node within the second preset time period in the future can be predicted relatively accurately, and then the load level of the database node after the second preset time period can be accurately predicted.

[0103] Furthermore, before the above step S221, the new load time series data can be differentiated to reduce the fluctuation of the new load time series data and reduce the dependence of the data on time, so that the new load time series data after differentiation is more stable. Then, the change trend of the time series data after differentiation is predicted through the autoregressive model to make the prediction result more accurate.

[0104] Furthermore, since the autoregressive model is based on the linear relationship between variables, the present embodiment can also verify the linear relationship, that is, the sequence stationarity of the load time series data, to further improve the accuracy of prediction based on the autoregressive model. If the current load change trend is increasing at a stable rate, the upcoming future moment can be predicted at the same rate, and when the predicted value reaches a preset threshold, the corresponding database node is determined to be a high-load node, and a database node high-load event can be output.

[0105] In addition, since autoregressive models are all based on stationary sequences, this embodiment can also perform a stationarity test before predicting the change trend information of the new load time series data of each database node within the second preset time period through the autoregressive model to ensure the accuracy of the autoregressive prediction.

[0106] In order to select a more appropriate sequence stationarity detection method, this embodiment studies the existing sequence stationarity detection methods, including graphic analysis, variance mean statistics, and multiple hypothesis testing methods. The research results show that graphic analysis is visual judgment and cannot be automatically detected. The variance mean statistics method divides the time series into multiple segments or multiple time windows, and observes whether the mean or variance of each time window is unchanged or in an approximate interval. In this way, the stability of the variance mean is transformed into a sequence stationarity problem again, resulting in poor detection effect and a large amount of calculation. The hypothesis testing methods, including DF test (Direction Finder), ADF test (Automatic Direction Finder), PP test (Projection Pursnit), DF-GLS test (Dickey-Fuller Test with GLS Detredding, test using generalized least squares to remove trend), KPSS test (Kwiatkowski-Phillips-Schmidt-Shin, python trend stability test), etc., can all be used to detect series stationarity, and these test methods can be used simultaneously or one or more to complement each other.

[0107] Among them, the ADF test is a method for detecting the linear stationarity of time series. It is based on the unit root test (Dickey-Fuller test) and adds a lag term to deal with the heteroscedasticity of the autocorrelation coefficient, thereby improving the accuracy and reliability of the test. The ADF test method is relatively simple and has a fast calculation speed, which is also suitable for the current time series linear stationarity test scenario.

[0108] In addition, this embodiment predicts the load of the database node, and the goal of the stability detection is to find a linearly rising sequence, not a linearly falling or linearly stable sequence. The method of distinguishing a linearly rising sequence from the other two sequences is that the first-order difference of the sequence is stable, and the sequence itself is not stable, or the sequence itself is linearly stable. In order to better achieve the detection purpose, this embodiment also conducts an in-depth study on the ADF test algorithm, and performs a variety of transformations on the autoregressive model of the ADF test algorithm in order to achieve a better detection effect. The research results show that: (1) without the α constant disturbance term and the βt time trend term, a stable sequence without a time trend can be identified, and the stability of the differential term can be identified for a linear sequence; (2) with the α constant disturbance term and without the βt time trend term, the recognition effect is not very good, and the number of false alarms is relatively high. (3) with the α constant disturbance term and the βt time trend term, with the time trend means that a linearly rising / declining stable sequence can be directly identified.

[0109] Based on the above research on the test algorithm, this embodiment can use the ADF test method to identify whether the new load time series data after noise smoothing is in a stable growth trend, that is, whether there is a linear relationship between the load values ​​of the database nodes at different times. Based on the above research results on the autoregressive model of the ADF test algorithm, a regression model without an α constant disturbance term and without a βt time trend term can be used to identify the stability of the differential term, and a regression model with an α constant disturbance term and a βt time trend term can be used to identify the stability of the load time series, so that the regression effect of the autoregressive model is better.

[0110] The null hypothesis of the ADF test is that the time series has a unit root, that is, non-stationarity. If the null hypothesis is rejected, the time series is considered to be stationary. The statistic of the ADF test is a t statistic, which is used to determine whether the null hypothesis is rejected by comparing the t value with the critical value. The specific test steps are as follows:

[0111] a) Differentiate the time series in the new load time series data to obtain a new time series. The differentiated time series is more stable and easy to predict the load change trend.

[0112] b) Establish a generalized autoregressive model to fit the time series after difference, and test whether the residual of the model has a unit root. Specifically, the autoregressive model can be shown as the following formula (2).

[0113] Δy t =α+βt+γy t-1 +δ1Δy t-1 +…+δ p-1 Δy t-p+1 +∈ t (2)

[0114] Among them, Δy t is the new load after differentiation corresponding to time t, y t-1 is the load corresponding to time (t-1), Δy t-1 is the new load after differentiation corresponding to time (t-1), δ1 is Δy t-1 The corresponding weight coefficient, Δy t-p+1 is the new load after differentiation corresponding to the time (t-p+1), δ p-1 is Δy t-p+1 The corresponding weight coefficient, ∈ t is a white noise sequence, α is a constant term, β is a time trend coefficient, p is the lag order of the autoregressive process, and the βt term represents a random walk with a drift term. When γ<0, it means that there is no unit root and the time series is stable. When γ=0, it means that there is at least one unit root and the sequence is non-stationary.

[0115] c) Calculate the t-value of the statistic and compare it with the critical value to determine whether the null hypothesis is rejected.

[0116] Specifically, the t value can be calculated according to the following formula (3), where SE is the standard error, is the predicted value of γ.

[0117]

[0118] It can be understood that the above-mentioned use of ADF to test the sequence stability of load time series data is only one implementation of this embodiment, and this embodiment is not limited to this. Those skilled in the art can also apply other hypothesis testing methods, such as DF test, PP test, DF-GLS test, KPSS test, etc., to perform sequence stability tests.

[0119] In addition, the premise of autoregression is that the sequence must be stationary, so a stationarity test is required. If the load time series of the database node does not meet the requirements of the stationarity test, it means that the current load is irregular. The load of the database node can be predicted without prediction. Load prediction is only performed when the load time series of the database node meets the requirements of the stationarity test.

[0120] Step S3: Based on the load change trend information of each database node, determine whether to add a standby read-only node to the cluster to which each database node belongs.

[0121] Among them, the standby read-only node can be called warm ro, which is a database standby read-only node that is invisible to users. When the database cluster reaches a high load, warm ro can be put online as a read-only node to bear read requests and reduce cluster pressure. Building warm ro read-only nodes can better cope with the performance loss caused by sudden high-load traffic in the cluster. In the scenario where the Serverless cluster reaches a high load, warm ro can be directly converted into a read-only node and officially put online, thereby reducing the minute-level performance loss of the existing extended read-only node to the second level, and then effectively improving the concurrency, reliability and overall performance of the database.

[0122] In some embodiments, Figure 6 As shown, the above step S3 may include the following processing: step S31, based on the load change trend information of each database node, respectively determine the predicted load level of each database node after the second preset time period; step S32, screen out high-load nodes whose predicted load level reaches the first preset threshold from each database node; step S33, based on the high-load nodes, determine whether to add a backup read-only node to the cluster to which the high-load nodes belong.

[0123] The first preset threshold may be a specific load value or a percentage of the total database node load. This embodiment does not specifically limit this and the specific value of the first preset threshold, which may be slightly less than the total database node load. If it is a percentage, it may be, for example, 70%, 75%, 80%, etc. A high-load node is a database node whose load value is greater than or equal to the first preset threshold.

[0124] It is understandable that the database nodes in the cluster include read-only nodes and read-write nodes, but each cluster does not necessarily include both read-only nodes and read-write nodes, and may only include read-only nodes, or only include read-write nodes, which is not specifically limited in this embodiment. Regardless of whether it is a read-only node or a read-write node, as long as the load level of the database node reaches the first preset threshold, it is a high-load node.

[0125] When determining whether to add warm ro, since what is added is a standby read-only node, when the read-write node is under high load, even adding a read-only standby read-only node will not relieve the load pressure brought by the write traffic of the read-write node. Therefore, when there are both read-only nodes and read-write nodes in the cluster, it is also possible to determine to add a read-only standby read-only node only when the predicted load level of the read-only node reaches the first preset threshold, so as to avoid frequent determination and execution of adding warm ro events due to the high load of the read-write nodes, resulting in certain resource waste, and even the subsequent inability to effectively delete warm ro due to the constant determination to add warm ro. When there are only read-write nodes in the cluster, if the read-write node reaches a high load, it is also determined to add warm ro. After that, the cluster includes read-write nodes and read-only nodes. When determining whether to add warm ro, it can be determined only based on the predicted load level of the read-only node.

[0126] In actual applications, there may be multiple database clusters. When it is predicted that a database node will reach a high load, the cluster to which the predicted high-load node belongs can be determined first, and then only the backup read-only node of the cluster can be considered.

[0127] Specifically, the load change trend information of each database is regarded as a linear change and linear regression fitting is performed. The predicted load level of each database node after the second preset time period can also be determined according to the slope of the load change trend. That is, the specific implementation method of the above step S3 is only one implementation method of the embodiment of the present application, and the present embodiment is not limited to this.

[0128] Furthermore, if Figure 7 As shown, the above step S33 may include the following steps: step S331, based on the high-load nodes, determine the number of high-load nodes belonging to the same cluster; step S332, based on the number of high-load nodes belonging to the same cluster, determine whether to increase the cluster's backup read-only nodes.

[0129] In this embodiment, the load conditions of all database nodes in the same cluster can be aggregated, and the high load events of each node can be aggregated into cluster dimension events. The node high load events are filtered and aggregated in the cluster dimension, and then the number of high load nodes belonging to the same cluster is counted. Based on the number of high load nodes, it is determined whether to increase the backup read-only node of the cluster.

[0130] In other embodiments, the above step S332 may further include the following processing: when the number of high-load nodes belonging to the same cluster is greater than or equal to a preset number, determining whether the total number of database nodes in the cluster reaches a preset upper limit threshold and whether there are any backup read-only nodes in the cluster; when the total number of database nodes in the cluster does not reach the preset upper limit threshold and there are no backup read-only nodes in the cluster, determining to increase the backup read-only nodes of the cluster.

[0131] The preset number may be greater than or equal to 1, and may be less than the total number of database nodes in the cluster, so as to avoid the extreme situation where the entire cluster is highly loaded, and to determine to add a standby read-only node as long as a database node reaches a high load, so as to ensure that the node addition operation can be completed in advance to cope with the performance loss under high load. The preset upper limit threshold may be any value greater than or equal to 1, and its value is related to the number of database nodes provided by the cloud service provider to the client, and this embodiment does not specifically limit its specific value.

[0132] In this embodiment, when multiple database nodes reach high load in the cluster dimension, that is, when there are multiple high-load nodes, only one increase warm ro event can be output, which effectively avoids the situation where multiple increase warm ro events are generated simultaneously in a database cluster, improves the precision and accuracy of the increase warm ro event, and reduces the burden on the node control end.

[0133] In addition, the database nodes that the client can use are limited. Before the center issues an instruction to add a standby read-only node, it can also first determine whether the database nodes in the corresponding cluster have reached the above-mentioned preset upper limit threshold. When the database nodes in the cluster have reached the above-mentioned preset upper limit threshold, the event of adding warm ro may not be triggered to avoid wasting resources and violating the upper limit of nodes agreed by the customer. When warm ro already exists in the cluster, it means that the current standby read-only node has not been officially launched, and there is no need to add a new node. The event of adding warm ro may not be triggered, thereby further saving resources and reducing costs.

[0134] In other embodiments, after the above determination of whether to add a standby read-only node to the cluster to which the high-load node belongs, the database node management method may also include the following processing: determining the total load level of the cluster and the enabled status of the standby read-only nodes in the cluster; and determining to delete the standby read-only node when the total load level of the cluster and the enabled status of the standby read-only nodes meet preset conditions.

[0135] The total load level of the cluster can be understood as the sum of the load conditions of all database nodes in the cluster, which can be the sum of the load values ​​of all database nodes in the cluster, or can refer to the total information including the load conditions of each database node in the cluster. It should be noted that the total load level of the cluster can include the load of read-only nodes and the load of read-write nodes, but each cluster does not necessarily include both read-only nodes and read-write nodes, and can also include only read-only nodes, or only read-write nodes, which is not specifically limited in this embodiment.

[0136] The above preset conditions correspond to the total load level of the cluster, which may include at least two groups. One group is that the total load level of the cluster is the sum of the load values ​​of all database nodes in the cluster. At this time, the preset conditions may include: no matter what the type of database node in the cluster is, as long as there is a certain amount of remaining load space in the cluster, that is, the total load within the third preset time period does not reach the second preset threshold, and the standby read-only node is never enabled, the standby read-only node can be deleted. Among them, the value of the second preset threshold can refer to the value method of the first preset threshold, which can be a specific load value or a percentage of the total load of the database node. This embodiment does not make specific limitations on this and the specific value of the second preset threshold. It can be less than the total load of all database nodes in the cluster. If it is a percentage, it can also be 50%, 55%, 60%, etc. Another group is that the total load level of the cluster is the total information including the load of each database node in the cluster. At this time, the preset condition may include: no matter what the type of database node in the cluster is, the load of all database nodes in the cluster does not reach the second preset threshold value within the third preset time, and the standby read-only node is never enabled, then the standby read-only node can be deleted. In particular, when there is no read-only node in the cluster, it can be required that the load of the read-write node does not reach the second preset threshold value within the third preset time. If there are both read-write nodes and read-only nodes in the cluster, it can only be required that the load of the read-only node does not reach the second preset threshold value within the third preset time. The value of the second preset threshold here can still refer to the value method of the first preset threshold, which can be a specific load value (the second preset threshold corresponding to each database node can be the same or different), or it can be a percentage of the load of each database node. This embodiment does not make specific restrictions on this and the specific value of the second preset threshold, which can be less than the upper limit of the load of the corresponding database node. If it is a percentage, it can also be 50%, 55%, 60%, etc.

[0137] The enabling status of the standby read-only node includes two types: enabled and not enabled. Putting the standby read-only node online can be considered as enabling the standby read-only node.

[0138] In actual applications, after warm ro is generated, it may not be used for a long time. For example, if the prediction is inaccurate, after adding warm ro, the load level of the database node may not reach the high load as predicted. In this case, warm ro will not be enabled and will remain idle. In view of this, the embodiment of the present application also designs a warm ro recovery mechanism. When the total load of the cluster does not reach the second preset threshold for a continuous third preset time period and the backup read-only node is never enabled, it is determined to delete the backup read-only node. The computing center can send a command to delete warm ro to the node control, and the node control will execute the operation of deleting warm ro accordingly to achieve resource recovery and reduce resource waste.

[0139] Specifically, Flink can be used to detect the load of the entire cluster in real time. When the cluster is in a low-level state for a long time and warm ro is never enabled, the deletion of warm ro is triggered.

[0140] In another embodiment, in combination with Figure 1 and Figure 8 , the complete life cycle of the standby read-only node (warm ro) in the embodiment of the present application is explained, that is, the process of adding, going online and deleting warm ro is as follows.

[0141] (I) Increase in warm ro

[0142] This embodiment uses Flink to design and implement a cluster load prediction model. By streaming and reading second-level performance indicator data, the cluster load is analyzed in real time, and the load that is on an upward trend is located in real time to predict whether it is about to break through the high level line. This can achieve low-latency load prediction to ensure the stability and efficiency of the entire system.

[0143] The cluster load prediction model can be divided into two major modules, namely the load prediction algorithm module and the logical instance dimension aggregation module.

[0144] The load prediction algorithm module is mainly responsible for predicting the load growth trend of each database node. The specific prediction process is as follows.

[0145] (1) Because the upcoming future trend is largely determined by the influence of the latest series trend, a small window can be used to save the latest load data in a short period of time. The algorithm module can use LOWESS series smoothing algorithms of different strengths to perform multiple transformations on the original load time series, so as to simultaneously extract the overall trend characteristics of the load time series and retain the detailed characteristics of each interval.

[0146] The LOWESS sequence smoothing algorithm is used. Based on the weighted least squares method, a weight is assigned to each data point to transform it into a weighted regression problem. For each data point, a time window can be selected, which includes some data points around the data point as the neighbor points of the data point. According to the distance d between the neighbor point and the data point, the weight w of each neighbor point is calculated according to formula (1). The closer the neighbor point, the greater the weight.

[0147] Then, for each data point, weighted least squares regression is performed according to the weights of the neighboring points to obtain the smoothed value of the data point.

[0148] The LOWESS algorithm does not require assumptions about the distribution of data and can be applied to a variety of data types. By adjusting the window size and the weight of neighbor points, the data can be locally smoothed, retaining the overall trend and detailed characteristics of the data. The degree of smoothing and the number of neighbor points can be adjusted to adapt to different data sets; and the calculation time is independent of the size of the data set, which is suitable for real-time data processing and online analysis.

[0149] (2) The ADF test can be used to identify whether the window sequence after noise smoothing is in a stable growth trend. If the current load change trend is increasing at a stable rate, the upcoming future moment can be predicted at the same rate. When the predicted value reaches the high load level threshold, the node high load event is output.

[0150] The null hypothesis of the ADF test is that the time series has a unit root, that is, non-stationarity. If the null hypothesis is rejected, the time series is considered to be stationary. The statistic of the ADF test is a t statistic, which determines whether the null hypothesis is rejected by comparing the t value with the critical value.

[0151] The inspection steps are as follows:

[0152] a) Differentiate the load time series to obtain a new load time series sequence and make it stationary.

[0153] b) Establish a generalized autoregressive model as shown in formula (2) above to fit the differenced series and test whether the residual of the model has a unit root.

[0154] c) Calculate the t-value of the statistic and compare it with the critical value to determine whether the null hypothesis is rejected. Calculate the t-value according to the above formula (3).

[0155] The advantage of the ADF test is that it has a fast calculation speed and is suitable for the current sequence linear stationarity test scenario.

[0156] The load prediction algorithm module can output the high load situation of each database node in real time. The logic instance dimension aggregation module is responsible for aggregating the high load events of each node into cluster dimension events, filtering and aggregating the node high load events in the cluster dimension. When multiple nodes meet the high load in the cluster dimension, an increase warm ro event will be output once; when the cluster has reached the preset upper limit threshold of the database node, the increase warm ro event will not be triggered; when the cluster already has warm ro, the increase warm ro event will not be triggered. This aggregation logic effectively avoids the simultaneous generation of multiple increase warm ro events in a database cluster, improves the accuracy and precision of events, and reduces the burden on the node control end. Accordingly, the node control end actually executes the increase warm ro task flow.

[0157] (II) Launch of warm ro

[0158] When the warm ro node is generated and not online, it is not visible to the user. When the application server actually detects that the current cluster is in a high-load state, it will trigger an event to add a node. At this time, if there is a warm ro node in the cluster, the node control end will officially bring the warm ro online as a read-only node. This process takes seconds, thereby reducing the time consumption of the original node addition operation. At this time, the warm ro node is visible to the user.

[0159] (III) Deletion of warm ro

[0160] After warm ro is generated, the warm ro recovery mechanism uses Flink to detect the load of the entire cluster in real time. When the cluster is in a low-level state for a long time and warm ro is never enabled, a warm ro deletion event will be triggered, and the node control end will delete warm ro accordingly to achieve resource recovery.

[0161] In summary, the database node management method provided in the embodiment of the present application can predict the load change trend information of each database node within a second preset time period in the future based on the load timing data of each database node within a first preset time period before the current moment and closest to the current moment, and determine whether to add a backup read-only node to the cluster to which each database node belongs based on the load change trend information. In this way, the embodiment of the present application is based on the load time series data of each database node in the recent period of time. According to the correspondence between load and time in the load time series data, a mathematical statistical method can be used to predict the load change trend information of each database node in the future period of time. The load change trend information of each database node can be used to predict the load level of each database node in the future period of time, that is, it can be predicted in time before the database node reaches a high load, so that the backup read-only node of the cluster to which the database node belongs can be added in advance, so that when the load of the database node where the client is located reaches a high load, the added backup read-only node can be directly put online and expanded for use, thereby reducing the node expansion operation at the minute level to the second level, greatly shortening the expansion operation time of the database node, and promptly responding to the performance loss problem when the database node is under high load, which can improve the concurrency, reliability, scalability and overall performance of the entire database cluster, and at the same time can achieve more accurate database node expansion operations to avoid waste of resources and costs.

[0162] Some embodiments of the present application further provide a database node management device, which is used to execute the database node management method provided in any of the above embodiments. Fig. 9 A schematic diagram of the database node management device is shown in FIG. Fig. 9 As shown, the database node management device includes:

[0163] A data acquisition module, used to acquire the load time series data of each database node before the current moment and within a first preset time period closest to the current moment;

[0164] A load prediction module, used to predict the load change trend information of each database node within a second preset time period in the future based on the load time series data of each database node;

[0165] The node management module is used to determine whether to add a standby read-only node to the cluster to which each database node belongs based on the load change trend information of each database node.

[0166] It can be understood that the database node management device provided in this embodiment and the database node management method provided in the embodiment of the present application are based on the same inventive concept, and can at least achieve the same beneficial effects of the database node management method, and the various implementation methods of the database node management method embodiment are also applicable to the embodiment of the database node management device, which will not be repeated here.

[0167] Some embodiments of the present application also provide a database node management system, such as Figure 1 As shown, the database node management system includes a computing center end, an application server end and a node control end. The computing center end is used for the above-mentioned database node management method, and the application server end is used to monitor the load level of the database node, and when the load level is greater than a preset first threshold, send an instruction to the node control end to go online; when the computing center end determines to add a spare read-only node, the node control end adds a spare read-only node to the cluster; and when receiving the instruction to go online, the node control end brings the spare read-only node online.

[0168] Specifically, Figure 1 and Figure 8 As shown, the application server is connected to the client of the application cloud database, and can collect the performance indicator data of the database node where the client is located in real time, and monitor the load level of the database node based on the collected performance indicator data, and send an instruction to add a node to the node control end when the load level is greater than the preset first threshold. The computing center end can send an instruction to add a standby read-only node to the node control end when it is determined to add a standby read-only node. The node control end is connected to the computing center end and the application server end respectively, and is used to actually perform operations such as adding database nodes, deleting database nodes, and bringing standby read-only nodes online. After receiving the instruction to add a standby read-only node sent by the computing center end, add a standby read-only node to the cluster; after receiving the instruction to add a node sent by the application server end, bring the standby read-only node online; and after receiving the instruction to delete the standby read-only node sent by the computing center end, delete the standby read-only node in the cluster; and when there is no standby read-only node in the cluster, add a read-only node.

[0169] It can be understood that the database node management system provided in this embodiment, by applying the above-mentioned database node management method, can not only achieve the same beneficial effects as the database node management method, but also can distribute different tasks, such as communication tasks with the client, computing tasks, and database node management tasks, to different execution ends for processing, further improving the parallelism, stability, reliability and prediction accuracy of the overall system. In addition, the various implementation methods of the database node management method embodiment are also applicable to the embodiment of the database node management device, which will not be repeated here.

[0170] It should be noted that the data involved in this application (including but not limited to data used for model training, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0171] The present application also provides an electronic device to execute the above database node management method. Fig.10 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. Fig.10 As shown, the electronic device 4 includes: a processor 400, a memory 401, a bus 402 and a communication interface 403. The processor 400, the communication interface 403 and the memory 401 are connected via the bus 402; the memory 401 stores a computer program that can be run on the processor 400, and when the processor 400 runs the computer program, it executes the database node management method provided in any of the aforementioned embodiments of the present application.

[0172] The memory 401 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the device network element and at least one other network element is realized through at least one communication interface 403 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.

[0173] The bus 402 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 401 is used to store programs, and the processor 400 executes the program after receiving an execution instruction. The database node management method or database node management method disclosed in any implementation of the aforementioned embodiment of the present application may be applied to the processor 400, or implemented by the processor 400.

[0174] The processor 400 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 400. The above processor 400 can be a general-purpose processor, including a processor (Central Processing Unit, referred to as CPU), a network processor (Network Processor, referred to as NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and completes the steps of the above method in combination with its hardware.

[0175] The electronic device provided in the embodiment of the present application and the database node management method or database node management method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.

[0176] The present application also provides a computer-readable storage medium corresponding to the database node management method provided in the above embodiment. Fig.11 The computer-readable storage medium shown is a CD 50 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the database node management method provided by any of the aforementioned embodiments will be executed.

[0177] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0178] An embodiment of the present application also provides a computer program product, including a computer program, which is executed by a processor to implement the database node management method of any of the above embodiments.

[0179] The computer-readable storage medium and computer program product provided in the above-mentioned embodiments of the present application are based on the same inventive concept as the database node management method provided in the embodiments of the present application, and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0180] It should be noted that:

[0181] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.

[0182] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be interpreted as reflecting the following schematic diagram: the claimed application requires more features than the features explicitly recited in the individual claims. More specifically, as reflected in the claims below, the inventive aspects are less than all the features of the individual embodiments disclosed above. Therefore, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.

[0183] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.

[0184] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A database node management method, characterized in that: include: Obtaining load time series data of each database node before the current moment and within a first preset time period closest to the current moment; Based on the load time series data of each database node, respectively predict the load change trend information of each database node within a second preset time period in the future; Based on the load change trend information of each database node, determine whether to add a standby read-only node to the cluster to which each database node belongs.

2. The method according to claim 1, characterized in that: The step of obtaining the load time series data of each database node within a first preset time period before the current moment and closest to the current moment includes: Obtaining performance indicator data of each database node at each moment within a first preset time period before the current moment and closest to the current moment; Based on the performance indicator data of each database node, the load time series data of each database node is determined.

3. The method according to claim 1, characterized in that The method of predicting the load change trend information of each database node within a second preset time period in the future based on the load time series data of each database node includes: Perform noise smoothing on the load time series data of each database node to obtain new load time series data of each database node; Based on the new load time series data of each database node, the load change trend information of each database node within a second preset time period in the future is predicted respectively.

4. The method according to claim 3, characterized in that The noise smoothing of the load time series data of each database node to obtain new load time series data of each database node includes: Determine the neighboring points of each data point in the load time series data of each database node; Based on the weights of each neighboring point of the data point, the smoothing value of the corresponding data point is calculated; Based on the smoothed value of each data point in each database node, new load time series data corresponding to each database node is formed.

5. The method according to claim 4, characterized in that Before calculating the smoothing value of the corresponding data point based on the weights of each neighboring point of the data point, the method further includes: Determine the time windows corresponding to each data point in the load time series data of each database node; The weight of each neighbor point is determined according to the distance between the neighbor point of each data point in the corresponding time window and the corresponding data point.

6. The method according to claim 3, characterized in that The method of predicting the load change trend information of each database node within a second preset time period in the future based on the new load time series data of each database node includes: Predicting the change trend information of the new load time series data of each database node within the second preset time period based on the autoregressive model; Based on the change trend information of the new load time series data, the load change trend information of the corresponding database node within the second preset time length is determined.

7. The method according to claim 1, characterized in that The determining whether to add a standby read-only node to the cluster to which each database node belongs based on the load change trend information of each database node includes: Based on the load change trend information of each database node, respectively determine the predicted load level of each database node after the second preset time period; Filter out high-load nodes whose predicted load levels reach a first preset threshold from each database node; Based on the high-load node, it is determined whether to add a standby read-only node to the cluster to which the high-load node belongs.

8. The method according to claim 7, characterized in that The determining, based on the high-load node, whether to add a standby read-only node to the cluster to which the high-load node belongs includes: Based on the high-load nodes, determining the number of high-load nodes belonging to the same cluster; Based on the number of high-load nodes belonging to the same cluster, it is determined whether to increase a standby read-only node of the cluster.

9. The method according to claim 8, characterized in that The determining whether to add a standby read-only node of the cluster based on the number of high-load nodes belonging to the same cluster includes: When the number of high-load nodes belonging to the same cluster is greater than or equal to a preset number, determining whether the total number of database nodes in the cluster reaches a preset upper limit threshold, and whether there are standby read-only nodes in the cluster; When the total number of database nodes in the cluster does not reach the preset upper limit threshold and there is no standby read-only node in the cluster, it is determined to increase the standby read-only node of the cluster.

10. The method according to claim 8, characterized in that After determining whether to add a standby read-only node to the cluster to which the high-load node belongs, the method further includes: Determining the total load level of the cluster and the activation of standby read-only nodes in the cluster; When the total load level of the cluster and the activation status of the standby read-only node meet preset conditions, determine to delete the standby read-only node; the preset conditions include: the total load in the cluster does not reach a second preset threshold within a third preset time period, and the standby read-only node is never activated.

11. A database node management device, characterized in that: include: A data acquisition module, used to acquire the load time series data of each database node before the current moment and within a first preset time period closest to the current moment; A load prediction module, used to predict the load change trend information of each database node within a second preset time period in the future based on the load time series data of each database node; The node management module is used to determine whether to add a standby read-only node to the cluster to which each database node belongs based on the load change trend information of each database node.

12. A database node management system, characterized in that: It comprises a computing center end, an application service end and a node control end, wherein the computing center end is used to execute the database node management method according to any one of claims 1 to 10; The application server is used to monitor the load level of the database node and send an instruction to the node control terminal to go online when the load level is greater than a preset first threshold; The node control end adds a spare read-only node to the cluster when the computing center determines to add a spare read-only node; and brings the spare read-only node online when receiving the instruction to add the node.

13. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method according to any one of claims 1 to 10.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to any one of claims 1 to 10.