Capacity expansion and contraction prediction method and device, electronic equipment and storage medium

By using a pre-trained scaling prediction model in the server cluster, and automatically adjusting resources based on real-time monitoring data, the problem of difficult to cope with rapidly changing load requirements in the existing technology is solved, and efficient scaling management is achieved.

CN120086086APending Publication Date: 2025-06-03DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411966038.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When the resource management of existing server clusters faces rapidly changing load requirements, it is difficult to meet service requirements in a timely manner, resulting in delivery pressure and user losses.

Method used

The scaling parameters, including scaling type and scale, are obtained by obtaining the target monitoring data (including CPU, memory, disk, network and SWAP monitoring data) and inputting it into the pre-trained scaling prediction model. This model extracts data features through the LSTM network layer, combines the dropout layer and the fully connected layer, and outputs the scaling prediction results.

Benefits of technology

Automatic scaling is realized, improving scaling is achieved, ensuring that the server cluster can adjust resources in a timely manner when facing load changes, and reducing delivery pressure and user losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086086A_ABST
    Figure CN120086086A_ABST
Patent Text Reader

Abstract

The invention provides a capacity expansion and contraction prediction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining target monitoring data, including CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data and SWAP monitoring data; inputting the target monitoring data into the capacity expansion and contraction prediction model to obtain capacity expansion and contraction parameters, wherein the capacity expansion and contraction parameters comprise a capacity expansion and contraction type and a capacity expansion and contraction scale; the capacity expansion and contraction prediction model is pre-trained based on sample monitoring data, the sample monitoring data comprises CPU, memory, disk, network and SWAP monitoring data, and the labels are corresponding capacity expansion and contraction parameters. The capacity expansion and contraction parameters are determined based on the cluster monitoring data by using the capacity expansion and contraction prediction model, so that the capacity expansion and contraction prediction efficiency and the capacity expansion and contraction efficiency are improved; in the training process, the used sample data includes various data included in the cluster operation process, so that the integrity of the sample data is improved, the learning of the model on the relationship between various operation data and capacity expansion and contraction parameters is improved, and the model prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and in particular, to a method, device, electronic device, and storage medium for scaling prediction. Background Art

[0002] At present, the resource management of server clusters is crucial for maintaining system performance and cost efficiency. Traditional resource scaling and relocation strategies often rely on predefined policies and static thresholds, and cannot effectively cope with rapidly changing load demands. When the capacity risk of a service is already very clear, and then scaling and asset relocation are carried out, it is difficult to meet the service requirements in a timely manner. While increasing the delivery pressure on asset and resource administrators, it may also cause losses to users due to failure to deliver on time. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, device, electronic device, and storage medium for scaling prediction, so as to improve the efficiency of scaling prediction, and further improve the efficiency of scaling.

[0004] According to an aspect of the present invention, a method for scaling prediction is provided. The method includes:

[0005] Obtaining target monitoring data, where the target monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data;

[0006] Inputting the target monitoring data into a pre-trained scaling prediction model, and obtaining scaling parameters output by the scaling prediction model, where the scaling parameters include a scaling type and a scaling scale;

[0007] Wherein, the scaling prediction model is pre-trained through the following steps:

[0008] Obtaining sample monitoring data and a label corresponding to the sample monitoring data, where the sample monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data, and the label is the scaling parameter corresponding to the sample monitoring data;

[0009] Inputting each sample monitoring data into an initial scaling prediction model, where the initial scaling prediction model is an LSTM model;

[0010] Obtaining a prediction result output by the initial scaling model, and calculating a target difference between the prediction result and the label of the sample monitoring data;

[0011] In the case where the target difference is higher than the preset difference threshold, adjust the model parameters of the initial scaling prediction model based on the target difference until the target is not higher than the preset difference threshold, and obtain the target scaling prediction model.

[0012] In one possible embodiment, the obtaining of the sample monitoring data and the label corresponding to the sample monitoring data includes:

[0013] Based on the preset monitoring metrics, obtain the initial monitoring data from the scaling history. The preset monitoring metrics include: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data;

[0014] Perform data cleaning on the initial monitoring data, and normalize the cleaned monitoring data to obtain the sample monitoring data; the data cleaning includes: missing value filling and outlier removal;

[0015] Based on the timestamp information of each sample monitoring data and the timestamp information of the scaling history, determine the label of each sample monitoring data.

[0016] In one possible embodiment, the method further includes:

[0017] For the normalized sample monitoring data, generate time series data according to the timestamp information of the sample monitoring data as the target sample monitoring data;

[0018] Divide the target sample monitoring data into a training set, a validation set, and a test set.

[0019] In one possible embodiment, the initial scaling prediction model includes multiple LSTM network layers, a dropout layer, a fully connected layer, and an output layer. The method further includes:

[0020] The initial scaling prediction model extracts the data features in the sample monitoring data through multiple LSTM network layers and inputs them into the dropout layer;

[0021] Process the data features through the dropout layer and input the processed data features into the fully connected layer;

[0022] Output the result vector output by the fully connected layer to the output layer, and output the prediction result based on the result vector through the output layer.

[0023] In one possible embodiment, the output layer includes a first neuron and a second neuron;

[0024] Outputting the prediction result by the output layer based on the result vector includes:

[0025] Performing binary classification on the result vector by the first neuron using the sigmoid activation function, and outputting the prediction result of the scaling type;

[0026] Outputting a continuous value by the second neuron using the linear activation function based on the result vector as the prediction result of the scaling scale.

[0027] In a possible embodiment, calculating the target difference between the prediction result and the label of the sample monitoring data includes:

[0028] Calculating the binary cross-entropy between the prediction result of the scaling type and the scaling type included in the label of the sample monitoring data;

[0029] Calculating the mean square error or root mean square error between the prediction result of the scaling scale and the scaling scale included in the label of the sample monitoring data.

[0030] In a possible embodiment, the initial scaling prediction model further includes: an Adam optimizer, and adjusting the model parameters of the initial scaling prediction model based on the target difference includes:

[0031] Updating the model parameters of the initial scaling prediction model by the Adam optimizer based on the target difference.

[0032] According to another aspect of the present invention, there is provided a scaling prediction device, and the device includes:

[0033] An acquisition module, configured to acquire target monitoring data, where the target monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data;

[0034] A prediction module, configured to input the target monitoring data into a pre-trained scaling prediction model, and acquire the scaling parameters output by the scaling prediction model, where the scaling parameters include a scaling type and a scaling scale;

[0035] A training module, configured to pre-train the scaling prediction model through the following steps:

[0036] Acquiring sample monitoring data and the label corresponding to the sample monitoring data, where the sample monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data, and the label is the scaling parameter corresponding to the sample monitoring data;

[0037] Input each of the sample monitoring data into an initial scaling prediction model, where the initial scaling prediction model is an LSTM model;

[0038] Obtain the prediction result output by the initial scaling prediction model, and calculate the target difference between the prediction result and the label of the sample monitoring data;

[0039] In the case where the target difference is higher than a preset difference threshold, adjust the model parameters of the initial scaling prediction model based on the target difference until the target is not higher than the preset difference threshold, and obtain a target scaling prediction model.

[0040] According to another aspect of the present invention, there is provided an electronic device, including:

[0041] A processor; and

[0042] A memory storing a program,

[0043] wherein the program includes instructions that, when executed by the processor, cause the processor to execute any of the above-mentioned scaling prediction methods.

[0044] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any of the above-mentioned scaling prediction methods.

[0045] One or more technical solutions provided in the embodiments of the present invention pre-train a scaling prediction model and use the scaling prediction model to determine the scaling parameters of the cluster based on the monitoring data during the operation of the cluster, realizing automatic scaling, improving the scaling prediction efficiency, and further improving the scaling efficiency; furthermore, during the training process of the scaling prediction model, the sample data used includes various data included in the operation of the cluster, improving the integrity of the sample data, and further improving the learning of the relationship between various operation data in the cluster and the scaling parameters by the scaling prediction model, and improving the prediction accuracy of the scaling prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:

[0047] Figure 1 is a schematic flowchart of a scaling prediction method provided by an embodiment of the present invention;

[0048] Figure 2 is another schematic flowchart of a scaling prediction method provided by an embodiment of the present invention;

[0049] Figure 3A structural schematic diagram of the scaling prediction device provided by an embodiment of the present invention;

[0050] Figure 4 A block diagram of an exemplary electronic device capable of implementing the embodiments of the present invention is shown. Detailed implementation manners

[0051] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0052] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0053] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or the interdependence relationship.

[0054] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0055] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0056] According to the current status of server resource management and delivery, the existing technologies mainly have the following two disadvantages:

[0057] (1) The current server relocation and expansion strategies are not intelligent and forward-looking enough, and resource bottleneck problems are likely to occur. At present, server relocation and expansion mostly rely on users to submit requirements, and then asset and resource administrators deliver according to user requirements. The asset delivery process is complex, involving multiple tasks such as relocation, expansion, testing, and acceptance. It requires coordinating manpower in various directions, and there are also many problems in the process adaptation of different resources. This brings huge delivery pressure to resource and asset administrators and also brings risks of resource shortage and full capacity to users.

[0058] (2) The current monitoring indicators of the server cluster are not fully utilized. The monitoring indicators in dimensions such as CPU, memory, disk, network, and SWAP swap partition (a virtual memory implementation method in the Linux system architecture) have been built very comprehensively, but in fact, only limited indicators such as CPU usage, memory usage, and network bandwidth are used in various problem analyses. This does not give full play to the role of monitoring.

[0059] Based on this, the embodiments of the present invention provide a method, device, electronic device, and storage medium for scaling prediction. The scaling prediction method provided by the embodiments of the present invention can be applied to any electronic device with scaling prediction function, and the electronic device can be a server, a computer, a mobile terminal, etc. The solution of the present invention will be described below with reference to the accompanying drawings:

[0060] Figure 1 FIG. is a schematic flowchart of a scaling prediction method provided by an embodiment of the present invention, which may include the following steps:

[0061] S101. Obtain target monitoring data, where the target monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data;

[0062] S102. Input the target monitoring data into a pre-trained scaling prediction model, and obtain the scaling parameters output by the scaling prediction model, where the scaling parameters include the scaling type and the scaling scale;

[0063] Among them, the scaling prediction model is pre-trained through the following steps:

[0064] S201. Obtain sample monitoring data and the labels corresponding to the sample monitoring data, where the sample monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data, and the label is the scaling parameter corresponding to the sample monitoring data;

[0065] S202. Input each of the sample monitoring data into an initial scaling prediction model, where the initial scaling prediction model is an LSTM model;

[0066] S203. Obtain the prediction result output by the initial scaling prediction model, and calculate the objective difference between the prediction result and the label of the sample monitoring data;

[0067] S204. When the objective difference is higher than a preset difference threshold, adjust the model parameters of the initial scaling prediction model based on the objective difference until the objective is not higher than the preset difference threshold, so as to obtain a target scaling prediction model.

[0068] Applying the embodiments of the present invention, by pre-training a scaling prediction model and determining the scaling parameters of a cluster based on the monitoring data during the operation of the cluster through this scaling prediction model, automated scaling is achieved, the scaling prediction efficiency is improved, and thus the scaling efficiency is improved; furthermore, during the training process of the scaling prediction model, the sample data used includes various data included in the operation process of the cluster, which improves the integrity of the sample data, and further improves the learning of the relationship between various operation data in the cluster and the scaling parameters by the scaling prediction model, and improves the prediction accuracy of the scaling prediction model.

[0069] The following gives an exemplary description of the above S101 - S102 and S201 - S204:

[0070] In a possible embodiment, the above target monitoring data can be obtained according to a preset time period, and the preset time period can be set according to the actual application scenario. For example, the time period can be set to 1 minute, that is, the target monitoring data is obtained every minute. In a possible embodiment, the time period can also be set according to the traffic change of the cluster. For example, the time period can be set to 1s during the traffic peak period of the cluster and 2 minutes during the traffic trough period, etc., so as to collect the target monitoring data more flexibly and achieve a balance between the collection efficiency and the amount of resources consumed by the collection.

[0071] In a possible embodiment, the above-mentioned target monitoring data may include: CPU monitoring data, memory monitoring data, network monitoring data, and SWAP monitoring data. Among them, the CPU monitoring data may include the server CPU load within a preset duration, the number of context switches per second, the number of CPU interrupts per second, the system CPU time ratio, the user CPU time ratio, and the waiting IO CPU time ratio. The above-mentioned preset duration can be set according to the actual application scenario. For example, it can be the server CPU load within the last one minute / five minutes / fifteen minutes. The above-mentioned memory monitoring data may include the memory usage, memory utilization rate, total memory, file system memory cache value, block device read / write memory buffer amount, etc. The network monitoring data may include the input / output bandwidth of the primary network card, the input / output packet rate of the primary network card, the number of established tcp connections, the number of tcp received packets, the number of tcp sent packets, and the number of tcp error packets. The disk monitoring data may include the total disk space of the entire server, the total disk usage of the entire server, the total HOME disk space, the HOME disk space usage, the HOME disk space utilization rate, the total root disk space, the root disk space usage, the root disk space utilization rate, the per-second disk I / O read volume of the entire server, and the per-second disk I / O write volume of the entire server. The SWAP monitoring data includes the total swap partition size, the swap partition usage, and the swap partition free space.

[0072] The above-mentioned monitoring data can be obtained through a cluster monitoring tool. Exemplarily, the running data in the cluster can be monitored in real time through the cluster monitoring tool Prometheus and stored in a database. Correspondingly, the corresponding monitoring data can be obtained from the database based on preset monitoring metrics.

[0073] After obtaining the above-mentioned target monitoring data, the target monitoring data can be input into a pre-trained scaling prediction model. The scaling prediction model determines whether to scale up or down currently based on the target monitoring data, as well as the corresponding scaling-up or scaling-down scale. The scaling-up or scaling-down scale refers to the number of servers to be scaled up or down, and can also be the number of CPU cores, memory, disks, etc. that need to be scaled up or down.

[0074] The above-mentioned scaling prediction model is pre-trained through the above S201-S204. The following is an exemplary description of the above steps:

[0075] In S201, the sample monitoring data can be obtained from the scaling history record. In a possible embodiment, after each scaling operation, the target monitoring data used in that operation and the corresponding scaling parameters can be stored in the scaling history record database. Exemplarily, the target monitoring data identifier and the scaling parameter identifier can be stored in correspondence. The target monitoring data identifier may include the server corresponding to the target monitoring data and the acquisition timestamp, and the scaling parameter identifier may include the corresponding server and the calculation timestamp.

[0076] During the model training process, the sample monitoring data can be obtained from the scaling history record database. In a possible embodiment, the sample monitoring data can be obtained through the following steps:

[0077] S211. Based on the preset monitoring metrics, obtain the initial monitoring data from the scaling history record. The preset monitoring metrics include: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data.

[0078] S212. Clean the initial monitoring data and normalize the cleaned monitoring data to obtain the sample monitoring data. The data cleaning includes: missing value filling and outlier removal.

[0079] S213. Based on the timestamp information of each sample monitoring data and the timestamp information of the scaling history record, determine the label of each sample monitoring data.

[0080] Exemplarily, all the data in the history record database can be obtained as the sample monitoring data, or a preset number of historical records can be obtained as the sample monitoring data through sampling algorithms such as the K-medoids algorithm. The labels of the above sample monitoring data can also be obtained from the scaling history record.

[0081] As a possible implementation, after obtaining the sample monitoring data, the corresponding scaling parameters can be obtained from the scaling history record based on the sample monitoring data identifier as the label of the sample monitoring data.

[0082] In a possible embodiment, after obtaining the sample monitoring data, it can be cleaned. Exemplarily, according to the characteristics and missing degree of the sample monitoring data, missing values can be filled and outliers can be removed. The characteristics of the above sample monitoring data can be its time characteristics. For example, the change period of sample monitoring data A is 1 minute. Therefore, if the sample monitoring data A at a certain moment is missing, the missing value can be filled based on the data records within 1 minute.

[0083] In a possible embodiment, since the scales of the monitoring metrics included in the sample monitoring data are inconsistent, the sample monitoring data can be normalized (also known as standardized). In the embodiments of the present invention, the sample monitoring data can be normalized by any feasible normalization method.

[0084] In a possible embodiment, the normalized sample monitoring data can be constructed into time series data according to the timestamp information, so as to obtain the target sample monitoring data, enabling the scaling prediction model to better learn the relationship between the temporal characteristics of the monitoring data and the scaling parameters.

[0085] Exemplarily, each sample monitoring data can be sorted from early to late according to the time indicated by the timestamp information to form a sample monitoring data matrix, where the number of columns of the matrix is the number of preset monitoring metrics, and the number of rows is the number of timestamps. Then, the matrix can be divided multiple times with the scaling timestamp as the time node to obtain each target sample monitoring data. Exemplarily, for the scaling timestamp 1, the sample monitoring data with timestamps before it can be used as the target sample monitoring data corresponding to this scaling timestamp.

[0086] In a possible embodiment, the above target sample monitoring data can be divided into a training set, a validation set, and a test set. As a possible implementation, the target sample monitoring data can be divided into a training set, a validation set, and a test set in a ratio of 6:2:2. Among them, the sample monitoring data in the training set is used to train the initial scaling prediction model, the sample monitoring data in the validation set is used to verify the training result of the scaling prediction model, and the test set is used to test the scaling prediction model that has passed the verification.

[0087] The above scaling prediction model is an LSTM (Long Short-Term Memory) model, whose core consists of LSTM units. Each unit includes three main gate structures: a forget gate, an input gate, and an output gate. These gating structures enable the LSTM to regulate the information flow, determine which information is retained or discarded, and thus effectively maintain and update the internal state of the unit. This LSTM model can extract the temporal features in the target monitoring data.

[0088] Metrics such as the capacity, load, and performance of a server cluster usually manifest as time series data, which has obvious temporal dependence and periodicity. LSTM, through its gating mechanism (including forget gate, input gate, and output gate), can effectively manage the long-term flow of information and is suitable for processing sequential data with long-term dependencies. Additionally, a major challenge in time series prediction is capturing long-term dependencies, that is, the current prediction may be affected by data points from a long time ago. Traditional neural networks and some simpler time series models (such as ARIMA (Autoregressive Integrated Moving Average Model), SVM (Support Vector Machine), etc.) often struggle to handle such dependencies. LSTM, through its internal memory cells, can "remember" a large amount of historical information and adjust its output based on the current input and its memory state, which makes LSTM usually more effective than other models when predicting time series data.

[0089] Furthermore, the LSTM model can simulate non-linear relationships, which is crucial for predicting server performance under different operating and load conditions. The load of a server cluster is usually affected by multiple factors, such as user behavior, accidental events, promotional activities, etc. The non-linear delivery of these factors may lead to sudden increases or decreases in load. LSTM can capture these non-linear relationships through its complex neural network, thereby providing more accurate predictions. Also, different from traditional machine learning models that require manual feature design, deep learning models such as LSTM can automatically learn useful features from a large amount of data. This automatic feature extraction ability enables the model to discover potential and complex patterns from the raw data, further improving the accuracy and robustness of the prediction.

[0090] In a possible embodiment, the LSTM model may include multiple layers of neural networks. Exemplarily, it may include multiple long short-term memory network layers, a dropout layer, a fully connected layer, and an output layer.

[0091] Among them, the above-mentioned multiple long short-term memory layers are used to extract the data features of the input target sample monitoring data. As a possible implementation, the LSTM model may further include an embedding layer, which can embed the target sample monitoring data to obtain a target vector corresponding to the target sample monitoring data. The multiple long short-term memory layers can extract the feature data of the target sample monitoring data based on the target vector and output the feature data to the dropout layer.

[0092] The Dropout layer is used to prevent overfitting during training. Specifically, each neuron in the Dropout layer is retained with probability p and stops working with probability 1 - p. This ensures that the neurons retained during each forward pass are different, making the model less dependent on certain local features and thus improving the model's generalization ability.

[0093] The Dropout layer can output the feature data to the fully connected layer, which can output the result vector to the output layer based on the feature data. The output layer is used to output the scaling parameter based on the result vector.

[0094] In a possible embodiment, the scaling parameter can be defined according to the actual application scenario. For example, the scaling parameter can be defined to include the scaling type, i.e., whether the output needs to be scaled up or down, or the scaling parameter can be defined to include the scaling type and the scaling scale.

[0095] In a possible embodiment, the output layer includes a first neuron and a second neuron; the first neuron and the second neuron are connected in parallel and both receive the input from the fully connected layer. The output of the prediction result by the output layer based on the result vector includes:

[0096] The first neuron uses the sigmod activation function to perform binary classification based on the result vector and outputs the prediction result of the scaling type;

[0097] The second neuron uses the linear activation function to output a continuous value based on the result vector as the prediction result of the scaling scale.

[0098] Correspondingly, the calculation of the objective difference between the prediction result and the label of the sample monitoring data includes:

[0099] Calculate the binary cross-entropy between the prediction result of the scaling type and the scaling type included in the label of the sample monitoring data;

[0100] Calculate the mean square error or root mean square error between the prediction result of the scaling scale and the scaling scale included in the label of the sample monitoring data.

[0101] In a possible embodiment, the above output layer can also include only one neuron. Exemplarily, when it is necessary to determine whether the current server cluster needs to be scaled up, the output layer can be designed as the first neuron, i.e., using the sigmoid activation function to perform binary classification based on the result vector to determine whether to scale up. When it is necessary to predict the scaling scale, the output layer is designed as the above second neuron, i.e., using the linear activation function to directly output a continuous value representing the number of scale-ups.

[0102] In a possible embodiment, when only the first neuron is included in the above output layer, binary cross-entropy can be used as the loss function to calculate the difference between the prediction result and the label, and the initial scaling prediction model can be trained based on this difference. When only the second neuron is included in the above output layer, mean squared error or root mean squared error can be used as the loss function to calculate the difference between the prediction result and the label, and the initial scaling prediction model can be trained based on this difference.

[0103] In a possible embodiment, when both the first neuron and the second neuron are included in the above output layer, the model parameters need to be adjusted based on binary cross-entropy and mean squared error or root mean squared error.

[0104] In a possible embodiment, the above output layer may further include an optimizer. Specifically, it may include an adam optimizer, which is an adaptive optimization algorithm that can adjust the learning rate through historical gradient information to improve the training efficiency.

[0105] When the above target difference converges, it can be determined that the model training is completed. The trained model can be verified using the above validation set. If the model can output correct results for all the data in the validation set, it can be determined that the model passes the verification, and the model can be tested using the test set. If the model can output correct results for all the data in the test set, it can be determined that the model passes the test and can be put into use.

[0106] If the model fails to pass the verification or the test, the training set is reconstructed to train the model until the model passes the test.

[0107] In a possible embodiment, when both the first neuron and the second neuron are included in the above output layer, it is necessary for both the above binary cross-entropy and mean squared error (root mean squared error) to converge before it can be determined that the model training is completed.

[0108] As Figure 2 shown, Figure 2 Another flowchart of the scaling prediction method provided by the embodiment of the present invention may include three parts: feature selection, feature engineering, and LSTM model prediction.

[0109] Feature selection refers to the input of the selection model, which can include CPU monitoring data, memory monitoring data, network monitoring data, disk monitoring data and SWAP monitoring data. Among them, CPU monitoring data includes CPU usage, server CPU load in the last one minute / five minutes / fifteen minutes, number of context switches per second, number of CPU interruptions per second, system CPU time ratio, user CPU time ratio, waiting IO CPU time ratio, etc. Memory monitoring data includes memory usage, memory usage, total memory, file system memory cache value, block device read and write memory buffer, etc. Network monitoring data includes main network card input / output bandwidth, main network card input / output packet rate, number of established TCP connections, number of TCP received packets, number of TCP sent packets, number of TCP error packets, etc. Disk monitoring data includes total disk space of the entire server, total disk usage of the entire server, total HOME disk space, HOME disk space usage, HOME disk space usage, total root disk space, root disk space usage, root disk space usage, disk I / O reads per second of the entire server, and disk I / O writes per second of the entire server. SWAP monitoring data includes the total amount of swap partitions, swap partition usage, swap partition free space, etc.

[0110] Feature engineering includes data cleaning, standardization / normalization operations, feature engineering construction, and data set construction. Through feature engineering, the input of the model can be obtained.

[0111] LSTM model prediction includes the LSTM network layer to capture the implicit mode of the model, the dropout layer to prevent overfitting, and the fully connected layer to obtain the output result vector. After that, different neurons are used to process the output results using different activation functions to obtain the scaling type and scale parameters.

[0112] By using the expansion and contraction prediction method provided by the embodiment of the present invention, a time series neural network model such as LSTM is applied to the relocation and expansion of the server, and the scale of server relocation and expansion is predicted, which is more timely and forward-looking, greatly reducing the capacity risk brought to the service due to untimely relocation and expansion, and recovering the losses brought to the company. In addition, the indicator information collected by the existing monitoring system is fully utilized, and the existing server capacity, load, performance and other indicators, as well as their event sequence attributes, are aggregated in a specified time window, and then the data is cleaned, the data quality is checked, and missing values ​​are removed and filled. Then, the data is standardized / normalized to avoid deviations caused by different dimensions during the training process of the model, construct feature engineering, select the features that have the greatest impact on the prediction target, and organize them into a mature and general data set. This fully utilizes the metadata of the existing monitoring information, provides training data sets for other tasks, and completes the reintegration of information.

[0113] Based on the same inventive concept, an embodiment of the present invention further provides a scaling prediction device, as Figure 3 shown. The device 300 may include:

[0114] An acquisition module 301, configured to acquire target monitoring data, where the target monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data;

[0115] A prediction module 302, configured to input the target monitoring data into a pre-trained scaling prediction model, and acquire scaling parameters output by the scaling prediction model, where the scaling parameters include a scaling type and a scaling scale;

[0116] A training module 303, configured to pre-train the scaling prediction model through the following steps:

[0117] Acquire sample monitoring data and a label corresponding to the sample monitoring data, where the sample monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data, and the label is the scaling parameter corresponding to the sample monitoring data;

[0118] Input each sample monitoring data into an initial scaling prediction model, where the initial scaling prediction model is an LSTM model;

[0119] Acquire a prediction result output by the initial scaling model, and calculate an objective difference between the prediction result and the label of the sample monitoring data;

[0120] In a case where the objective difference is higher than a preset difference threshold, adjust model parameters of the initial scaling prediction model based on the objective difference until the objective is not higher than the preset difference threshold, so as to obtain a target scaling prediction model.

[0121] In a possible embodiment, the acquiring the sample monitoring data and the label corresponding to the sample monitoring data includes:

[0122] Based on preset monitoring metrics, acquire initial monitoring data from a scaling history record, where the preset monitoring metrics include: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data;

[0123] Perform data cleaning on the initial monitoring data, and perform normalization on the cleaned monitoring data to obtain sample monitoring data; the data cleaning includes: missing value filling and outlier removal;

[0124] Determine the labels of each sample monitoring data based on the timestamp information of each sample monitoring data and the timestamp information of the scaling history record.

[0125] In a possible embodiment, the method further includes:

[0126] For the normalized sample monitoring data, generate time series data according to the timestamp information of the sample monitoring data as the target sample monitoring data;

[0127] Divide the target sample monitoring data into a training set, a validation set, and a test set.

[0128] In a possible embodiment, the initial scaling prediction model includes multiple LSTM network layers, a dropout layer, a fully connected layer, and an output layer. The method further includes:

[0129] The initial scaling prediction model extracts data features in the sample monitoring data through multiple LSTM network layers and inputs them into the dropout layer;

[0130] Process the data features through the dropout layer and input the processed data features into the fully connected layer;

[0131] Output the result vector output by the fully connected layer to the output layer, and output the prediction result through the output layer based on the result vector.

[0132] In a possible embodiment, the output layer includes a first neuron and a second neuron;

[0133] The outputting the prediction result through the output layer based on the result vector includes:

[0134] Perform binary classification on the result vector through the first neuron using the sigmod activation function and output the scaling type prediction result;

[0135] Output a continuous value through the second neuron using the linear activation function based on the result vector as the scaling scale prediction result.

[0136] In a possible embodiment, the calculating the target difference between the prediction result and the label of the sample monitoring data includes:

[0137] Calculate the binary cross-entropy between the scaling type prediction result and the scaling type included in the label of the sample monitoring data;

[0138] Calculate the mean square error or root mean square error between the predicted scaling capacity result and the scaling capacity included in the label of the sample monitoring data.

[0139] In a possible embodiment, the initial scaling prediction model further includes: an adam optimizer, and adjusting the model parameters of the initial scaling prediction model based on the target difference includes:

[0140] Updating the model parameters of the initial scaling prediction model based on the target difference through the adam optimizer.

[0141] Wherein, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the present invention all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0142] An exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.

[0143] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.

[0144] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.

[0145] Reference Figure 4 , the structural block diagram of the electronic device 400 that can be used as the server or client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0146] As Figure 4As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0147] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device capable of inputting information into the electronic device 400. The input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 407 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include, but is not limited to, magnetic disks and optical discs. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0148] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, the above-mentioned scaling prediction method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute the above-mentioned scaling prediction method in any other appropriate manner (e.g., by means of firmware).

[0149] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0150] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0152] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0153] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0154] A computer system can include clients and servers. The clients and servers are generally far apart from each other and typically interact through a communication network. The client - server relationship is created by computer programs that run on the respective computers and have a client - server relationship with each other.

Claims

1. A method for predicting expansion and contraction, characterized in that: The method comprises: Acquire target monitoring data, the target monitoring data including: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data and SWAP monitoring data; Input the target monitoring data into a pre-trained scaling prediction model, and obtain scaling parameters output by the scaling prediction model, wherein the scaling parameters include scaling type and scaling scale; The scaling prediction model is pre-trained through the following steps: Obtain sample monitoring data and labels corresponding to the sample monitoring data, wherein the sample monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data and SWAP monitoring data, and the labels are expansion and contraction parameters corresponding to the sample monitoring data; Inputting each of the sample monitoring data into an initial expansion and contraction prediction model, wherein the initial expansion and contraction prediction model is an LSTM model; Obtaining a prediction result output by the initial scaling model, and calculating a target difference between the prediction result and a label of the sample monitoring data; When the target difference is higher than a preset difference threshold, the model parameters of the initial scaling prediction model are adjusted based on the target difference until the target is no higher than the preset difference threshold, thereby obtaining a target scaling prediction model.

2. The method according to claim 1, characterized in that: The obtaining of sample monitoring data and labels corresponding to the sample monitoring data includes: Based on the preset monitoring indicators, initial monitoring data is obtained from the expansion and contraction history records, wherein the preset monitoring indicators include: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data, and SWAP monitoring data; The initial monitoring data is cleaned and the cleaned monitoring data is normalized to obtain sample monitoring data; the data cleaning includes: filling missing values ​​and removing outliers; Based on the timestamp information of each sample monitoring data and the timestamp information of the expansion and contraction history records, a label of each sample monitoring data is determined.

3. The method according to claim 2, characterized in that The method further comprises: For the normalized sample monitoring data, generating time series data according to the timestamp information of the sample monitoring data as target sample monitoring data; The target sample monitoring data is divided into a training set, a validation set and a test set.

4. The method according to claim 1, characterized in that: The initial expansion and contraction prediction model includes multiple LSTM network layers, a dropout layer, a fully connected layer and an output layer, and the method further includes: The initial expansion and contraction prediction model extracts data features from the sample monitoring data through a plurality of LSTM network layers and inputs the data features into the dropout layer; Processing the data features through the dropout layer, and inputting the processed data features into the fully connected layer; The result vector output by the fully connected layer is output to the output layer, and the prediction result is output based on the result vector through the output layer.

5. The method according to claim 4, characterized in that The output layer includes a first neuron and a second neuron; Outputting the prediction result based on the result vector through the output layer includes: Perform binary classification based on the result vector by using a sigmoid activation function through the first neuron, and output a prediction result of the expansion and contraction type; The second neuron uses a linear activation function to output a continuous value based on the result vector as a scaling prediction result.

6. The method according to claim 5, characterized in that The calculating a target difference between the prediction result and the label of the sample monitoring data includes: Calculating a binary cross entropy between the scaling type prediction result and the scaling type included in the label of the sample monitoring data; The mean square error or root mean square error between the expansion and contraction scale prediction result and the expansion and contraction scale included in the label of the sample monitoring data is calculated.

7. The method according to claim 6, characterized in that The initial expansion and contraction prediction model further includes: an adam optimizer, and the adjusting of model parameters of the initial expansion and contraction prediction model based on the target difference includes: The model parameters of the initial scaling prediction model are updated based on the target difference by the adam optimizer.

8. A device for predicting expansion and contraction, characterized in that: The device comprises: An acquisition module is used to acquire target monitoring data, wherein the target monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data and SWAP monitoring data; A prediction module, used to input the target monitoring data into a pre-trained scaling prediction model, and obtain scaling parameters output by the scaling prediction model, wherein the scaling parameters include scaling type and scaling scale; The training module is used to pre-train the scaling prediction model through the following steps: Obtain sample monitoring data and labels corresponding to the sample monitoring data, wherein the sample monitoring data includes: CPU monitoring data, memory monitoring data, disk monitoring data, network monitoring data and SWAP monitoring data, and the labels are expansion and contraction parameters corresponding to the sample monitoring data; Inputting each of the sample monitoring data into an initial expansion and contraction prediction model, wherein the initial expansion and contraction prediction model is an LSTM model; Obtaining a prediction result output by the initial scaling model, and calculating a target difference between the prediction result and a label of the sample monitoring data; When the target difference is higher than a preset difference threshold, the model parameters of the initial scaling prediction model are adjusted based on the target difference until the target is no higher than the preset difference threshold, thereby obtaining a target scaling prediction model.

9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-7.