Cluster service level automatic scaling method and device, computer equipment, medium and product
Through variational modal decomposition and neural network prediction technology, the problem of traditional HPA response lag is solved, and the number of cluster service replicas is accurately adjusted and resource utilization is improved.
Patent Information
- Application Number
- CN202510641017.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The static threshold triggering mechanism of traditional Kubernetes native HPA has a response lag, making it difficult to adjust the number of service instances in a timely manner, resulting in service overload or resource waste.
The service indicator data of the cluster is obtained through the exporter, and the variational modal decomposition is performed to obtain trend terms, periodic terms and random terms. These components are predicted using the neural network model, and the number of service replicas is adjusted according to the prediction results.
The number of cluster service replicas has been accurately adjusted, resource utilization has been improved, and adaptive expansion of service pressure and high availability of services has been ensured.
Smart Images

Figure CN120179417A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of containerization technology, and particularly to a method, device, computer device, medium, and product for automatically scaling the cluster service level. Background Art
[0002] With the popularization of cloud computing and microservices architecture, containerization technology has become the mainstream way of modern application deployment. As the current mainstream container orchestration system, the container orchestration engine (Kubernetes) dynamically adjusts the number of service instances through the Horizontal Pod Autoscaler (HPA) mechanism to meet the ever-changing workload requirements. HPA monitors key metrics (such as CPU and memory usage) and compares them with preset target values, and automatically triggers scaling operations to ensure the high availability and resource utilization of the service.
[0003] In traditional methods, the native HPA of Kubernetes adopts a static threshold triggering mechanism. By continuously collecting the resource usage metrics of Pods and comparing them with the static thresholds preset by users, when the metrics exceed or are lower than the thresholds, HPA sends requests to the application programming interface server (API Server) to increase or decrease the number of Pod replicas. However, due to the inherent response lag of the static threshold mechanism, it is difficult to make timely adjustments in case of sudden load changes, resulting in service overload or resource waste. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer device, medium, and product for automatically scaling the cluster service level that can improve resource utilization in response to the above technical problems.
[0005] In a first aspect, this application provides a method for automatically scaling the cluster service level, including:
[0006] During the process of performing instantiation operations through an exporter, obtain the service metric data of the cluster; the metric data includes CPU data and memory data;
[0007] Perform variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors;
[0008] Predict the target components through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are skip-connected;
[0009] Adjust the number of service replicas in the cluster according to the target prediction result to scale the service in the cluster up or down.
[0010] In one embodiment, the steps of performing variational mode decomposition on the service metric data to obtain the target component include:
[0011] Establish a variational constraint model; the variational constraint model includes an objective function and a constraint function;
[0012] Update the objective function according to the service metric data;
[0013] Based on the updated objective function, perform iterative solution on the variational constraint model until the constraint function meets the preset condition and stop. According to the current solution of the variational constraint model, obtain the target component.
[0014] In one embodiment, the steps of predicting the target component through a neural network model to obtain the target prediction result include:
[0015] Process the target component through the network layer of the neural network model to obtain the network layer hidden state parameters;
[0016] Process the network layer hidden state parameters through the self-attention layer of the neural network model to obtain the self-attention layer state parameters;
[0017] Use the target component and the self-attention layer state parameters as the input data of the fully connected layer of the neural network model, and obtain the target prediction result according to the output data of the fully connected layer.
[0018] In one embodiment, the steps of obtaining the target prediction result according to the output data of the fully connected layer include:
[0019] Process the target component and the self-attention layer state parameters through the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to the trend term, a second prediction result corresponding to the periodic term, and a third prediction result corresponding to the random term;
[0020] Perform time series addition processing on the first prediction result, the second prediction result, and the third prediction result to obtain the target prediction result.
[0021] In one embodiment, the steps of adjusting the number of service replicas in the cluster according to the target prediction result include:
[0022] Publish the target prediction result to the cluster in a preset format through a preset interface, so that the horizontal auto-scaler instance can obtain the target prediction result through the cluster and adjust the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
[0023] In one embodiment, the method further includes:
[0024] Obtaining sample components according to the sample service metric data of the cluster service, and training a neural network model based on the sample components;
[0025] When the training of the neural network model is completed, deploying the neural network model to the cluster to predict the target component through the neural network model to obtain a target prediction result.
[0026] In a second aspect, the present application also provides a cluster service level automatic scaling device, including:
[0027] A data acquisition module, configured to acquire the service metric data of the cluster during the process of performing the instantiation operation through an exporter; the metric data includes CPU data and memory data;
[0028] A variational mode decomposition module, configured to perform variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors;
[0029] A model prediction module, configured to predict the target component through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are jump-connected;
[0030] A replica adjustment module, configured to adjust the number of service replicas in the cluster according to the target prediction result to expand or contract the services in the cluster.
[0031] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the method steps of any one of the first aspect are implemented.
[0032] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method steps of any one of the first aspect are implemented.
[0033] In a fifth aspect, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method steps of any one of the first aspect are implemented.
[0034] The above cluster service level automatic scaling method, device, computer device, medium and product obtain service metric data of the cluster during the instantiation operation executed by the exporter, perform variational mode decomposition on the service metric data to obtain target components, predict the target components through a neural network model to obtain a target prediction result, and adjust the number of service replicas in the cluster according to the target prediction result to scale the service in the cluster up or down, which can accurately adjust the number of service replicas in the cluster, improve resource utilization rate, thereby realizing adaptive scaling of service pressure and ensuring high availability of the service and stability of the cluster. Brief Description of the Drawings
[0035] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is an application environment diagram of the cluster service level automatic scaling method in one embodiment;
[0037] Figure 2 It is a flowchart of the cluster service level automatic scaling method in one embodiment;
[0038] Figure 3 It is a schematic diagram of the target component in one embodiment;
[0039] Figure 4 It is a schematic diagram of the structure of the neural network model in one embodiment;
[0040] Figure 5 It is a flowchart of the cluster service level automatic scaling method in another embodiment;
[0041] Figure 6 It is a structural block diagram of the cluster service level automatic scaling device in one embodiment;
[0042] Figure 7 It is an internal structure diagram of a computer device in one embodiment. Detailed Embodiments
[0043] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] The cluster service level automatic scaling method provided by the embodiments of the present application can be applied to such asFigure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. Among them, the terminal 102 is used to obtain the service metric data of the cluster during the process of performing the instantiation operation through the exporter, perform variational mode decomposition on the service metric data to obtain the target component, predict the target component through the neural network model to obtain the target prediction result, and adjust the number of service replicas in the cluster according to the target prediction result to expand or contract the services in the cluster. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0045] In an exemplary embodiment, as Figure 2 shown, a method for automatically scaling the cluster service level is provided. Taking the terminal 102 in Figure 1 as an example, the method includes the following steps 202 to 208. Among them:
[0046] S202: Obtain the service metric data of the cluster during the process of performing the instantiation operation through the exporter; the metric data includes CPU data and memory data.
[0047] Optionally, during the process of performing the instantiation operation through the exporter, collect the relevant metric data of the services in the cluster, including CPU data and memory data, which reflect the service running status and resource usage. Among them, for the GPU devices in the cluster, the metric data also includes GPU video memory, etc.
[0048] Exemplarily, the Exportor collects the service CPU and memory metrics respectively by instantiating the Target, and stores the corresponding metric data in a polling manner. Among them, the Exportor is a program used to collect metric data of specific applications, services, or system components. The Target represents the specific object to be monitored. Instantiating the Target means creating a corresponding instance for each object to be monitored, enabling the Exportor to clearly know where to collect the metric data. In this process, the Exportor will collect data for the two key metrics of the service, namely CPU and memory, respectively.
[0049] S204: Perform variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors.
[0050] Optionally, through variational mode decomposition, the service metric data is decomposed into different target components, including a trend term, a periodic term, and a random term. Among them, the trend term can reflect the overall change trend of the service metric data over time; the periodic term can reflect the internal connection between the service metric data and the service period; the random term reflects the situation where the service metric data is affected by external uncertain factors or other unplanned events.
[0051] S206: Predict the target components through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are skip-connected.
[0052] Optionally, a neural network model with skip connections between the input layer and the output layer is used to predict the decomposed target components. This skip connection structure helps the model better capture long-term dependencies and complex patterns in the data, solve the "degradation" problem of the network model, and at the same time achieve the fusion of data features at different scales. By predicting the trend term, the periodic term, and the random term, an estimate of the future service metric data, that is, the target prediction result, can be obtained.
[0053] S208: Adjust the number of service replicas in the cluster according to the target prediction result to scale the service in the cluster up or down.
[0054] Optionally, adjust the number of service replicas in the cluster according to the target prediction result. When the prediction result shows that the load of the service may increase, increase the number of service replicas to improve the processing capacity of the service and meet higher request volumes; when the load is expected to decrease, reduce the number of service replicas to avoid waste of resources.
[0055] In the above method for automatically scaling the cluster service level, during the process of instantiating through the exporter, service metric data of the cluster is obtained, variational mode decomposition is performed on the service metric data to obtain target components, the target components are predicted through a neural network model to obtain a target prediction result, and the number of service replicas in the cluster is adjusted according to the target prediction result to expand or contract the service in the cluster, which can accurately adjust the number of service replicas in the cluster, improve resource utilization rate, and thus achieve self-adaptive expansion of service pressure, ensuring high availability of the service and stability of the cluster.
[0056] In an exemplary embodiment, the step of performing variational mode decomposition on the service metric data to obtain target components includes: establishing a variational constraint model; the variational constraint model includes an objective function and a constraint function; updating the objective function according to the service metric data; based on the updated objective function, iteratively solving the variational constraint model until the constraint function satisfies a preset condition and then stopping, and obtaining the target components according to the current solution of the variational constraint model.
[0057] Optionally, for CPU metric data, when using the variational mode decomposition algorithm to perform time series signal decomposition, the goal is to decompose the data into different components to reflect different characteristics. In the variational constraint model, the objective function is usually related to decomposing the signal into each component and minimizing a certain error or cost, so as to decompose the CPU metric data into components such as a trend term, a periodic term, and a random term. Among them, the constraint function is a limiting condition for the decomposition process. In variational mode decomposition, the constraint condition ensures the rationality and stability of the decomposition process. Taking CPU metric data as an example, since the objective function is related to the specific signal being decomposed, when actual CPU metric data is available, the objective function will be adjusted accordingly according to these data, making the objective function more suitable for the current service metric data for more accurate decomposition. Based on the updated objective function, the variational constraint model is iteratively solved until the constraint function satisfies the preset condition and stops. In each iteration, the component and its central frequency are alternately updated. When the iteration stops, the current solution of the variational constraint model at this time corresponds to each decomposed target component, that is, the trend term, the periodic term, and the random term. Through this process, the original CPU metric data is successfully decomposed into these three components that can reflect different characteristics.
[0058] Exemplarily, as Figure 3 shown, using the Variational Mode Decomposition (VMD) algorithm to perform time series signal decomposition, a trend term reflecting the long-term system change of the service, a periodic term reflecting the associated change with the service cycle such as a scheduled task, and a random term reflecting the influence of external factors or other unplanned events are obtained respectively. Taking CPU data Taking it as an example for illustration, if it is decomposed into three components: a trend term, a periodic term, and a random term, a variational constraint equation can be constructed as follows:
[0059]
[0060] Among them, and are the k-th component and its central frequency respectively. Among them, .
[0061] Introduce the Lagrange multiplier operator Transform the above constrained problem into an unconstrained problem, and the expression is:
[0062]
[0063] Introduce the alternating direction multiplier method and the Fourier equidistant transformation to solve until convergence. Finally, the trend term, periodic term, and random term components of each index are obtained, and the expression is:
[0064]
[0065] Among them, the component and its central frequency are updated alternately until acceptance, and the trend term, periodic term, and random term components of the CPU index data are obtained.
[0066] In this embodiment, by establishing a variational constraint model, according to the service index data, the objective function is updated. Based on the updated objective function, the variational constraint model is iteratively solved until the constraint function meets the preset conditions and stops. According to the current solution of the variational constraint model, the target component is obtained, and the change trend of the service index data can be accurately obtained, so as to accurately adjust the number of service replicas in the cluster and improve resource utilization.
[0067] In an exemplary embodiment, the steps of predicting the target component through a neural network model to obtain a target prediction result include: processing the target component through the network layer of the neural network model to obtain the network layer hidden state parameters; processing the network layer hidden state parameters through the self-attention layer of the neural network model to obtain the self-attention layer state parameters; using the target component and the self-attention layer state parameters as the input data of the fully connected layer of the neural network model, and obtaining the target prediction result according to the output data of the fully connected layer.
[0068] Optionally, the network layer of the neural network model usually contains multiple neurons. The target component is passed as input data to the network layer. Each neuron will perform a weighted sum on the input and perform a non-linear transformation through an activation function, so as to extract the feature information in the target component. After being processed by the network layer, the hidden state parameters of the network layer are obtained. These parameters contain the feature representation of the target component after being processed by one or more layers of neurons. The role of the self-attention layer is to further analyze and weight the hidden state parameters of the network layer. It will calculate the correlation between the hidden state parameters at each position and those at other positions, so as to assign a weight to each position. This weight reflects the importance of this position in the entire sequence. In this way, the self-attention layer can highlight key information and suppress unimportant information, and obtain the self-attention layer state parameters. After that, the target component and the self-attention layer state parameters are fused and further processed through the fully connected layer. Each neuron in the fully connected layer is connected to all the features of the input. It will perform a linear combination on the input and perform a non-linear transformation through the activation function to obtain the final output data.
[0069] Exemplarily, taking the neural network model as the Long Short-Term Memory (LSTM) model as an example for illustration, the attention mechanism is introduced into the traditional LSTM model to calculate the similarity between the current state and all historical states, improving the long-term memory ability of the model. At the same time, a skip connection is established between the current input and the output of the self-attention layer to solve the "degradation" problem of the network model while realizing the fusion of data features at different scales. The structure of the Skip Connection LSTM Self-attention (SC-LSTM-SA) neural network model is as Figure 4 shown. For the SC-LSTM-SA model at time t, each index component ( , ) is trained and learned through the LSTM network layer to obtain the hidden state of the LSTM network layer. To capture the long-term change features of each index component, the hidden state of the network layer is input into the self-attention layer to obtain the self-attention layer state containing the long-term similarity information of the component sequence. At the same time, to realize the fusion of data at different scales, the index component and the self-attention layer state are used as the input of the fully connected layer, and finally the prediction result of the model is obtained, realizing the long-term multi-scale feature fusion of each index data.
[0070] The overall calculation of the model can be expressed as:
[0071]
[0072] Among them, LSTM is a model network constructed based on LSTM units, SA is a self-attention layer, and FC is a fully connected layer.
[0073] Since the VMD frequency division result satisfies the time series addition model, adding the prediction results of the index components can obtain the prediction result of the overall index, which can be expressed as:
[0074]
[0075] Among them, , , and n is the length of the index sequence.
[0076] In this embodiment, the target component is processed through the network layer of the neural network model to obtain the hidden state parameters of the network layer. The hidden state parameters of the network layer are processed through the self-attention layer of the neural network model to obtain the state parameters of the self-attention layer. The target component and the state parameters of the self-attention layer are used as the input data of the fully connected layer of the neural network model, and the target prediction result is obtained according to the output data of the fully connected layer, which can better capture the complex features and relationships in the target component, thereby improving the prediction accuracy.
[0077] In an exemplary embodiment, the step of obtaining the target prediction result according to the output data of the fully connected layer includes: processing the target component and the state parameters of the self-attention layer through the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to the trend term, a second prediction result corresponding to the periodic term, and a third prediction result corresponding to the random term; performing time series addition processing on the first prediction result, the second prediction result, and the third prediction result to obtain the target prediction result.
[0078] Optionally, the fully connected layer will comprehensively process the input target component and the state parameters of the self-attention layer to obtain a reference prediction result. Among them, the reference prediction result is a vector containing multiple parts, corresponding to the prediction results of the trend term, the periodic term, and the random term respectively, that is, the first prediction result, the second prediction result, and the third prediction result. Since the trend term, the periodic term, and the random term are different components obtained by performing variational mode decomposition on service index data (such as CPU index data), they respectively reflect the characteristics and change laws of the data from different angles. When predicting, combining the prediction results of these three components through time series addition can comprehensively consider factors such as the long-term trend, periodic changes, and random fluctuations of the data, thereby obtaining a more comprehensive and accurate target prediction result to reflect the overall future change situation of the service index data.
[0079] In this embodiment, the fully connected layer of the neural network model processes the target component and the state parameters of the self-attention layer to obtain a first prediction result corresponding to the trend term, a second prediction result corresponding to the periodic term, and a third prediction result corresponding to the random term. By performing sequential addition processing on the first prediction result, the second prediction result, and the third prediction result, the target prediction result is obtained, which can accurately capture the changing trend of resource requirements, thereby precisely adjusting the number of service replicas in the cluster and improving resource utilization rate.
[0080] In an exemplary embodiment, the step of adjusting the number of service replicas in the cluster according to the target prediction result includes: publishing the target prediction result to the cluster in a preset format through a preset interface, so that the horizontal pod autoscaler instance can obtain the target prediction result through the cluster and adjust the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
[0081] Optionally, the target prediction result contains important information such as the predicted service load of the cluster. To ensure that each component in the cluster can correctly receive and process this information, it is necessary to package and publish the target prediction result in a preset format. The horizontal pod autoscaler (HPA) instance in the cluster continuously monitors relevant data in the cluster. When the target prediction result is published to the cluster, the HPA instance can obtain this information through its connection channel with the cluster. The HPA instance will parse data such as the predicted service load in the target prediction result, and then decide whether to adjust the number of service replicas according to a preset algorithm. If the predicted service load indicates that the load is about to be too high, that is, the resource requirements of the system are expected to exceed the processing capacity provided by the current service replicas, the HPA instance will send an instruction to the corresponding deployment to increase the number of Pod replicas, so that more Pods can participate in service processing, thereby improving the overall processing capacity of the system to cope with the upcoming high load. If the predicted service load indicates that the current number of service replicas may exceed the actual demand, that is, there is idle or wasted system resources, the HPA instance will instruct the deployment to reduce the number of Pod replicas. In this way, under the premise of ensuring the smoothness and high availability of service operation, the resource usage efficiency can be optimized, and excessive occupation and waste of resources can be avoided.
[0082] In this embodiment, by publishing the target prediction result to the cluster in a preset format through a preset interface, the horizontal pod autoscaler instance can obtain the target prediction result through the cluster and adjust the number of service replicas according to the target prediction result, which can precisely adjust the number of service replicas in the cluster, improve resource utilization rate, thereby realizing adaptive expansion of service pressure and ensuring the high availability of the service and the stability of the cluster.
[0083] In an exemplary embodiment, the method further includes: obtaining sample components according to the sample service metric data of the cluster service, and training a neural network model based on the sample components; when the training of the neural network model is completed, deploying the neural network model to the cluster to predict the target components through the neural network model to obtain a target prediction result.
[0084] Optionally, by collecting service metric data from the cluster as samples, processing the sample service metric data using methods such as variational mode decomposition, decomposing it into different sample components, training the neural network model based on the sample components, and when the training of the neural network model is completed, deploying the neural network model to the cluster so that it can run in an actual production environment, thereby predicting the target components through the neural network model to obtain a target prediction result.
[0085] In this embodiment, by obtaining sample components according to the sample service metric data of the cluster service, training the neural network model based on the sample components, and when the training of the neural network model is completed, deploying the neural network model to the cluster to predict the target components through the neural network model to obtain a target prediction result, it is possible to accurately adjust the number of service replicas in the cluster, improve resource utilization, thereby realizing the adaptive expansion of service pressure and ensuring the high availability of the service and the stability of the cluster.
[0086] In an exemplary embodiment, as Figure 5 shown, a method for automatically scaling the cluster service level is provided, and the method includes the following steps:
[0087] During the process of performing instantiation operations through the exporter, obtain the service metric data of the cluster; the metric data includes CPU data and memory data.
[0088] Establish a variational constraint model; the variational constraint model includes an objective function and a constraint function; update the objective function according to the service metric data; based on the updated objective function, perform iterative solution on the variational constraint model until the constraint function meets the preset conditions and stop, and obtain the target components according to the current solution of the variational constraint model.
[0089] Among them, the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors.
[0090] Obtain sample components according to the sample service metric data of the cluster service, and train a neural network model based on the sample components; when the training of the neural network model is completed, deploy the neural network model to the cluster.
[0091] Process the target component through the network layer of the neural network model to obtain the network layer hidden state parameters; process the network layer hidden state parameters through the self-attention layer of the neural network model to obtain the self-attention layer state parameters; use the target component and the self-attention layer state parameters as the input data of the fully connected layer of the neural network model, and process the target component and the self-attention layer state parameters through the fully connected layer of the neural network model to obtain the reference prediction result; the reference prediction result includes a first prediction result corresponding to the trend term, a second prediction result corresponding to the periodic term, and a third prediction result corresponding to the random term; perform a time series addition process on the first prediction result, the second prediction result, and the third prediction result to obtain the target prediction result.
[0092] Among them, the input layer and the output layer of the neural network model are jump-connected.
[0093] Publish the target prediction result to the cluster in a preset format through a preset interface, so that the horizontal auto-scaler instance can obtain the target prediction result through the cluster, and adjust the number of service replicas according to the target prediction result to expand or contract the services in the cluster.
[0094] Among them, the target prediction result includes the predicted service load of the cluster.
[0095] In this embodiment, during the process of performing the instantiation operation through the exporter, obtain the service metric data of the cluster, perform variational mode decomposition on the service metric data to obtain the target component, predict the target component through the neural network model to obtain the target prediction result, and adjust the number of service replicas in the cluster according to the target prediction result to expand or contract the services in the cluster, which can accurately adjust the number of service replicas in the cluster, improve resource utilization, and thus achieve the adaptive expansion of service pressure and ensure the high availability of the service and the stability of the cluster.
[0096] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0097] Based on the same inventive concept, an embodiment of the present application further provides an automatic scaling device for cluster service level to implement the automatic scaling method for cluster service level involved above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the automatic scaling device for cluster service level provided below can refer to the limitations on the automatic scaling method for cluster service level in the above text, and will not be elaborated here.
[0098] In an exemplary embodiment, as Figure 6 shown, an automatic scaling device for cluster service level is provided, including: a data acquisition module 10, a modal decomposition module 20, a model prediction module 30, and a replica adjustment module 40, where:
[0099] The data acquisition module 10 is configured to obtain service metric data of the cluster during the instantiation operation performed by the exporter; the metric data includes CPU data and memory data.
[0100] The modal decomposition module 20 is configured to perform variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors.
[0101] The model prediction module 30 is configured to predict the target components through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are jump-connected.
[0102] The replica adjustment module 40 is configured to adjust the number of service replicas in the cluster according to the target prediction result to expand or contract the service in the cluster.
[0103] In an exemplary embodiment, the modal decomposition module 20 is further configured to establish a variational constraint model; the variational constraint model includes an objective function and a constraint function; update the objective function according to the service metric data; based on the updated objective function, perform iterative solution on the variational constraint model until the constraint function meets the preset condition and stop, and obtain the target components according to the current solution of the variational constraint model.
[0104] In an exemplary embodiment, the model prediction module 30 is further configured to process the target components through the network layer of the neural network model to obtain network layer hidden state parameters; process the network layer hidden state parameters through the self-attention layer of the neural network model to obtain self-attention layer state parameters; use the target components and the self-attention layer state parameters as the input data of the fully connected layer of the neural network model, and obtain the target prediction result according to the output data of the fully connected layer.
[0105] In an exemplary embodiment, the model prediction module 30 is further configured to process the target component and the self-attention layer state parameters through the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to the trend term, a second prediction result corresponding to the periodic term, and a third prediction result corresponding to the random term; perform a temporal addition process on the first prediction result, the second prediction result, and the third prediction result to obtain a target prediction result.
[0106] In an exemplary embodiment, the replica adjustment module 40 is further configured to publish the target prediction result to the cluster in a preset format through a preset interface, so that the horizontal auto-scaler instance can obtain the target prediction result through the cluster and adjust the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
[0107] In an exemplary embodiment, the model prediction module 30 is further configured to obtain sample components according to the sample service metric data of the cluster service, and train the neural network model based on the sample components; when the neural network model training is completed, deploy the neural network model to the cluster to predict the target component through the neural network model to obtain a target prediction result.
[0108] Each module in the above cluster service level auto-scaling device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0109] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for automatically scaling the cluster service level. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, a touchpad, or a mouse, etc.
[0110] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0111] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: During the process of performing instantiation operations through an exporter, obtain the service metric data of the cluster; the metric data includes CPU data and memory data; perform variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors; predict the target components through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are jump-connected; adjust the number of service replicas in the cluster according to the target prediction result to expand or contract the services in the cluster.
[0112] In one embodiment, when the processor executes a computer program, variational mode decomposition is performed on service metric data to obtain target components, including: establishing a variational constraint model; the variational constraint model includes an objective function and a constraint function; updating the objective function according to the service metric data; based on the updated objective function, iteratively solving the variational constraint model until the constraint function satisfies a preset condition and then stopping, and obtaining the target components according to the current solution of the variational constraint model.
[0113] In one embodiment, when the processor executes a computer program, the target components are predicted through a neural network model to obtain a target prediction result, including: processing the target components through the network layer of the neural network model to obtain network layer hidden state parameters; processing the network layer hidden state parameters through the self-attention layer of the neural network model to obtain self-attention layer state parameters; using the target components and the self-attention layer state parameters as input data for the fully connected layer of the neural network model, and obtaining the target prediction result according to the output data of the fully connected layer.
[0114] In one embodiment, obtaining the target prediction result according to the output data of the fully connected layer includes: processing the target components and the self-attention layer state parameters through the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to a trend term, a second prediction result corresponding to a periodic term, and a third prediction result corresponding to a random term; performing time series addition processing on the first prediction result, the second prediction result, and the third prediction result to obtain the target prediction result.
[0115] In one embodiment, when the processor executes a computer program, adjusting the number of service replicas in the cluster according to the target prediction result includes: publishing the target prediction result in a preset format to the cluster through a preset interface, so that the horizontal auto-scaler instance obtains the target prediction result through the cluster and adjusts the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
[0116] In one embodiment, when the processor executes a computer program, the following steps are further implemented: obtaining sample components according to the sample service metric data of the cluster service, and training the neural network model based on the sample components; when the neural network model training is completed, deploying the neural network model to the cluster to predict the target components through the neural network model to obtain the target prediction result.
[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: during the process of performing instantiation operations through an exporter, obtain service metric data of a cluster; the metric data includes CPU data and memory data; perform variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors; predict the target components through a neural network model to obtain a target prediction result; there is a skip connection between the input layer and the output layer of the neural network model; adjust the number of service replicas in the cluster according to the target prediction result to scale out or scale in the services in the cluster.
[0118] In one embodiment, the variational mode decomposition of the service metric data to obtain target components, which is involved when the computer program is executed by a processor, includes: establishing a variational constraint model; the variational constraint model includes an objective function and a constraint function; update the objective function according to the service metric data; based on the updated objective function, perform iterative solution on the variational constraint model until the constraint function meets a preset condition and then stop, and obtain the target components according to the current solution of the variational constraint model.
[0119] In one embodiment, the prediction of the target components through a neural network model to obtain a target prediction result, which is involved when the computer program is executed by a processor, includes: process the target components through the network layer of the neural network model to obtain network layer hidden state parameters; process the network layer hidden state parameters through the self-attention layer of the neural network model to obtain self-attention layer state parameters; use the target components and the self-attention layer state parameters as input data for the fully connected layer of the neural network model, and obtain the target prediction result according to the output data of the fully connected layer.
[0120] In one embodiment, the obtaining of the target prediction result according to the output data of the fully connected layer, which is involved when the computer program is executed by a processor, includes: process the target components and the self-attention layer state parameters through the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to the trend term, a second prediction result corresponding to the periodic term, and a third prediction result corresponding to the random term; perform temporal addition processing on the first prediction result, the second prediction result, and the third prediction result to obtain the target prediction result.
[0121] In one embodiment, when the computer program is executed by a processor, adjusting the number of service replicas in a cluster according to a target prediction result includes: publishing the target prediction result to the cluster in a preset format through a preset interface, so that an instance of a horizontal auto-scaler obtains the target prediction result through the cluster and adjusts the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
[0122] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining sample components according to the sample service metric data of the cluster service, and training a neural network model based on the sample components; when the neural network model training is completed, deploying the neural network model to the cluster to predict target components through the neural network model to obtain a target prediction result.
[0123] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the following steps: obtaining service metric data of a cluster during the process of performing an instantiation operation through an exporter; the metric data includes CPU data and memory data; performing variational mode decomposition on the service metric data to obtain target components; the target components include a trend term, a periodic term, and a random term; the trend term is used to characterize the change trend of the service metric data, the periodic term is used to characterize the correlation between the service metric data and the service period, and the random term is used to characterize the correlation between the service metric data and external influencing factors; predicting the target components through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are skip-connected; adjusting the number of service replicas in the cluster according to the target prediction result to expand or contract the services in the cluster.
[0124] In one embodiment, when the computer program is executed by a processor, performing variational mode decomposition on service metric data to obtain target components includes: establishing a variational constraint model; the variational constraint model includes an objective function and a constraint function; updating the objective function according to the service metric data; iteratively solving the variational constraint model based on the updated objective function until the constraint function meets a preset condition and then stopping, and obtaining the target components according to the current solution of the variational constraint model.
[0125] In one embodiment, when the computer program is executed by a processor, predicting the target components through a neural network model to obtain a target prediction result includes: processing the target components through the network layer of the neural network model to obtain network layer hidden state parameters; processing the network layer hidden state parameters through the self-attention layer of the neural network model to obtain self-attention layer state parameters; using the target components and the self-attention layer state parameters as input data for the fully connected layer of the neural network model, and obtaining the target prediction result according to the output data of the fully connected layer.
[0126] In one embodiment, obtaining a target prediction result based on the output data of the fully connected layer when the computer program is executed by a processor includes: processing a target component and self-attention layer state parameters through the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to a trend term, a second prediction result corresponding to a periodic term, and a third prediction result corresponding to a random term; performing a temporal addition process on the first prediction result, the second prediction result, and the third prediction result to obtain the target prediction result.
[0127] In one embodiment, adjusting the number of service replicas in a cluster based on the target prediction result when the computer program is executed by a processor includes: publishing the target prediction result to the cluster in a preset format through a preset interface, so that the horizontal autoscaler instance obtains the target prediction result through the cluster and adjusts the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
[0128] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining a sample component according to the sample service metric data of the cluster service, and training the neural network model based on the sample component; when the training of the neural network model is completed, deploying the neural network model to the cluster to predict the target component through the neural network model to obtain the target prediction result.
[0129] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0131] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A cluster service level automatic scaling method, characterized in that: The method comprises: In the process of executing the instantiation operation through the exporter, the service indicator data of the cluster is obtained; the indicator data includes CPU data and memory data; Performing variational mode decomposition on the service indicator data to obtain a target component; the target component includes a trend item, a period item and a random item; the trend item is used to characterize the change trend of the service indicator data, the period item is used to characterize the correlation between the service indicator data and the service period, and the random item is used to characterize the correlation between the service indicator data and the external influencing factors; The target component is predicted by a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are jump-connected; The number of service replicas in the cluster is adjusted according to the target prediction result to expand or reduce the capacity of the services in the cluster.
2. The method according to claim 1, characterized in that The performing variational mode decomposition on the service indicator data to obtain the target component includes: Establishing a variational constraint model; the variational constraint model includes an objective function and a constraint function; According to the service indicator data, updating the objective function; Based on the updated objective function, the variational constraint model is iteratively solved until the constraint function meets a preset condition, and the target component is obtained according to the current solution of the variational constraint model.
3. The method according to claim 1, characterized in that The method of predicting the target component by using a neural network model to obtain a target prediction result includes: Processing the target component through the network layer of the neural network model to obtain the hidden state parameters of the network layer; Processing the hidden state parameters of the network layer through the self-attention layer of the neural network model to obtain the self-attention layer state parameters; The target component and the self-attention layer state parameters are used as input data of the fully connected layer of the neural network model, and the target prediction result is obtained according to the output data of the fully connected layer.
4. The method according to claim 3, characterized in that: The obtaining the target prediction result according to the output data of the fully connected layer includes: The target component and the state parameter of the self-attention layer are processed by the fully connected layer of the neural network model to obtain a reference prediction result; the reference prediction result includes a first prediction result corresponding to the trend item, a second prediction result corresponding to the period item, and a third prediction result corresponding to the random item; The first prediction result, the second prediction result and the third prediction result are subjected to time series addition processing to obtain a target prediction result.
5. The method according to claim 1, characterized in that The adjusting the number of service replicas in the cluster according to the target prediction result includes: The target prediction result is published to the cluster in a preset format through a preset interface, so that the horizontal autoscaler instance obtains the target prediction result through the cluster and adjusts the number of service replicas according to the target prediction result; the target prediction result includes the predicted service load of the cluster.
6. The method according to claim 1, characterized in that The method further comprises: Acquire sample components according to the sample service indicator data of the cluster service, and train the neural network model based on the sample components; When the neural network model is trained, the neural network model is deployed to the cluster to predict the target component through the neural network model to obtain a target prediction result.
7. A cluster service level automatic scaling device, characterized in that: The device comprises: A data acquisition module, used to acquire service indicator data of the cluster during the process of executing the instantiation operation through the exporter; the indicator data includes CPU data and memory data; A modal decomposition module, used to perform variational modal decomposition on the service indicator data to obtain a target component; the target component includes a trend item, a period item and a random item; the trend item is used to characterize the change trend of the service indicator data, the period item is used to characterize the correlation between the service indicator data and the service period, and the random item is used to characterize the correlation between the service indicator data and external influencing factors; A model prediction module, used to predict the target component through a neural network model to obtain a target prediction result; the input layer and the output layer of the neural network model are jump-connected; The replica adjustment module is used to adjust the number of service replicas in the cluster according to the target prediction result to expand or reduce the capacity of the services in the cluster.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Load prediction type elastic telescoping system and method based on Kubernetes
CN116610416A
Resource management method and device, computer equipment and storage medium
CN117909076A
Automatic operation and maintenance method based on container and big data
CN117971384A
Automatic scaling method and system based on mobile edge computing server cluster
CN119383664A