Model reasoning service elastic load prediction method, device and medium

By building a transformer-based timing prediction model on the Kubernetes platform, processing and predicting the load data of model inference services, the shortcomings of traditional load prediction methods when processing complex timing data are solved, and more accurate load prediction and intelligent resource scheduling are achieved.

CN119938481APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510121754.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional load prediction methods cannot effectively handle nonlinear relationships and multi-level features in complex timing data, resulting in inaccurate load prediction and increasing manual intervention and operation and maintenance costs.

Method used

Using the model inference service based on the Kubernetes platform, Prometheus collects operation data, cleanses and organizes data, builds a transformer-based timing prediction model, generates load predictions for preset time periods, and delivers the prediction results to the HPA of the Kubernetes platform to adjust the load count.

Benefits of technology

It improves the accuracy and reliability of load prediction, reduces manual intervention and operation and maintenance costs, and realizes more intelligent resource scheduling and flexible load management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938481A_ABST
    Figure CN119938481A_ABST
Patent Text Reader

Abstract

The invention discloses a model reasoning service elastic load prediction method and device and a medium, and belongs to the technical field of elastic load prediction. The method comprises the following steps: acquiring operation data of a model reasoning service based on Prometheus of the Kubernetes platform; cleaning and sorting the operation data to generate standard operation data; a time sequence prediction model based on transform is constructed; processing the standard operation data based on the time sequence prediction model to generate load prediction of a preset time period; and the load prediction in a set time period is transmitted to the HPA of the Kubernetes platform, and the load number of the model reasoning service is adjusted based on the HPA. According to the method, the technical effects of more accurately predicting load change, reducing manual intervention and operation and maintenance cost and more intelligently scheduling resources are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of elastic load prediction, and in particular to a method, device and medium for elastic load prediction of a model reasoning service. Background Art

[0002] Kubernetes (K8s for short) is an open source container orchestration platform for automating the deployment, expansion, and management of containerized applications. Kubernetes has elastic scaling and self-healing functions, that is, Kubernetes can automatically adjust the number of containers according to the application load to achieve elastic scaling. It can also achieve self-healing of applications by automatically restarting or replicating containers.

[0003] In Kubernetes, HPA (Horizontal Pod Autoscaler) automatically adjusts the number of Pod copies according to the current load situation. HPA is an important part of Kubernetes' resource orchestration and automatic management functions, and aims to optimize resource utilization and application performance by automatically adjusting the number of Pod copies. During peak loads, HPA can automatically increase the number of Pod copies to improve processing power; during low loads, HPA can reduce the number of Pod copies to save resources. However, traditional load prediction methods often rely on rules or simple statistical models, and cannot handle nonlinear relationships and multi-level features in complex time series data.

[0004] Therefore, how to more accurately predict load changes, reduce manual intervention and operation and maintenance costs, and achieve more intelligent resource scheduling has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The embodiments of the present application provide a model reasoning service elastic load prediction method, device and medium to solve the following technical problems: how to more accurately predict load changes, reduce manual intervention and operation and maintenance costs, and achieve more intelligent resource scheduling.

[0006] In the first aspect, an embodiment of the present application provides a model inference service elastic load prediction method, which is applied to the model inference service of the Kubernetes platform, and the method includes: collecting operation data of the model inference service based on Prometheus of the Kubernetes platform; wherein the operation data includes CPU utilization, memory utilization, request processing time, request success rate and network traffic; cleaning and organizing the operation data to generate standard operation data; constructing a transformer-based timing prediction model; processing the standard operation data based on the timing prediction model to generate a load prediction for a preset time period; transmitting the load prediction for the preset time period to the HPA of the Kubernetes platform, and adjusting the load quantity of the model inference service based on the HPA.

[0007] In one implementation of the present application, Prometheus based on the Kubernetes platform collects the operating data of the model reasoning service, specifically including: configuring Prometheus; wherein Prometheus is used to monitor the Pods and nodes of the model reasoning service; based on the ServiceMonitor in Prometheus, preliminary operating data of the model reasoning service is collected; based on the query language PromQL of Prometheus, the preliminary operating data is compared with a preset data retention database to determine the operating data.

[0008] In one implementation of the present application, the operation data is cleaned and organized to generate standard operation data, specifically including: removing or filling in outliers and missing values ​​in the operation data based on a preset linear interpolation method to generate completed data; standardizing the completed data to eliminate the dimensional differences between different indicators in the completed data to generate completed standard data; sorting the completed standard data in the order of time series to form standard operation data containing timestamps, operation data names and indicator values.

[0009] In one implementation of the present application, a transformer-based time series prediction model is constructed, specifically including: constructing a transformer model to be trained; wherein the transformer model to be trained includes an input embedding layer, a position encoding layer, multiple transformer encoding layers and an output layer; representing the input embedding layer based on an embedded vector to convert standard operating data into a high-dimensional vector; representing the position encoding layer based on a sine function encoding to add position information to the vector; multiple transformer encoding layers are set to 6 layers, each layer contains 8 attention heads, the hidden layer dimension is 512, and the dropout rate is 0.1; the output layer adopts a fully connected layer; and training the transformer model to be trained based on a preset data set and standard operating data to generate a time series prediction model.

[0010] In one implementation of the present application, a transformer model to be trained is trained based on a preset data set and standard operating data to generate a time series prediction model, specifically including: dividing the data set into a training set and a validation set based on a cross-validation algorithm; initializing the parameters of the transformer model to be trained based on a preset Xavier initialization algorithm; iteratively training the transformer model to be trained based on the training set, and validating the transformer model to be trained based on the validation set, until the transformer model to be trained converges to generate a time series prediction model.

[0011] In one implementation of the present application, standard operating data is processed based on a timing prediction model to generate a load forecast for a preset time period, specifically including: inputting the standard operating data into the timing prediction model for prediction; and predicting the load forecast for the preset time period through forward propagation of the timing prediction model.

[0012] In one implementation of the present application, the load forecast for a preset time period is delivered to the HPA of the Kubernetes platform, and the load quantity of the model reasoning service is adjusted based on the HPA, specifically including: delivering the load forecast result for the preset time period to the HPA of the Kubernetes platform through the RESTful API interface of the Kubernetes platform; when the load forecast for the preset time period exceeds a first threshold, increasing the number of Pods of the model reasoning service; when the load forecast for the preset time period is lower than a second threshold, reducing the number of Pods of the model reasoning service.

[0013] In one implementation of the present application, the method also includes: setting alarm rules based on the PrometheusAlertmanager of the Kubernetes platform; notifying relevant personnel when the performance of the model reasoning service reaches a preset performance threshold and / or abnormal fluctuations occur.

[0014] In the second aspect, an embodiment of the present application also provides a model reasoning service elastic load prediction device, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: collect operation data of the model reasoning service based on Prometheus of the Kubernetes platform; wherein the operation data includes CPU usage, memory usage, request processing time, request success rate and network traffic; clean and organize the operation data to generate standard operation data; build a transformer-based timing prediction model; process the standard operation data based on the timing prediction model to generate a load prediction for a preset time period; transmit the load prediction for the preset time period to the HPA of the Kubernetes platform, and adjust the load quantity of the model reasoning service based on the HPA.

[0015] On the third aspect, an embodiment of the present application also provides a non-volatile computer storage medium for elastic load prediction of a model reasoning service, storing computer executable instructions, and the computer executable instructions are set to: collect operation data of the model reasoning service based on Prometheus of the Kubernetes platform; wherein the operation data includes CPU utilization, memory utilization, request processing time, request success rate, and network traffic; clean and organize the operation data to generate standard operation data; build a transformer-based timing prediction model; process the standard operation data based on the timing prediction model to generate a load prediction for a preset time period; transmit the load prediction for the preset time period to the HPA of the Kubernetes platform, and adjust the load quantity of the model reasoning service based on the HPA.

[0016] The embodiments of the present application provide a method, device and medium for elastic load prediction of a model reasoning service, which collects the operating data of the model reasoning service through the Prometheus of the Kubernetes platform, cleans and organizes it, generates standard operating data, and improves the quality and availability of the data. A transformer-based time series prediction model is constructed, and the model is used to process standard operating data to generate accurate load predictions for preset time periods, thereby improving the accuracy and reliability of load predictions. The load prediction results are transmitted to the HPA of the Kubernetes platform, and the load quantity of the model reasoning service is dynamically adjusted according to the prediction results, thereby realizing elastic load management and optimizing resource utilization. The present invention can more accurately predict load changes, reduce manual intervention and operation and maintenance costs, and realize more intelligent resource scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A flow chart of a model reasoning service elastic load prediction method provided in an embodiment of the present application;

[0019] Figure 2 A schematic diagram of the internal structure of a model reasoning service elastic load prediction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0021] The embodiments of the present application provide a model reasoning service elastic load prediction method, device and medium to solve the following technical problems: how to more accurately predict load changes, reduce manual intervention and operation and maintenance costs, and achieve more intelligent resource scheduling.

[0022] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.

[0023] Figure 1 A model inference service elastic load prediction flow chart provided in an embodiment of the present application. Figure 1 As shown, a model reasoning service elastic load prediction method provided in an embodiment of the present application specifically includes the following steps:

[0024] Step 1: Prometheus on the Kubernetes platform collects the operating data of the model inference service. The operating data includes CPU usage, memory usage, request processing time, request success rate, and network traffic.

[0025] Step 11. Configure Prometheus. Prometheus is used to monitor the Pods and nodes of the model inference service.

[0026] In order for Prometheus to accurately monitor the Pods and nodes of the model inference service, we need to configure it accordingly.

[0027] In the Prometheus configuration file, define the Pod and node of the model inference service as the monitoring target. Set the interval for Prometheus to capture indicator data to ensure the real-time and accuracy of the data.

[0028] Step 12: Collect preliminary operation data of the model inference service based on ServiceMonitor in Prometheus.

[0029] ServiceMonitor is a custom resource of Prometheus in the Kubernetes environment, which is used to define how to monitor specific services.

[0030] Prometheus automatically discovers the ServiceMonitor resource in the Kubernetes cluster and starts to capture preliminary operating data of the model inference service based on its definition.

[0031] Step 13: Based on Prometheus' query language PromQL, compare the preliminary operation data with the preset data retention database to determine the operation data.

[0032] PromQL is a powerful query language provided by Prometheus that allows users to perform complex queries and analyses on collected indicator data.

[0033] According to the operation data requirements of the model inference service, write the corresponding PromQL query statement to extract preliminary operation data from Prometheus. Compare the preliminary operation data with the data retention database to determine whether the preliminary operation data is the data that needs to be monitored.

[0034] In the embodiments of the present application:

[0035] Deploy a Kubernetes cluster and install the Prometheus monitoring system in the cluster. Deploy the model inference service in the Kubernetes cluster and configure the corresponding Service and Pod.

[0036] In the Prometheus configuration file, add the Pod and node of the model inference service as monitoring targets.

[0037] Set the crawling interval to 15 seconds to ensure the real-time nature of the data. Configure the data storage path and retention policy of Prometheus to meet the needs of long-term monitoring.

[0038] Create a ServiceMonitor resource for the model inference service, specify the monitoring port as 8080, the path as / metrics, and the crawling interval consistent with the Prometheus configuration.

[0039] Write PromQL query statements to extract operating data such as CPU usage, memory usage, request processing time, request success rate, and network traffic of the model inference service.

[0040] Step 2: Clean and organize the operation data to generate standard operation data.

[0041] Step 21: Remove or fill in outliers and missing values ​​in the running data based on a preset linear interpolation method to generate completed data.

[0042] Outliers are data that differ significantly from most data points, which may be caused by system errors, sensor failures, or other abnormal conditions. To remove outliers, statistical methods or machine learning algorithms can be used to identify and remove them from the data set.

[0043] Missing values ​​refer to the missing data at certain time points in a data set. In order to fill the missing values, the present invention adopts a preset linear interpolation method. Linear interpolation is a simple and effective interpolation method that estimates missing values ​​through a linear relationship based on the values ​​of adjacent data points. Specifically, for each missing value, the two adjacent data points before and after it are found, and then the values ​​of these two data points and the time interval between them are used to calculate the estimated value of the missing value through a linear interpolation formula.

[0044] After outlier processing and missing value filling, the completed data are generated.

[0045] Step 22: Standardize the completed data to eliminate the dimensional differences between different indicators in the completed data to generate completed standard data.

[0046] Since the operation data contains multiple indicators, such as CPU usage, memory usage, request processing time, etc., these indicators have different dimensions, and direct comparison or calculation will lead to inaccurate results. Therefore, it is necessary to standardize the supplementary data to eliminate the dimension differences between different indicators.

[0047] The present invention adopts the Z-score standardization method, that is, the value of each indicator is subtracted from its mean value, and then divided by its standard deviation. After such processing, the data has a mean value of 0 and a standard deviation of 1, and different indicators are comparable.

[0048] After standardization, complete standard data is generated, in which each indicator has the same dimension and distribution.

[0049] Step 23: Sort the completed standard data in the order of time series to form standard operation data including timestamp, operation data name and indicator value.

[0050] Since the operation data is collected in chronological order, the supplementary standard data needs to be sorted in the order of the time series to ensure the time sequence of the data.

[0051] The sorted data contains three parts: timestamp, running data name and indicator value. The timestamp indicates the data collection time, the running data name indicates the indicator name of the data, such as CPU usage, memory usage, etc., and the indicator value indicates the value of the corresponding indicator at that time point.

[0052] After time series sorting and formatting, standard operating data are generated.

[0053] Step 3: Build a transformer-based time series prediction model.

[0054] The Ransformer model has sequence modeling and parallel processing capabilities, and performs well in the field of time series prediction.

[0055] Step 31, construct a transformer model to be trained; wherein the transformer model to be trained includes an input embedding layer, a position encoding layer, multiple transformer encoding layers and an output layer; the input embedding layer is represented based on the embedded vector to convert the standard running data into a high-dimensional vector; the position encoding layer is represented based on the sine function encoding to add position information to the vector; the multiple transformer encoding layers are set to 6 layers, each layer contains 8 attention heads, the hidden layer dimension is 512, and the dropout rate is 0.1; the output layer adopts a fully connected layer.

[0056] The input embedding layer is responsible for converting the standard operating data into a high-dimensional vector. An embedding vector is a fixed-length vector representing the input data, which can map discrete data points into a continuous high-dimensional space, thereby facilitating subsequent neural network processing. In the present invention, the input embedding layer converts standard operating data (such as CPU usage, memory usage, etc.) into corresponding high-dimensional vectors.

[0057] Since the transformer model itself does not have the ability to process position information in sequence data, it is necessary to add position information to the vector through the position encoding layer. The present invention adopts a sine function encoding method, which can generate a unique code for each position, thereby helping the model understand the sequential relationship between data points.

[0058] The present invention is set to 6 layers of transformer encoding layers, each layer contains 8 attention heads, the hidden layer dimension is 512, and the dropout rate is 0.1. The transformer encoding layer is the core part of the transformer model, which captures the long-distance dependencies in the sequence data through the self-attention mechanism. The stacking of multiple encoding layers can further enhance the expressiveness of the model.

[0059] A fully connected layer is used as the output layer to map the output of the transformer encoding layer to the prediction target (such as the load value). The fully connected layer is a commonly used neural network layer that can multiply the input vector with the weight matrix and add the bias term to obtain the output.

[0060] Step 32: Train the transformer model to be trained based on the preset data set and standard operating data to generate a time series prediction model.

[0061] Step 321: divide the data set into a training set and a validation set based on a cross-validation algorithm.

[0062] Cross-validation is a commonly used model evaluation method that evaluates the performance of a model by dividing a data set into multiple subsets and taking one of the subsets as a validation set and the remaining subsets as training sets in turn. The present invention uses a cross-validation algorithm to divide a preset data set into a training set and a validation set to ensure the generalization ability of the model.

[0063] Step 322: Initialize the parameters of the transformer model to be trained based on the preset Xavier initialization algorithm.

[0064] Parameter initialization is an important step in the neural network training process, which affects the convergence speed and final performance of the model. The Xavier initialization algorithm is a commonly used parameter initialization method, which automatically adjusts the initial value of the parameter according to the input and output dimensions, thereby making the neural network more stable during the training process. The present invention uses the Xavier initialization algorithm to initialize the parameters of the transformer model to be trained.

[0065] Step 323: iteratively train the transformer model to be trained based on the training set, and verify the transformer model to be trained based on the verification set until the transformer model to be trained converges, thereby generating a time series prediction model.

[0066] During the training process, the present invention adopts an iterative training method, that is, repeatedly using the training set to update the parameters of the model until the performance of the model on the validation set is no longer improved. Specifically, each iteration calculates the loss function value of the model on the training set, and updates the parameters of the model through the back propagation algorithm. At the same time, the performance of the model is regularly evaluated on the validation set to monitor the generalization ability of the model. When the performance of the model on the validation set reaches a preset threshold, the model is considered to have converged, and a time series prediction model can be generated at this time.

[0067] Step 4: Process the standard operating data based on the time series forecasting model to generate a load forecast for a preset time period.

[0068] Step 41: Input the standard operation data into the time series prediction model for prediction.

[0069] Step 42: Predict the load forecast for a preset time period through forward propagation of the time series prediction model.

[0070] In this step, the standard operation data passes through the forward propagation of the time series prediction model, and is processed by the input embedding layer, position encoding layer, multiple transformer encoding layers, and output layer to finally generate the load prediction result.

[0071] Before forward propagation, the prediction period needs to be specified. This can be set according to actual needs, such as predicting the load in the next hour, day, or week.

[0072] After the forward propagation is completed, the model will output the load forecast results within the forecast period. These results are usually presented in the form of a sequence, with each element representing the load forecast value at a time point.

[0073] Step 5: Deliver the load forecast for the preset time period to the HPA of the Kubernetes platform, and adjust the load quantity of the model inference service based on the HPA.

[0074] Step 51: The load forecast result for the preset time period is transmitted to the HPA of the Kubernetes platform through the RESTful API interface of the Kubernetes platform.

[0075] First, you need to ensure that the load forecast results are organized in a suitable format, usually a dataset containing time series and corresponding load forecast values. The Kubernetes platform provides RESTful API interfaces that allow external systems to interact with it. In this step, these API interfaces will be used to deliver the load forecast results to the HPA. The load forecast results are sent programmatically (such as using HTTP requests) to a specific endpoint of the Kubernetes API server, which is associated with the HPA for receiving and adjusting scaling policies.

[0076] Step 52: When a load forecast in a preset time period exceeds a first threshold, increase the number of Pods of the model reasoning service.

[0077] According to the performance and resource requirements of the model inference service, a load threshold (first threshold) is set. When the predicted load exceeds this threshold, it means that the number of Pods needs to be increased to cope with the upcoming load peak. When HPA receives a load prediction result that exceeds the first threshold, it triggers the scaling strategy to increase the number of Pods of the model inference service. The number of increases can be determined based on the specific scaling strategy and resource configuration.

[0078] Step 53: When a load forecast in a preset time period is lower than a second threshold, reduce the number of Pods of the model reasoning service.

[0079] Similarly, according to the performance and resource utilization efficiency of the model reasoning service, another load threshold (second threshold) is set. When the predicted load is lower than this threshold, it means that the number of Pods can be reduced to save resources. When HPA receives a load prediction result lower than the second threshold, it triggers the scaling policy to reduce the number of Pods of the model reasoning service. The number of reductions is also determined by the scaling policy and resource configuration.

[0080] In an embodiment of the present application: a load forecast result for the next 24 hours is obtained, and the result is presented in the form of a time series and a corresponding load forecast value.

[0081] Understandably, the load prediction results are processed and converted into JSON format to be compatible with the Kubernetes API.

[0082] Use programming methods (such as Python's requests library) to send HTTP requests and deliver the load prediction results to HPA through the RESTful API interface of the Kubernetes platform. Set the first threshold and the second threshold based on the performance and resource requirements of the model inference service. For example, increase the number of Pods when the load exceeds 80%, and reduce the number of Pods when the load is less than 30%.

[0083] In addition, this application also includes the following methods:

[0084] A1. Set alert rules based on Prometheus Alertmanager on the Kubernetes platform.

[0085] Alertmanager is a component in the Prometheus ecosystem that is responsible for processing alerts sent by Prometheus and performing deduplication, grouping, suppression, and routing to appropriate recipients (such as email, Slack, etc.) according to configured rules.

[0086] A2. When the performance of the model inference service reaches the preset performance threshold and / or experiences abnormal fluctuations, relevant personnel are notified.

[0087] In the Prometheus configuration file, use PromQL to define alert rules. For example, set a rule to trigger an alert when the response time of the model inference service exceeds 500 milliseconds for 5 minutes, or when the success rate is lower than 95%.

[0088] In addition, set a rule to detect abnormal fluctuations, such as when the change rate of response time exceeds 20%, an alarm is also triggered.

[0089] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a model reasoning service elastic load prediction device, whose structure is as follows: Figure 2 shown.

[0090] Figure 2 A schematic diagram of the internal structure of a model reasoning service elastic load prediction device provided in an embodiment of the present application. Figure 2 As shown, the device includes:

[0091] at least one processor 201;

[0092] and, a memory 202 communicatively connected to the at least one processor;

[0093] The memory 202 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 201 to enable at least one processor 201 to:

[0094] Prometheus based on the Kubernetes platform collects the operating data of the model inference service; the operating data includes CPU usage, memory usage, request processing time, request success rate and network traffic; cleans and organizes the operating data to generate standard operating data; builds a transformer-based time series prediction model; processes the standard operating data based on the time series prediction model to generate a load forecast for a preset time period; transmits the load forecast for the preset time period to the HPA of the Kubernetes platform, and adjusts the load quantity of the model inference service based on the HPA.

[0095] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for elastic load prediction of a model inference service, storing computer executable instructions, wherein the computer executable instructions are set to:

[0096] Prometheus based on the Kubernetes platform collects the operating data of the model inference service; the operating data includes CPU usage, memory usage, request processing time, request success rate and network traffic; cleans and organizes the operating data to generate standard operating data; builds a transformer-based time series prediction model; processes the standard operating data based on the time series prediction model to generate a load forecast for a preset time period; transmits the load forecast for the preset time period to the HPA of the Kubernetes platform, and adjusts the load quantity of the model inference service based on the HPA.

[0097] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0098] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.

[0099] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0100] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0101] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0103] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0104] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0105] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0106] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0107] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A model inference service elastic load prediction method, applied to the model inference service of the Kubernetes platform, characterized in that: The method comprises: Prometheus on the Kubernetes platform collects the operating data of the model inference service; wherein the operating data includes CPU usage, memory usage, request processing time, request success rate, and network traffic; Cleaning and arranging the operation data to generate standard operation data; Build a transformer-based time series prediction model; Processing the standard operation data based on the time series prediction model to generate a load prediction for a preset time period; The load forecast for a preset time period is delivered to the HPA of the Kubernetes platform, and the load quantity of the model inference service is adjusted based on the HPA.

2. A model reasoning service elastic load prediction method according to claim 1, characterized in that: Prometheus based on the Kubernetes platform collects the operation data of the model inference service, specifically including: Configure Prometheus; wherein, Prometheus is used to monitor the Pods and nodes of the model inference service; Collect preliminary operation data of the model reasoning service based on the ServiceMonitor in the Prometheus; Based on Prometheus' query language PromQL, the preliminary operation data is compared with a preset data retention database to determine the operation data.

3. According to claim 1, a model reasoning service elastic load prediction method is characterized in that: Cleaning and arranging the operation data to generate standard operation data, specifically including: Removing or filling in outliers and missing values ​​in the operating data based on a preset linear interpolation method to generate supplementary data; Standardizing the supplemented data to eliminate dimensional differences between different indicators in the supplemented data to generate supplemented standard data; The completed standard data is sorted in the order of time series to form standard operation data including timestamp, operation data name and indicator value.

4. According to claim 1, a model reasoning service elastic load prediction method is characterized in that: Build a transformer-based time series prediction model, including: Constructing a transformer model to be trained; wherein the transformer model to be trained includes an input embedding layer, a position encoding layer, multiple transformer encoding layers and an output layer; Representing the input embedding layer based on the embedding vector to convert the standard operation data into a high-dimensional vector; The position encoding layer is represented based on a sine function encoding to add position information to the vector; The multiple transformer encoding layers are set to 6 layers, each layer contains 8 attention heads, the hidden layer dimension is 512, and the dropout rate is 0.1; The output layer adopts a fully connected layer; The transformer model to be trained is trained based on a preset data set and the standard operating data to generate a time series prediction model.

5. A model reasoning service elastic load prediction method according to claim 4, characterized in that: Training the transformer model to be trained based on the preset data set and the standard operation data to generate a time series prediction model specifically includes: Dividing the data set into a training set and a validation set based on a cross-validation algorithm; Initializing the parameters of the transformer model to be trained based on a preset Xavier initialization algorithm; The transformer model to be trained is iteratively trained based on the training set, and the transformer model to be trained is verified based on the verification set until the transformer model to be trained converges, thereby generating the time series prediction model.

6. A model reasoning service elastic load prediction method according to claim 1, characterized in that: Processing the standard operation data based on the time series prediction model to generate a load prediction for a preset time period specifically includes: Inputting the standard operation data into the time series prediction model to perform prediction; The load forecast for a preset time period is predicted through the forward propagation of the time series prediction model.

7. A model reasoning service elastic load prediction method according to claim 1, characterized in that: The load forecast for a preset time period is transmitted to the HPA of the Kubernetes platform, and the load quantity of the model inference service is adjusted based on the HPA, specifically including: The load forecast results for the preset time period are transmitted to the HPA of the Kubernetes platform through the RESTful API interface of the Kubernetes platform; When a load forecast in a preset time period exceeds a first threshold, increasing the number of Pods of the model reasoning service; When there is a load forecast in the preset time period that is lower than the second threshold, the number of Pods of the model reasoning service is reduced.

8. A model reasoning service elastic load prediction method according to claim 1, characterized in that: The method further comprises: Set alert rules based on Prometheus Alertmanager on the Kubernetes platform; When the performance of the model reasoning service reaches a preset performance threshold and / or experiences abnormal fluctuations, relevant personnel are notified.

9. A model reasoning service elastic load prediction device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Prometheus on the Kubernetes platform collects the operating data of the model inference service; wherein the operating data includes CPU usage, memory usage, request processing time, request success rate, and network traffic; Cleaning and arranging the operation data to generate standard operation data; Build a transformer-based time series prediction model; Processing the standard operation data based on the time series prediction model to generate a load prediction for a preset time period; The load forecast for a set time period is delivered to the HPA of the Kubernetes platform, and the load quantity of the model inference service is adjusted based on the HPA.

10. A non-volatile computer storage medium for elastic load prediction of a model reasoning service, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Prometheus on the Kubernetes platform collects the operating data of the model inference service; wherein the operating data includes CPU usage, memory usage, request processing time, request success rate, and network traffic; Cleaning and arranging the operation data to generate standard operation data; Build a transformer-based time series prediction model; Processing the standard operation data based on the time series prediction model to generate a load prediction for a preset time period; The load forecast for a preset time period is delivered to the HPA of the Kubernetes platform, and the load quantity of the model inference service is adjusted based on the HPA.

Citation Information

Cited By

  • Performance index prediction method and device for Pod and medium

    CN121958053A

  • A performance index prediction method, device and medium for a pod

    CN121958053B