Server load prediction method, electronic device, storage medium, and program product

By acquiring multi-layered data from the server cluster and using temporal convolutional networks and graph attention networks to predict load, the problem of resource exhaustion and latency caused by sudden load increases in the server cluster is solved, thereby improving the accuracy of load prediction and the quality of service.

CN120723591BActive Publication Date: 2025-11-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511244893.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-21
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

The server cluster experiences a sudden surge in load due to factors such as fluctuations in business traffic, leading to resource exhaustion and service response delays, which affects service quality.

Method used

By acquiring data from the physical, virtual, and application layers of the server cluster, the relationships and data characteristics between the servers are determined. Temporal convolutional networks and graph attention networks are used to predict the load in future periods, and resource allocation is adjusted to avoid sudden increases in load.

Benefits of technology

Accurately predict server cluster load, reduce service response latency, and improve service quality and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723591B_ABST
    Figure CN120723591B_ABST
Patent Text Reader

Abstract

The application discloses a kind of server load prediction method, electronic equipment, storage medium and program product, it is related to computer technical field, including obtaining the physical layer data, virtual layer data and application layer data of server cluster;According to the first relationship between each server in server cluster determined by physical layer data, virtual layer data and application layer data;According to the feature of the data corresponding to each node in first relationship, the second feature corresponding to server cluster is determined, and the second feature includes the space relationship and time relationship between the data corresponding to each node.In the above method, the space relationship and time relationship between the data corresponding to each node in server cluster can accurately predict the load of server cluster in future period, so that server cluster can adjust resource allocation according to predicted load value, so as to reduce service response delay, improve the service quality of server cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to server load prediction methods, electronic devices, storage media, and program products. Background Technology

[0002] A server cluster consists of multiple independent servers connected by a network. Under unified management and scheduling, these servers can integrate and share computing power and storage resources. Server clusters are widely used in cloud computing platforms, large-scale distributed systems, and other applications.

[0003] However, during the operation of a server cluster, the server load often increases suddenly due to factors such as fluctuations in business traffic. Increased load can easily lead to server resource exhaustion and server processing capacity saturation, which in turn can cause problems such as service response delays, seriously affecting the service quality of the server cluster. Summary of the Invention

[0004] This application provides a server load prediction method, electronic device, storage medium, and program product to at least solve the problem of slow service response in server clusters in related technologies.

[0005] This application provides a server load prediction method, including:

[0006] Obtain physical layer data, virtual layer data, and application layer data of the server cluster;

[0007] Based on physical layer data, virtual layer data and application layer data, the first relationship between each server in the server cluster is determined. The first relationship includes the nodes corresponding to the first features of each type of data, as well as the edges between each node. The edges are used to indicate the association between the data corresponding to the nodes.

[0008] Based on the characteristics of the data corresponding to each node in the first relationship, the second characteristics of the server cluster are determined. The second characteristics include the spatial and temporal relationships between the data corresponding to each node.

[0009] Based on the second feature, predict the load of the server cluster in future time periods.

[0010] This application also provides a server load prediction device, comprising: an acquisition module, a first determination module, a second determination module, and a prediction module, wherein:

[0011] The acquisition module is used to acquire physical layer data, virtual layer data, and application layer data of the server cluster.

[0012] The first determining module is used to determine the first relationship between servers in the server cluster based on physical layer data, virtual layer data and application layer data. The first relationship includes the nodes corresponding to the first features of various types of data, as well as the edges between the nodes. The edges are used to indicate the association between the data corresponding to the nodes.

[0013] The second determining module is used to determine the second feature corresponding to the server cluster based on the characteristics of the data corresponding to each node in the first relationship. The second feature includes the spatial relationship and temporal relationship between the data corresponding to each node.

[0014] The prediction module is used to predict the load of the server cluster in future time periods based on the second feature.

[0015] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described server load prediction methods.

[0016] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described server load prediction methods.

[0017] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described server load prediction methods.

[0018] This application enables the accurate prediction of the server cluster's load in future periods by utilizing the spatial and temporal relationships between the data corresponding to each node in the server cluster. This allows the server cluster to adjust resource allocation based on the predicted load value, thereby solving the technical problem of slow service response in server clusters and achieving the technical effect of reducing service response latency and improving the service quality of server clusters. Attached Figure Description

[0019] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating the server load prediction method provided in this application embodiment;

[0021] Figure 2 A flowchart illustrating a method for determining a second feature provided in an embodiment of this application;

[0022] Figure 3 A flowchart illustrating a method for determining a first relationship provided in an embodiment of this application;

[0023] Figure 4 A flowchart illustrating a method for acquiring various types of data provided in an embodiment of this application;

[0024] Figure 5 A method for adjusting resource allocation is provided for embodiments of this application;

[0025] Figure 6 This is a schematic diagram of the server load prediction device provided in the embodiments of this application;

[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0028] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0029] Definitions:

[0030] Time step: A time step is a fixed time interval unit set to unify the time granularity of various types of data. For example, a time step is 1 second (s), indicating that various types of data are collected and unified at 1-second intervals.

[0031] Feature dimension: Feature dimension refers to the number of core feature indicators aggregated by a single node at a single time step. For example, a feature dimension of 3 indicates that a single node can describe its own state at a single time step through 3 values ​​(e.g., fan speed, power consumption, memory usage).

[0032] A server cluster consists of multiple independent servers connected via a network. Under unified management and scheduling, these servers can integrate and share computing power and storage resources.

[0033] However, during the operation of a server cluster, the load status of the server cluster often increases suddenly due to factors such as fluctuations in business traffic (such as peak access, time-limited events), service dependency linkage (such as cascading requests in microservice call chains), or changes in hardware status (such as node performance degradation). Increased load can easily lead to exhaustion of server resources and saturation of server processing capacity, which in turn can cause problems such as service response delays, seriously affecting the service quality of the server cluster.

[0034] To address the aforementioned issues, this embodiment of the application acquires physical layer data, virtual layer data, and application layer data of the server cluster. Based on the physical layer data, virtual layer data, and application layer data, a first relationship between the servers in the server cluster is determined. The first relationship includes nodes corresponding to first features of various types of data, and edges between nodes, where edges indicate the association between the data corresponding to nodes. Based on the features of the data corresponding to each node in the first relationship, a second feature corresponding to the server cluster is determined. The second feature includes spatial and temporal relationships between the data corresponding to each node. Based on the second feature, the load of the server cluster in future time periods is predicted. In the above method, by using the spatial and temporal relationships between the data corresponding to each node in the server cluster, the load of the server cluster in future time periods can be accurately predicted. This allows the server cluster to adjust resource allocation based on the predicted load value to avoid server resource exhaustion due to sudden increases in load, thereby reducing service response latency and improving the service quality of the server cluster.

[0035] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0036] Figure 1 This is a flowchart illustrating the server load prediction method provided in the embodiments of this application, as shown below. Figure 1 As shown, embodiments of this application provide a server load prediction method, which is described in detail below:

[0037] S101. Obtain physical layer data, virtual layer data, and application layer data of the server cluster.

[0038] The execution subject of this application embodiment can be an electronic device or a server load prediction device installed in an electronic device. The server load prediction device can be implemented by software or by a combination of software and hardware.

[0039] In one possible implementation, the physical layer data of the server cluster includes, but is not limited to: the temperature, power consumption, fan speed of the central processing unit (CPU) of each node in the server, and the identifier corresponding to the physical node;

[0040] The virtual layer data of the server includes, but is not limited to: CPU utilization, memory usage, input / output (I / O) throughput of each node in the server, and the identifiers corresponding to the virtual nodes;

[0041] The application layer data of the server includes, but is not limited to: the request time of the microservice call chain of each node in the server, the dependency relationship between each node, and the identifier corresponding to the application service node;

[0042] Physical nodes refer to the physical hardware devices of a server cluster (such as physical servers), virtual nodes refer to logical nodes abstracted from physical nodes through virtualization technologies (such as virtual hypervisors), such as virtual machines and containers, and application service nodes refer to the logical nodes in a server cluster that host application, service, or microservice instances.

[0043] S102. Based on physical layer data, virtual layer data and application layer data, determine the first relationship between each server in the server cluster. The first relationship includes the nodes corresponding to the first features of each type of data, as well as the edges between each node. The edges are used to indicate the association between the data corresponding to the nodes.

[0044] In one possible implementation, the node corresponding to the first feature of the physical layer data is a physical node, the node corresponding to the first feature of the virtual layer data is a virtual node, and the node corresponding to the first feature of the application layer data is an application service node.

[0045] In one possible implementation, the first feature corresponds to the core attributes of the three types of data in S101, and integrates the spatial correlation features between each node.

[0046] The core attributes of the three types of data include: the encoding information of physical layer data, the encoding information of virtual layer data, and the encoding information of application layer data. The encoding information is the feature vector after the data is encoded.

[0047] In one possible implementation, the format of the first relation is a mathematical, structured data format that a computer can store, compute, and process. For example, the format of the first relation is a multidimensional vector, matrix, dictionary, or table.

[0048] For example, taking a dictionary format based on the first relation, each node, its corresponding first feature, and edges can be stored using key-value pairs. In one implementation, the nodes corresponding to the first features of various data types can be stored as keys, and the first features and edges corresponding to each node can be stored as values.

[0049] S103. Based on the characteristics of the data corresponding to each node in the first relationship, determine the second feature corresponding to the server cluster. The second feature includes the spatial and temporal relationships between the data corresponding to each node.

[0050] In one possible implementation, the spatial relationship between the data corresponding to each node refers to the topological association between each node and the data corresponding to each node, such as the dependency relationship between each node and the association relationship between each node.

[0051] In one possible implementation, the temporal relationship between the data corresponding to each node refers to the dynamic evolution of the data corresponding to each node in the time dimension. The dynamic evolution refers to the fluctuation characteristics of the node data over time, that is, the fluctuation characteristics of the node data at different time steps, such as the daily peak and trough of physical node power consumption, the periodic increase of virtual node memory usage, etc.

[0052] S104. Based on the second feature, predict the load of the server cluster in future time periods.

[0053] In one possible implementation, the load on the server cluster in future time periods can be predicted based on a second feature of the server cluster at the current time step and historical time steps.

[0054] In one possible implementation, the second feature can be predicted based on a dynamic weight fusion algorithm to obtain the load.

[0055] In this embodiment, physical layer data, virtual layer data, and application layer data of the server cluster are obtained. Based on the physical layer data, virtual layer data, and application layer data, a first relationship between the servers in the server cluster is determined. The first relationship includes nodes corresponding to first features of various types of data, and edges between nodes, where edges indicate the association between the data corresponding to nodes. Based on the features of the data corresponding to each node in the first relationship, a second feature corresponding to the server cluster is determined. The second feature includes spatial and temporal relationships between the data corresponding to each node. Based on the second feature, the load of the server cluster in future periods is predicted. In the above method, by using the spatial and temporal relationships between the data corresponding to each node in the server cluster, the load of the server cluster in future periods can be accurately predicted. This allows the server cluster to adjust resource allocation based on the predicted load value to avoid server resource exhaustion due to sudden increases in load, thereby reducing service response latency and improving the service quality of the server cluster.

[0056] Based on any of the above embodiments, the following, in conjunction with Figure 2 The method for determining the second feature ( Figure 1 The embodiment of S103 will be described in detail.

[0057] Figure 2 A flowchart illustrating a method for determining a second feature provided in an embodiment of this application is shown below. Figure 2 The method may include:

[0058] S201. Based on the first feature corresponding to each node, determine the third feature. The third feature includes the spatial and temporal fusion information between the data corresponding to each node in the first relationship.

[0059] In some embodiments, time features are extracted from the first feature corresponding to each node to obtain the fourth feature corresponding to each node; for any node, the first feature corresponding to the node and the fourth feature corresponding to the node are fused to obtain the fused feature corresponding to the node; and the third feature is determined based on the fused feature corresponding to each node.

[0060] In one possible implementation, the fourth feature corresponding to each node includes the dynamic evolution of the first feature corresponding to each node in the time dimension, that is, the pattern and trend of the first feature of each node changing over time.

[0061] In one possible implementation, a temporal convolutional network (TCN) can be used to extract the temporal features of the first feature corresponding to each node, thereby obtaining the fourth feature corresponding to each node.

[0062] In one possible implementation, multiple dilated causal convolutional layers can be stacked in the temporal convolutional network to cover a sufficiently long time step. Alternatively, residual connections can be set in the temporal convolutional network (e.g., the output of the first convolutional layer is added directly to the output of the second convolutional layer through the residual path). This allows the gradient to propagate directly back to the shallow layers through the residual path, avoiding gradient vanishing and enabling the deep layers of the temporal convolutional network to effectively learn the temporal features of each node.

[0063] In one possible implementation, the fused features can be expressed as follows:

[0064]

[0065] in, Indicates the gate value, This represents element-wise multiplication. Indicates the first feature, The fourth feature is represented by a threshold value that can be preset and can take values ​​in the range [0,1]. For example, the threshold value is 0.5.

[0066] In one possible implementation, the third feature can be expressed as follows:

[0067]

[0068] in, Indicates will and Perform vertical splicing. This represents the fusion characteristics corresponding to each physical node. This represents the fusion characteristics corresponding to each virtual node. N represents the total number of physical and virtual nodes, T represents the number of time steps, and d represents the feature dimension.

[0069] In this embodiment of the application, since the first feature includes the spatial association features between nodes and the fourth feature includes the dynamic evolution law of the first feature corresponding to each node in the time dimension, the fusion feature (and the third feature) includes the spatial and temporal fusion information between the data corresponding to each node.

[0070] S202. Determine the second feature based on the third feature.

[0071] In some embodiments, multiple first weights can be determined among the nodes, the first weights being used to indicate the importance of the nodes at multiple time steps to load prediction; then, a second feature is determined based on the third feature corresponding to each node and the multiple first weights.

[0072] The specific implementation method for determining the second feature based on the third feature corresponding to each node and multiple first weights is as follows:

[0073] In some embodiments, the third features corresponding to each node are weighted according to multiple first weights to obtain the weighted features corresponding to each node; a first matrix is ​​determined according to the weighted features corresponding to each node, the first matrix being used to indicate the degree of dependency between nodes; and temporal features are extracted from the first matrix to obtain the second features.

[0074] In one possible implementation, the importance or influence of each node at each time step can be determined by analyzing the third feature in the time dimension, and the influence or importance of each node at each time step can be represented by the first weight.

[0075] In one possible implementation, the first weight can be calculated in one or more of the following ways:

[0076] The importance of different time steps can be estimated by analyzing patterns such as periodicity and trends in historical data. If a certain period of time frequently experiences peak load, then the data of that period may be considered more important, and the first weight value corresponding to that period will be larger.

[0077] The importance distribution over time is learned through the attention mechanism in a deep learning framework, and the weights at different time steps are adaptively adjusted to determine the first weight.

[0078] The weights of each node at different time steps are determined by statistical models (e.g., autoregressive integrated moving average model, exponential smoothing model, etc.).

[0079] In one possible implementation, the third feature corresponding to each node is weighted according to the first weight, which can highlight the importance of data at certain time steps for load prediction.

[0080] In one possible implementation, the attention matrix between nodes can be determined based on the weighted features, the spatial weight vector of each node can be determined based on the attention matrix between nodes, and then the first matrix can be determined based on the spatial weight vector of each node.

[0081] The attention matrix between nodes can be expressed as follows:

[0082]

[0083] in, This represents the spatial attention weight from node j to node i, used to measure the degree of influence of node j on node i. and Both are linear transformation matrices, used to generate the query vector for node i. and the key vector of node j , and Let i and j represent the feature vectors of node i and node j in the weighted features, respectively, and sqrt(d) represents the square root operation performed on the feature dimension d.

[0084] For example, the spatial weight vector of node i can be expressed as follows:

[0085]

[0086] in, This represents the spatial attention weights between node i and the features of its associated nodes. Indicates to Perform max pooling to determine node j that has the most significant impact on node i. This indicates that the result of max pooling is normalized.

[0087] In one possible implementation, the spatial weight vectors of each node can be concatenated, and the concatenated matrix can be used as the first matrix.

[0088] In one possible implementation, the temporal features of the first matrix can be extracted using a temporal convolutional network to obtain the second feature.

[0089] It should be noted that in S104, the load of the server cluster in future time periods can be predicted based on the dynamic weight fusion processor and the second features of the server cluster at the current time step and historical time steps. Specifically, the dynamic weight fusion processor can fuse the first matrix and the second feature, and then predict the load based on the second feature.

[0090] In this embodiment, by combining graph attention networks and temporal convolutional networks, joint modeling of spatial topology and time series is achieved, which can effectively extract the correlation between nodes and their evolution over time. In addition, by fusing the first feature (the spatial correlation features between nodes) and the fourth feature (the dynamic evolution law of the first feature in the time dimension) through gating values, the robustness and expressive power of feature representation can be improved, so that feature representation can still be effectively performed in complex server cluster environments (e.g., burst traffic (flash sales, etc.)).

[0091] Furthermore, in this embodiment, by combining the first weight (the importance of nodes at multiple time steps to load prediction) with the first matrix (the degree of dependence between nodes) to perform weighted fusion of the third feature, the influence of key historical moments can be highlighted in the time dimension, and the role of important nodes can be strengthened in the spatial dimension, thereby improving the accuracy of load prediction. The dynamic weight fusion device further integrates multi-dimensional features, making the prediction results more real-time and forward-looking, which helps to achieve a better resource scheduling strategy.

[0092] Based on any of the above embodiments, the following, in conjunction with Figure 3 The method for determining the first relation ( Figure 1 S102 in the embodiments will be described in detail.

[0093] Figure 3 A flowchart illustrating a method for determining a first relationship provided in an embodiment of this application is shown below. Figure 3 The method may include:

[0094] S301. Based on the physical layer data, virtual layer data, and application layer data, determine the relationship diagram. The relationship diagram includes the nodes corresponding to each type of data and the edges between each node.

[0095] In some embodiments, the initial edges between nodes can be obtained based on the amount of data transmitted between various types of data; the initial edges can be processed based on the dependencies between various types of data to obtain the edges between nodes; and the relationship graph can be determined based on the edges between nodes and the nodes corresponding to the characteristics of various types of data.

[0096] In one possible implementation, the amount of data transfer between physical nodes and virtual nodes can be determined by measuring the amount of real-time data interaction (such as the amount of CPU resource scheduling instructions transmitted and the amount of memory page swapping data) when the physical server allocates resources to virtual machines or containers.

[0097] The amount of data transferred between virtual nodes and application service nodes can also be determined by measuring the amount of data transferred between the container and the microservice instance (such as the size of application programming interface (API) request packets and the amount of data transferred in response).

[0098] The amount of data transferred between application service nodes can also be determined by determining the amount of data transferred across instances in the microservice call chain (such as the amount of response data).

[0099] In one possible implementation, the dependency relationship between physical nodes and virtual nodes can be determined. For example, if a virtual node is completely dependent on the hardware resources of a physical node (such as a virtual machine being bound to a physical CPU core), the corresponding initial edge will be increased by 20%-50%.

[0100] It can also determine the dependencies between virtual nodes and application service nodes. For example, if the application service node is only deployed on a specific virtual node (such as a POD binding container), the corresponding initial edge can be boosted.

[0101] In one possible implementation, a dynamic graph can be determined based on the edges between nodes acquired in real time and the nodes corresponding to the characteristics of various types of data. Then, the dynamic graph can be structurally represented by a graph feature encoding method, transforming the dynamic graph from an abstract conceptual graph into a relational graph. The format of the relational graph is a mathematical and structured data format that can be stored, calculated, and processed by a computer.

[0102] S302. Based on the relationship diagram, determine the first relationship.

[0103] In one possible implementation, feature extraction of the relation graph can be performed based on a graph attention network (GAT) to determine the first relation. GAT can obtain the local structural characteristics of a node (such as the distribution of neighboring nodes), its global position (importance in the entire graph), and its own attribute information (such as CPU utilization, memory usage, etc.). Here, neighboring nodes are one or more nodes adjacent to the current node.

[0104] In this embodiment, by using a dynamic graph construction method to model each node, not only can the static topology be captured, but the dynamic changes in the relationships between nodes during operation can also be reflected. By analyzing the data transmission volume and dependencies between nodes, prediction bias is reduced and the semantic expressive power of edges is enhanced, which helps to more accurately predict cascading failure scenarios, thereby more accurately depicting the actual interaction patterns between servers and providing high-quality graph structure input for subsequent feature extraction and load prediction.

[0105] Based on any of the above embodiments, the following, in conjunction with Figure 4 Methods for obtaining physical layer data, virtual layer data, and application layer data of server clusters ( Figure 1 The S101 example will be described in detail.

[0106] Figure 4 A flowchart illustrating a method for acquiring various types of data provided in this application embodiment is shown below for details. Figure 4 The method may include:

[0107] S401. Collect physical layer data, virtual layer data, and application layer data.

[0108] In one possible implementation, physical layer data collected by sensors on the physical hardware of the server cluster can be read through an intelligent platform management interface (IPMI). For example, the storage format of each piece of physical layer data can be physical node identifier + physical indicator name + physical value + collection timestamp (accurate to milliseconds). The physical indicator name can be the CPU temperature mentioned in the above embodiment, and the corresponding physical value can be 60 degrees Celsius (°C).

[0109] Virtual layer data can also be collected periodically through kernel-based virtual machines (KVM). For example, the storage format of each piece of virtual layer data can be virtual node identifier + virtual metric name + virtual value + collection timestamp. The virtual metric name can be the CPU utilization mentioned in the above embodiment, and the corresponding virtual value can be 80%.

[0110] Application layer data can also be collected through distributed application performance monitoring tools (SkyWalking), distributed tracing tools (Pinpoint), etc., or through event tracking. For example, the storage format of each piece of application layer data can be application node identifier + application metric name + application value + collection timestamp. The application metric name can be the request time of the microservice call chain mentioned in the above embodiment, and the corresponding application value can be 150 milliseconds (ms).

[0111] S402. Based on the timestamps in the physical layer data, the virtual layer data, and the application layer data, perform time alignment on the physical layer data, the virtual layer data, and the application layer data.

[0112] In one possible implementation, the time step grid can be divided into units of preset time steps (e.g., 10 seconds), and the timestamp of each data item can be mapped to the corresponding time step grid, thereby mapping the data corresponding to the timestamp to the corresponding time step grid.

[0113] In one possible implementation, for time steps with no data (e.g., no virtual layer data within a certain 10s), the data can be filled with data from the previous time step or a value of 0.

[0114] In one possible implementation, for multiple similar data points within the same time step (e.g., the CPU temperature of a physical node was collected 3 times within 10 seconds), the average or maximum value of the multiple similar data points can be taken as the data of that type within that time step.

[0115] S403. Standardize the aligned data to obtain the physical layer data, virtual layer data, and application layer data of the server cluster.

[0116] In one possible implementation, the aligned data can be standardized using methods such as min-max normalization or z-score standardization.

[0117] In this embodiment, by collecting multi-level data from the physical layer, virtual layer, and application layer, and performing timestamp normalization on the data, the time deviation problem between different data sources can be effectively eliminated, improving the accuracy of subsequent modeling and prediction. In addition, standardizing various types of data can eliminate differences in the dimensions of various indicators, making different types of data comparable and fusionable in subsequent models, laying the foundation for building unified data.

[0118] Based on the above embodiments, after predicting the load of the server cluster in future time periods according to the second feature, the resource configuration of the server cluster can also be adjusted. For example, Figure 5 A method for adjusting resource configuration is provided as an embodiment of this application, detailed in [reference needed]. Figure 5 The method is as follows:

[0119] S501. Based on the first strategy generation model, determine multiple control commands corresponding to the load. These multiple control commands are used to adjust the resource configuration of the server cluster.

[0120] In one possible implementation, the multiple control commands may include: commands to adjust instances, commands to adjust resource quotas, and commands to reserve resources.

[0121] Among them, adjusting instances refers to increasing or decreasing the number of virtual machines, containers and / or application layer service instances or migrating their locations; adjusting resource quotas refers to adjusting the deployment method of various services or applications on different nodes; reserving resources refers to reserving CPU cores and memory for critical physical nodes, and / or reserving resource pools for specific virtual instances or application services, etc.

[0122] S502. Sort multiple control instructions to obtain the execution order of the multiple control instructions.

[0123] In one possible implementation, multiple control commands determined in S501 can be processed based on a control command priority arbitration method to obtain the execution order of the multiple control commands. The control command priority arbitration method may include control command priority rules and conflict detection.

[0124] In one possible implementation, a set of explicit control command priority rules can be pre-set. For example, when resources are scarce, ensuring the stable operation of core business services may be given the highest priority; while when resources are sufficient, optimizing costs may become the main consideration.

[0125] For example, the control command priority rule can be:

[0126] First priority: Commands to reserve resources; Second priority: Commands to adjust instance settings; Third priority: Commands to adjust resource quotas.

[0127] Setting the first priority to the command that reserves resources can avoid the problem that adding instances later may fail due to a lack of available resources if resources are not locked in advance; setting the second priority to the command that adjusts instance quotas can avoid the problem that if resource quotas are adjusted first, but the total number of instances is insufficient to cope with high loads; setting the third priority to the command that adjusts resource quotas can adjust quotas based on the load pressure of individual instances after the number of instances is determined, avoiding resource waste (if the total number of instances is sufficient, there is no need to expand resource quotas first).

[0128] In one possible implementation, for each control command, it's possible to check for potential conflicts. For example, one control command might increase the number of instances to handle high load, while another might reduce resource allocation to save costs. In this case, a pre-defined priority rule would determine which control command should be executed.

[0129] It should be noted that with new data input (such as the latest load forecast values, real-time acquired data, etc.), electronic devices can automatically re-evaluate and adjust the priority rules of control commands.

[0130] S503. Execute multiple control instructions according to the execution order.

[0131] In one possible implementation, an ordered sequence of control instructions can be determined first, based on the execution order of multiple control instructions, and then the multiple control instructions can be executed according to the ordered sequence.

[0132] In one possible implementation, an ordered sequence of control instructions is determined based on a defined execution order, so that the control instructions are executed safely in sequence without causing system instability or other negative effects.

[0133] In one possible implementation, executing the instruction to reserve resources will yield a buffered resource configuration; executing the instruction to adjust instances will yield an updated node list; and executing the instruction to adjust resource quotas will yield a new resource allocation scheme.

[0134] The updated node list includes the available compute nodes in the server cluster and their current status (such as load, health status, etc.); the new resource allocation scheme refers to the deployment method of each service or application on different nodes.

[0135] In this embodiment, by sorting multiple control instructions and executing the sorted control instructions in an orderly manner, conflicts and fluctuations during resource adjustment are avoided. By determining a reasonable execution order through control instruction priority rules and combining instance adjustment, quota change and resource reservation strategies, fine-grained control of resource allocation is achieved.

[0136] based on Figure 5 In this embodiment, after executing multiple control instructions, the following can also be done: based on the sorted multiple control instructions, determine configuration parameters, including the number of instances, resource quota threshold, and resource reservation ratio; and process the resource configuration of the server cluster according to the configuration parameters.

[0137] In one possible implementation, configuration parameters refer to specific settings or strategies used to optimize the energy efficiency of the server cluster. Configuration parameters may include operations that control the hardware level, such as adjusting memory frequency and fan speed.

[0138] Energy-saving configuration parameters can determine specific energy-saving measures based on the status of each node. For example, for nodes with lower loads in the updated node list, a more aggressive energy-saving strategy can be adopted, such as setting the nodes with lower loads to deep sleep mode or performing CPU frequency reduction operations on the nodes with lower loads. For nodes with higher loads, more conservative energy-saving measures can be adopted to achieve better performance.

[0139] Energy-saving configuration parameters can determine specific energy-saving measures based on resource allocation schemes. For example, if some nodes are designated as backup nodes to handle burst traffic, these nodes can be set to low-power standby state and woken up only when they are actually needed.

[0140] Energy-saving configuration parameters can determine specific energy-saving measures based on buffer resource configuration. For example, reserved resources that have not been used for a long time in the buffer resource configuration (such as those without sudden load occupation within 2 hours) can be set to deep sleep mode.

[0141] In this embodiment, by configuring energy-saving parameters, hardware collaborative optimization is achieved, thereby improving the energy efficiency ratio and balancing performance and energy-saving requirements.

[0142] based on Figure 5 In this embodiment, before determining the multiple control commands corresponding to the load based on the first strategy generation model, the electronic device may further: process the load according to the second strategy generation model to predict the multiple predicted control commands corresponding to the load; and determine the first strategy generation model based on the multiple predicted control commands.

[0143] In one possible implementation, the structure of the first policy generation model is the same as that of the second policy generation model, but the parameters of the first policy generation model and the second policy generation model can be different.

[0144] In one possible implementation, the load output at historical time S104 can be input into the second strategy generation model to determine the corresponding multiple predicted control commands. It should be understood that this embodiment employs a phased training approach. In phase one, the command to output fixed reserved resources remains unchanged. Based on the load at historical time, commands for adjusting the load corresponding to the instance and commands for adjusting the load's resource quota are determined. In phase two, based on the load at historical time, commands for adjusting the load corresponding to the instance, commands for adjusting the load's resource quota, and commands for reserving resources corresponding to the load are determined.

[0145] In this phase, once the training in Phase 1 reaches a certain level of convergence, that is, when the second strategy generation model performs relatively stably in determining the instructions for adjusting the instance corresponding to the load and the instructions for adjusting the resource quota corresponding to the load, and can meet certain performance requirements, Phase 2 training can proceed.

[0146] In the embodiments of this application, a multi-stage training method is adopted in the process of predicting multiple control commands corresponding to the load. This can gradually increase the complexity of adjusting resources, which is beneficial for the second strategy generation model to converge quickly in the early stage of training and establish a stable basic strategy.

[0147] In some embodiments, determining a first strategy generation model based on multiple predicted control commands may include: determining first parameters corresponding to the multiple predicted control commands according to the multiple predicted control commands and a multi-objective function; constraining the multiple predicted control commands and determining second parameters based on the first parameters; and determining the first strategy generation model based on the multiple predicted control commands when the second parameter is maximized.

[0148] In one possible implementation, the multi-objective function may include performance metrics and energy efficiency metrics. Performance metrics may include, for example, load fulfillment rate and response latency compliance rate, while energy efficiency metrics may include, for example, resource utilization rate and power consumption reduction rate. For example, the multi-objective function can be expressed as follows:

[0149]

[0150] in, This indicates the weight of the performance metrics. This indicates the weight of the energy efficiency indicators. This represents the mean of performance metrics. This represents the mean of energy efficiency indicators, where... .

[0151] In one possible implementation, the first parameter is the quantized score of multiple predicted control commands under a multi-objective function. The first parameter is used to evaluate the degree to which the predicted control commands are adapted to the dual objectives of improving performance and optimizing energy efficiency. Corresponding to the first stage mentioned above, the first parameter can calculate the incremental contribution of the commands for adjusting instances and the commands for adjusting resource quotas to improving performance and optimizing energy efficiency. Corresponding to the second stage mentioned above, the first parameter can calculate the incremental contribution of the commands for adjusting instances, the commands for adjusting resource quotas, and the commands for reserving resources to improving performance and optimizing energy efficiency.

[0152] In one possible implementation, the second parameter is the score obtained after combining the first parameter with multiple constraint processing. These constraint processing may include, for example, checking whether control instructions meet node dependencies (e.g., scaling up a POD to physical node A requires node A to be a non-high-load node to avoid violating the low-load node priority deployment rule), checking whether control instructions exceed resource quotas (e.g., reserving 4 CPU cores for physical node B requires a small reservation ratio to avoid crowding out core business resources), and checking whether control instructions meet minimum business requirements (e.g., scaling down a POD to one requires a high query per second (QPS) rate to avoid poor service quality).

[0153] In one possible implementation, the second parameter can be expressed as follows:

[0154]

[0155] in, 'b' represents the first parameter, and 'b' represents the constraint weight. The more the control command conforms to the above constraints, the larger the constraint weight; conversely, the less the control command conforms to the above constraints, the smaller the constraint weight.

[0156] For example, if the first parameter is 85 points but does not meet the above constraints (e.g., POD is deployed to a high-load node), the constraint weight can be 0.8, and the second parameter is 68 (85 x 0.8) points; if the first parameter is high and meets the above constraints, the second parameter can be the same as the first parameter, and further, the second parameter can be slightly larger than the first parameter (e.g., the constraint weight can be 1.2).

[0157] In one possible implementation, the parameters in the second strategy generation model can be adjusted in reverse using the second parameter. When the second parameter is at its maximum, the first strategy generation model can be determined based on the parameters in the second strategy generation model and the structure of the second strategy generation model.

[0158] In one possible implementation, after determining the first strategy generation model, the first strategy generation model can be compressed using methods such as distillation, thereby saving the resources used to store the first strategy generation model.

[0159] In this embodiment, the optimization direction of the policy is guided by a multi-objective function. By combining constraint processing and distillation methods, the safety of the policy is guaranteed, while the deployment efficiency and lightweight level of the first policy generation model are improved, thus achieving a balance between performance and energy efficiency.

[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0161] Figure 6 This is a schematic diagram of the server load prediction device provided in an embodiment of this application. Figure 6 As shown, embodiments of this application also provide a server load prediction device 60, including: an acquisition module 61, a first determination module 62, a second determination module 63, and a prediction module 64, wherein:

[0162] The acquisition module 61 is used to acquire physical layer data, virtual layer data and application layer data of the server cluster.

[0163] The first determining module 62 is used to determine the first relationship between servers in the server cluster based on physical layer data, virtual layer data and application layer data. The first relationship includes the nodes corresponding to the first features of various types of data, as well as the edges between the nodes. The edges are used to indicate the association between the data corresponding to the nodes.

[0164] The second determining module 63 is used to determine the second feature corresponding to the server cluster based on the characteristics of the data corresponding to each node in the first relationship. The second feature includes the spatial relationship and temporal relationship between the data corresponding to each node.

[0165] Prediction module 64 is used to predict the load of the server cluster in future time periods based on the second feature.

[0166] The server load prediction device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0167] In one possible implementation, the second determining module 63 is further configured to:

[0168] Based on the first feature corresponding to each node, the third feature is determined. The third feature includes the spatial and temporal fusion information between the data corresponding to each node in the first relationship.

[0169] Based on the third feature, determine the second feature.

[0170] In one possible implementation, the second determining module 63 is further configured to:

[0171] The time features of the first feature corresponding to each node are extracted to obtain the fourth feature corresponding to each node;

[0172] For any given node, the first feature corresponding to the node is fused with the fourth feature corresponding to the node to obtain the fused feature corresponding to the node.

[0173] The third feature is determined based on the fusion features corresponding to each node.

[0174] In one possible implementation, the second determining module 63 is further configured to:

[0175] Determine multiple first weights among the nodes, which are used to indicate the importance of the nodes at multiple time steps to load forecasting;

[0176] The second feature is determined based on the third feature corresponding to each node and multiple first weights.

[0177] In one possible implementation, the second determining module 63 is further configured to:

[0178] Based on multiple first weights, the third features corresponding to each node are weighted to obtain the weighted features corresponding to each node.

[0179] Based on the weighted features corresponding to each node, a first matrix is ​​determined. The first matrix is ​​used to indicate the degree of dependency between nodes.

[0180] The temporal features of the first matrix are extracted to obtain the second feature.

[0181] In one possible implementation, the first determining module 62 is further configured to:

[0182] Based on physical layer data, virtual layer data, and application layer data, a relationship graph is determined. The relationship graph includes nodes corresponding to each type of data, as well as edges between each node.

[0183] Based on the relationship diagram, determine the first relationship.

[0184] In one possible implementation, the first determining module 62 is further configured to:

[0185] Based on the amount of data transmitted between different types of data, obtain the initial edges between each node;

[0186] Based on the dependencies between various types of data, the initial edges are processed to obtain the edges between each node;

[0187] The relationship graph is determined based on the edges between nodes and the nodes corresponding to the characteristics of various data types.

[0188] In one possible implementation, the acquisition module 61 is used for:

[0189] Retrieve physical layer data, virtual layer data, and application layer data of the server cluster, including:

[0190] Collect physical layer data, virtual layer data, and application layer data;

[0191] Time alignment is performed on physical layer data, virtual layer data, and application layer data based on timestamps in physical layer data, virtual layer data, and application layer data.

[0192] The aligned data is standardized to obtain the physical layer data, virtual layer data, and application layer data of the server cluster.

[0193] In one possible implementation, the server load prediction device 60 further includes: an execution module 65, the execution module 65 being configured to:

[0194] Based on the first strategy generation model, multiple control commands corresponding to the load are determined. These multiple control commands are used to adjust the resource configuration of the server cluster.

[0195] Multiple control instructions are sorted to obtain their execution order.

[0196] Multiple control instructions are executed according to the execution order.

[0197] In one possible implementation, the server load prediction device 60 further includes a processing module 66, which is configured to:

[0198] Based on the sorted control commands, the configuration parameters are determined, including the number of instances, resource quota threshold, and resource reservation ratio.

[0199] The resource configuration of the server cluster is processed according to the configuration parameters.

[0200] In one possible implementation, the server load prediction device 60 further includes: a third determining module 67, the third determining module 67 being configured to:

[0201] Based on the second strategy generation model, the load is processed, and multiple predicted control commands corresponding to the load are predicted.

[0202] Based on multiple predicted control commands, a first strategy generation model is determined.

[0203] In one possible implementation, the third determining module 67 is further configured to:

[0204] Based on multiple predicted control commands and multiple objective functions, determine the first parameter corresponding to the multiple predicted control commands;

[0205] Constrain multiple predicted control commands and determine the second parameter based on the first parameter;

[0206] When the second parameter is at its maximum, the first strategy generation model is determined based on multiple predicted control commands.

[0207] The server load prediction device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0208] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 71 and a memory 72. Optionally, the electronic device 70 further includes a communication component 73. The processor 71, memory 72, and communication component 73 are connected via a bus.

[0209] In a specific implementation, at least one processor 71 executes computer execution instructions stored in memory 72, causing at least one processor 71 to execute the above-described server load prediction method embodiment.

[0210] The specific implementation process of processor 71 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0211] In the above embodiments, it should be understood that the processor can be a CPU, or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules within the processor.

[0212] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0213] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0214] Embodiments of this application also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described server load prediction method embodiments at runtime.

[0215] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0216] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described server load prediction method embodiments.

[0217] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described server load prediction method embodiments.

[0218] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0219] The above provides a detailed description of a server load prediction method and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server load prediction method, characterized by, The method comprises: acquiring physical layer data, virtual layer data and application layer data of a server cluster; determining a first relationship between servers in the server cluster according to the physical layer data, the virtual layer data and the application layer data, the first relationship comprising nodes corresponding to first features of various types of data and edges between nodes, the edges being used to indicate an association relationship between data corresponding to the nodes; determining a second feature corresponding to the server cluster according to features of data corresponding to nodes in the first relationship, the second feature comprising spatial and temporal relationships between data corresponding to the nodes; predicting a load of the server cluster in a future period according to the second feature; the determining of the second feature corresponding to the server cluster according to the features of data corresponding to the nodes in the first relationship comprises: determining a third feature according to the first features corresponding to the nodes, the third feature comprising fusion information of space and time between data corresponding to the nodes in the first relationship; determining the second feature according to the third feature; the determining of the third feature according to the first features corresponding to the nodes comprises: extracting time features of the first features corresponding to the nodes to obtain fourth features corresponding to the nodes; fusing the first features corresponding to the nodes with the fourth features corresponding to the nodes to obtain fusion features corresponding to the nodes for any one node; determining the third feature according to the fusion features corresponding to the nodes; the determining of the second feature according to the third feature comprises: determining a plurality of first weights between the nodes, the first weights being used to indicate an importance degree of the nodes in a plurality of time steps for the load prediction; determining the second feature according to the third features corresponding to the nodes and the plurality of first weights.

2. The method of claim 1, wherein, the determining of the second feature according to the third features corresponding to the nodes and the plurality of first weights comprises: weighting the third features corresponding to the nodes according to the plurality of first weights to obtain weighted features corresponding to the nodes; determining a first matrix according to the weighted features corresponding to the nodes, the first matrix being used to indicate a dependence degree between the nodes; extracting time sequence features of the first matrix to obtain the second feature.

3. The method of claim 1, wherein, the determining of the first relationship between servers in the server cluster according to the physical layer data, the virtual layer data and the application layer data comprises: determining a relationship graph according to the physical layer data, the virtual layer data and the application layer data, the relationship graph comprising nodes corresponding to various types of data and edges between the nodes; determining the first relationship according to the relationship graph.

4. The method of claim 3, wherein, the determining of the relationship graph according to the physical layer data, the virtual layer data and the application layer data comprises: acquiring initial edges between the nodes according to data transmission amounts between various types of data; processing the initial edges based on dependence relationships between various types of data to acquire edges between the nodes; determining the relationship graph according to the edges between the nodes and nodes corresponding to features of various types of data.

5. The method of claim 1, wherein, The physical layer data, the virtual layer data and the application layer data of the server cluster are acquired, including: The physical layer data, the virtual layer data and the application layer data are collected; The physical layer data, the virtual layer data and the application layer data are time-aligned based on the time stamp in the physical layer data, the time stamp of the virtual layer data and the time stamp of the application layer data; The aligned data is standardized to acquire the physical layer data, the virtual layer data and the application layer data of the server cluster.

6. The method of claim 1, wherein, After the second feature, the load of the server cluster in the future period is predicted, and the method further includes: Based on the first strategy generation model, a plurality of control instructions corresponding to the load are determined, and the plurality of control instructions are used to adjust the resource configuration of the server cluster; The plurality of control instructions are sorted to obtain an execution order of the plurality of control instructions; The plurality of control instructions are executed according to the execution order.

7. The method of claim 6, wherein, After the plurality of control instructions are executed, the method further includes: Based on the plurality of control instructions after sorting, a configuration parameter is determined, and the configuration parameter includes an instance number, a resource quota threshold and a resource reservation ratio; The resource configuration of the server cluster is processed according to the configuration parameter.

8. The method of claim 6, wherein, Before the first strategy generation model is determined based on the plurality of control instructions corresponding to the load, the method further includes: According to the second strategy generation model, the load is processed to predict a plurality of predicted control instructions corresponding to the load; The first strategy generation model is determined based on the plurality of predicted control instructions.

9. The method of claim 8, wherein, The first strategy generation model is determined based on the plurality of predicted control instructions, including: According to the plurality of predicted control instructions and a multi-objective function, a first parameter corresponding to the plurality of predicted control instructions is determined; The plurality of predicted control instructions are constrained and processed, and a second parameter is determined based on the first parameter; In the case that the second parameter is maximum, the first strategy generation model is determined based on the plurality of predicted control instructions.

10. An electronic device, comprising: It includes: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the server load prediction method according to any one of claims 1 to 9.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the server load prediction method according to any one of claims 1 to 9.

12. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the server load prediction method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Service node performance prediction method, device, equipment, readable medium and product

    CN119718860A

  • Method for deploying an application workload on a cluster

    US20200404076A1