Server load prediction method, electronic device, storage medium and program product
By acquiring and analyzing multi-layer data of the server cluster, using temporal convolutional networks and graph attention networks to predict load and adjust resource allocation, the service delay problem caused by sudden increase in load in the server cluster is solved, and the service quality and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202511244893.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-09-02
AI Technical Summary
The server cluster's load increases due to factors such as business traffic fluctuations, leading to resource exhaustion and service response delays, affecting service quality.
By obtaining the physical layer, virtual layer and application layer data of the server cluster, the association and characteristics between the servers are determined, and the load in the future period is predicted using the temporal convolutional network and graph attention network, and resource allocation is adjusted to avoid sudden increases in load.
Accurately predict server cluster load, reduce service response delay, and improve service quality and resource utilization efficiency.
Smart Images

Figure CN120723591A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a server load prediction method, electronic device, storage medium, and program product. Background Art
[0002] A server cluster is formed by multiple independent servers connected through a network. Under unified management and scheduling, these servers can integrate and share computing power and storage resources. Server clusters are widely used in cloud computing platforms, large-scale distributed systems, etc.
[0003] However, during the operation of a server cluster, the server load often increases suddenly due to factors such as business traffic fluctuations. The increase in load can easily lead to server resource exhaustion and saturation of server processing capacity, which in turn causes problems such as service response delays, seriously affecting the service quality of the server cluster. Summary of the Invention
[0004] The present application provides a server load prediction method, electronic device, storage medium and program product to at least solve the problem of slow service response of server clusters in related technologies.
[0005] This application provides a server load prediction method, including:
[0006] Obtain physical layer data, virtual layer data, and application layer data of the server cluster;
[0007] Determining, based on the physical layer data, the virtual layer data, and the application layer data, a first relationship between servers in the server cluster, the first relationship including nodes corresponding to first features of each type of data and edges between the nodes, where the edges are used to indicate associations between the data corresponding to the nodes;
[0008] Determining a second feature corresponding to the server cluster based on features of the data corresponding to each node in the first relationship, where the second feature includes a spatial relationship and a temporal relationship between the data corresponding to each node;
[0009] According to the second feature, the load of the server cluster in the future period is predicted.
[0010] The present application also provides a server load prediction device, comprising: an acquisition module, a first determination module, a second determination module, and a prediction module, wherein:
[0011] The acquisition module is used to obtain the physical layer data, virtual layer data and application layer data of the server cluster.
[0012] The first determination module is used to determine the first relationship between the servers in the server cluster based on the physical layer data, the virtual layer data and the application layer data. The first relationship includes the nodes corresponding to the first characteristics of each type of data and the edges between the nodes. The edges are used to indicate the association relationship between the data corresponding to the nodes.
[0013] The second determining module is used to determine a second feature corresponding to the server cluster according to features of the data corresponding to each node in the first relationship, where the second feature includes a spatial relationship and a temporal relationship between the data corresponding to each node.
[0014] The prediction module is used to predict the load of the server cluster in a future period according to the second feature.
[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server load prediction methods when executing the computer program.
[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server load prediction methods are implemented.
[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server load prediction methods when executed by a processor.
[0018] Through this application, since the spatial and temporal relationships between the data corresponding to each node in the server cluster can be used to accurately predict the load of the server cluster in the future period, the server cluster can adjust resource allocation according to the predicted load value, thereby solving the technical problem of slow service response of the server cluster, and achieving the technical effect of reducing service response delay and improving the service quality of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A flow chart of a server load prediction method provided in an embodiment of the present application;
[0021] Figure 2 A schematic flow chart of a method for determining a second feature provided in an embodiment of the present application;
[0022] Figure 3 A flowchart of a method for determining a first relationship provided in an embodiment of the present application;
[0023] Figure 4 A flowchart of a method for obtaining various types of data provided in an embodiment of the present application;
[0024] Figure 5 A method for adjusting resource configuration provided in an embodiment of the present application;
[0025] Figure 6 A schematic diagram of the structure of a server load prediction device provided in an embodiment of the present application;
[0026] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0028] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0029] Glossary:
[0030] Time step: A time step is a fixed time interval unit set to unify the time granularity of various types of data. For example, a time step of 1 second (s) indicates that various types of data are collected and unified at intervals of 1 second.
[0031] Feature dimension: The feature dimension refers to the number of core feature indicators aggregated by a single node in a single time step. For example, a feature dimension of 3 indicates that a single node can describe its own status through three values (for example, fan speed, power consumption, and memory usage) in a single time step.
[0032] A server cluster is formed by multiple independent servers connected through a network. Under unified management and scheduling, these servers can integrate and share computing power and storage resources.
[0033] However, during the operation of a server cluster, the load status of the server cluster often increases suddenly due to factors such as business traffic fluctuations (such as peak access, limited-time activities), service dependency linkage (such as cascading requests in the microservice call chain), or hardware status changes (such as node performance degradation). The increase in load can easily lead to server resource exhaustion and saturation of server processing capacity, which in turn causes service response delays and other problems, seriously affecting the service quality of the server cluster.
[0034] In response to the above problems, in an embodiment of the present application, the physical layer data, virtual layer data and application layer data of the server cluster are obtained; based on the physical layer data, virtual layer data and application layer data, the first relationship between the servers in the server cluster is determined, the first relationship includes nodes corresponding to the first characteristics of each type of data, and edges between the nodes, and the edges are used to indicate the association relationship between the data corresponding to the nodes; based on the characteristics of the data corresponding to each node in the first relationship, the second characteristics corresponding to the server cluster are determined, the second characteristics include the spatial relationship and time relationship between the data corresponding to each node; based on the second characteristics, the load of the server cluster in the future time period is predicted. In the above method, the load of the server cluster in the future time period can be accurately predicted through the spatial relationship and time relationship between the data corresponding to each node in the server cluster, so that the server cluster can adjust resource allocation according to the predicted load value to avoid server resource exhaustion caused by a sudden increase in load, thereby reducing service response delay and improving the service quality of the server cluster.
[0035] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0036] Figure 1 A flow chart of the server load prediction method provided in the embodiment of the present application is shown as follows: Figure 1 As shown, an embodiment of the present application provides a server load prediction method, which is described in detail as follows:
[0037] S101. Obtain physical layer data, virtual layer data, and application layer data of a server cluster.
[0038] The execution subject of the embodiment of the present application can be an electronic device, or a server load prediction device set in the electronic device. The server load prediction device can be implemented by software or a combination of software and hardware.
[0039] In one possible implementation, the physical layer data of the server cluster includes, but is not limited to: the central processing unit (CPU) temperature, power consumption, fan speed of each node in the server, and the identifier corresponding to the physical node;
[0040] The server's virtual layer data includes, but is not limited to, CPU utilization, memory usage, input / output (I / O) throughput, and identifiers of each virtual node in the server.
[0041] The application layer data of the server includes but is not limited to: the request duration of the microservice call chain of each node in the server, the dependency relationship between each node, and the corresponding identifier of the application service node;
[0042] Among them, physical nodes refer to the physical hardware devices of the server cluster (such as physical servers), virtual nodes refer to logical nodes abstracted from physical nodes through virtualization technology (such as virtual monitors (Hypervisors)), such as virtual machines and containers, and application service nodes refer to logical nodes in the server cluster that carry application programs, services, or microservice instances.
[0043] S102. Determine a first relationship between servers in a server cluster based on physical layer data, virtual layer data, and application layer data. The first relationship includes nodes corresponding to first features of each type of data and edges between the nodes. The edges are used to indicate an association relationship between data corresponding to the nodes.
[0044] In one possible implementation, the node corresponding to the first feature of the physical layer data is a physical node, the node corresponding to the first feature of the virtual layer data is a virtual node, and the node corresponding to the first feature of the application layer data is an application service node.
[0045] In a possible implementation, the first feature corresponds to the core attributes of the three types of data in S101 and integrates the spatial correlation features between the nodes.
[0046] The core attributes of the three types of data include: encoding information of physical layer data, encoding information of virtual layer data, and encoding information of application layer data. The encoding information is the feature vector after encoding the data.
[0047] In one possible implementation, the format of the first relation is a mathematical and structured data format that can be stored, calculated and processed by a computer. Exemplarily, the format of the first relation is a multi-dimensional vector, matrix, dictionary or table.
[0048] For example, taking the format of the first relation as a dictionary, each node, the first feature corresponding to each node, and the edge can be stored through key-value pairs. In one implementation, the nodes corresponding to the first feature of each type of data can be stored through keys, and the first feature corresponding to each node and the edge can be stored through values.
[0049] S103: Determine a second feature corresponding to the server cluster based on features of the data corresponding to each node in the first relationship, where the second feature includes a spatial relationship and a temporal relationship between the data corresponding to each node.
[0050] In a possible implementation, the spatial relationship between the data corresponding to each node refers to: the topological association relationship between each node and the data corresponding to each node, for example, the dependency relationship between each node, the association relationship between each node, etc.
[0051] In one possible implementation, the temporal relationship between the data corresponding to each node refers to the dynamic evolution pattern of the data corresponding to each node in the time dimension. The dynamic evolution pattern refers to the fluctuation characteristics of the node data over time, that is, the fluctuation characteristics of the node data at different time steps, such as the daily peaks and troughs of physical node power consumption and the periodic growth of virtual node memory usage.
[0052] S104: Predict the load of the server cluster in the future period based on the second feature.
[0053] In a possible implementation, the load of the server cluster in the future period can be predicted based on the second features of the server cluster at the current time step and the historical time steps.
[0054] In a possible implementation, the second feature can be predicted based on a dynamic weight fuser to obtain the load.
[0055] In an embodiment of the present application, the physical layer data, virtual layer data and application layer data of the server cluster are obtained; based on the physical layer data, virtual layer data and application layer data, a first relationship between the servers in the server cluster is determined, the first relationship includes nodes corresponding to the first characteristics of each type of data, and edges between the nodes, the edges are used to indicate the association relationship between the data corresponding to the nodes; based on the characteristics of the data corresponding to each node in the first relationship, a second characteristic corresponding to the server cluster is determined, the second characteristic includes the spatial relationship and time relationship between the data corresponding to each node; based on the second characteristic, the load of the server cluster in the future time period is predicted. In the above method, the load of the server cluster in the future time period can be accurately predicted through the spatial relationship and time relationship between the data corresponding to each node in the server cluster, so that the server cluster can adjust resource allocation according to the predicted load value to avoid server resource exhaustion caused by a sudden increase in load, thereby reducing service response delay and improving the service quality of the server cluster.
[0056] Based on any of the above embodiments, Figure 2 , the method for determining the second feature ( Figure 1 S103 in the embodiment is described in detail.
[0057] Figure 2 For a flow chart of a method for determining a second feature provided in an embodiment of the present application, please refer to Figure 2 , the method may include:
[0058] S201. Determine a third feature based on the first feature corresponding to each node, where the third feature includes spatial and temporal fusion information between the data corresponding to each node in the first relationship.
[0059] In some embodiments, the time feature of the first feature corresponding to each node is extracted to obtain the fourth feature corresponding to each node; for any node, the first feature corresponding to the node and the fourth feature corresponding to the node are fused to obtain the fused feature corresponding to the node; and the third feature is determined based on the fused feature corresponding to each node.
[0060] In one possible implementation, the fourth feature corresponding to each node includes the dynamic evolution law of the first feature corresponding to each node in the time dimension, that is, the pattern and trend of the first feature of each node changing over time.
[0061] In one possible implementation, a temporal convolutional network (TCN) may be used to extract the time feature of the first feature corresponding to each node to obtain the fourth feature corresponding to each node.
[0062] In one possible implementation, multiple layers of dilated causal convolutional layers can be stacked in the temporal convolutional network to cover a sufficiently long time step. Residual connections can also be set in the temporal convolutional network (for example, the output of the first layer of convolution, in addition to entering the second layer of convolution, is also directly added to the output of the second layer of convolution through the residual path). This allows the gradient to be directly backpropagated to the shallow layer through the residual path, avoiding gradient vanishing, so that the deep network of the temporal convolutional network can also effectively learn the temporal characteristics of each node.
[0063] In one possible implementation, the fusion feature can be expressed as follows:
[0064]
[0065] in, represents the gate value, represents element-wise multiplication, Represents the first feature, It represents the fourth feature, wherein the gate value can be pre-set, and the gate value can take a value in the interval [0,1]. Exemplarily, the gate value is 0.5.
[0066] In one possible implementation, the third feature can be expressed as follows:
[0067]
[0068] in, Indicates that and For vertical splicing, Indicates the fusion features corresponding to each physical node, Indicates the fusion features corresponding to each virtual node, , N represents the total number of physical nodes and virtual nodes, T represents the number of time steps, and d represents the feature dimension.
[0069] In the embodiment of the present application, since the first feature includes the spatial correlation features between each node, and the fourth feature includes the dynamic evolution law of the first feature corresponding to each node in the time dimension, the fusion feature (and the third feature) includes the spatial and temporal fusion information between the data corresponding to each node.
[0070] S202: Determine the second feature based on the third feature.
[0071] In some embodiments, multiple first weights between nodes can be determined, where the first weights are used to indicate the importance of nodes to load prediction at multiple time steps; and the second feature is determined based on the third feature corresponding to each node and the multiple first weights.
[0072] The specific implementation method of determining the second feature according to the third feature and multiple first weights corresponding to each node is as follows:
[0073] In some embodiments, the third features corresponding to each node are weighted according to multiple first weights to obtain weighted features corresponding to each node; a first matrix is determined according to the weighted features corresponding to each node, and the first matrix is used to indicate the degree of dependence between nodes; and the time series features of the first matrix are extracted to obtain the second features.
[0074] In one possible implementation, the importance or influence of each node at each time step can be determined by analyzing the third feature in the time dimension, and the influence or importance of each node at each time step can be represented by a first weight.
[0075] In one possible implementation, the first weight may be calculated by one or more of the following methods, but not limited to:
[0076] The importance of different time steps is estimated by analyzing patterns such as periodicity and trends in historical data. If peak loads frequently occur in a certain time period, the data during this period may be considered more important, and the first weight value corresponding to this period will be larger.
[0077] The importance distribution in the time dimension is learned through the attention mechanism in the deep learning framework, and the weights of different time steps are adaptively adjusted to determine the first weight.
[0078] The weight of each node at different time steps, that is, the first weight, is determined by a statistical model (for example, an autoregressive integrated moving average model (ARIMA) or an exponential smoothing model).
[0079] In a possible implementation, weighted processing is performed on the third feature corresponding to each node according to the first weight, so as to highlight the importance of data at certain time steps for load prediction.
[0080] In one possible implementation, the inter-node attention matrix can be determined based on the weighted features, the spatial weight vector of each node can be determined based on the inter-node attention matrix, and then the first matrix can be determined based on the spatial weight vector of each node.
[0081] The inter-node attention matrix can be expressed as follows:
[0082]
[0083] in, represents the spatial attention weight from node j to node i, which is used to measure the influence of node j on node i. and Both are linear transformation matrices, used to generate the query vector of node i and the key vector of node j , and They represent the feature vectors of node i and node j in the weighted features respectively, and sqrt(d) represents the square root operation performed on the feature dimension d.
[0084] Exemplarily, the spatial weight vector of node i can be expressed as follows:
[0085]
[0086] in, represents the spatial attention weight between node i and the features of each associated node, Express Perform a maximum pooling operation to determine the node j that has the most significant impact on node i. Indicates normalization of the maximum pooling result.
[0087] In a possible implementation, the spatial weight vectors of each node may be concatenated, and the concatenated matrix may be determined as the first matrix.
[0088] In one possible implementation, the time series features of the first matrix can be extracted through a time series convolutional network to obtain the second features.
[0089] It should be noted that in S104, the load of the server cluster in the future period can be predicted based on the dynamic weight fuser and the second feature of the server cluster at the current time step and the historical time step. The dynamic weight fuser can fuse the first matrix and the second feature, predict the second feature, and obtain the load.
[0090] In an embodiment of the present application, by combining the graph attention network and the temporal convolutional network, the joint modeling of spatial topology and time series is realized, which can effectively extract the correlation between nodes and their characteristics evolving over time. In addition, by fusing the first feature (the spatial correlation feature between nodes) and the fourth feature (the dynamic evolution law of the first feature in the time dimension) through the gated value, the robustness and expressiveness of the feature representation can be improved, so that feature representation can still be effectively performed in a complex server cluster environment (for example: bursty traffic (flash sale activities, etc.)).
[0091] In addition, in an embodiment of the present application, by combining the first weight (the importance of nodes in multiple time steps to load prediction) with the first matrix (the degree of dependence between nodes) to perform weighted fusion on the third feature, the impact of key historical moments can be highlighted in the time dimension, and the role of important nodes can be strengthened in the spatial dimension, thereby improving the accuracy of load prediction. The dynamic weight fuser further integrates multi-dimensional features to make the prediction results more real-time and forward-looking, which helps to achieve better resource scheduling strategies.
[0092] Based on any of the above embodiments, Figure 3 , the method for determining the first relationship ( Figure 1 S102 in the embodiment is described in detail.
[0093] Figure 3 For a flow chart of a method for determining a first relationship provided in an embodiment of the present application, please refer to Figure 3 , the method may include:
[0094] S301. Determine a relationship graph based on physical layer data, virtual layer data, and application layer data. The relationship graph includes nodes corresponding to various types of data and edges between the nodes.
[0095] In some embodiments, the initial edges between the nodes can be obtained based on the data transmission volume between the various types of data; the initial edges are processed based on the dependency relationship between the various types of data to obtain the edges between the nodes; and the relationship graph is determined based on the edges between the nodes and the nodes corresponding to the characteristics of the various types of data.
[0096] In one possible implementation, the data transmission volume between the physical node and the virtual node can be determined by determining the real-time data interaction volume (such as the CPU resource scheduling instruction transmission volume and the memory page exchange data volume) when the physical server allocates resources to the virtual machine or container;
[0097] You can also determine the data volume between the virtual node and the application service node by determining the call data volume between the container and the microservice instance (such as the application programming interface (API) request packet size and response data volume);
[0098] You can also determine the amount of data transferred between application service nodes by determining the amount of cross-instance data transferred in the microservice call chain (such as the amount of response data).
[0099] In one possible implementation, the dependency between physical nodes and virtual nodes can be determined. For example, if a virtual node is completely dependent on the hardware resources of a physical node (such as a virtual machine bound to a physical CPU core), the corresponding initial edge will be increased by 20%-50%;
[0100] The dependency relationship between the virtual node and the application service node may also be determined. For example, if the application service node is only deployed on a specific virtual node (such as a POD-bound container), the corresponding initial edge may be promoted.
[0101] In one possible implementation, a dynamic graph can be determined based on the edges between nodes obtained in real time and the nodes corresponding to the features of various types of data. The dynamic graph is then structured and represented using a graph feature encoding method. The dynamic graph is processed from an abstract conceptual graph into a relationship graph in a mathematical, structured data format that can be stored, calculated, and processed by computers.
[0102] S302: Determine a first relationship according to the relationship diagram.
[0103] In one possible implementation, a graph attention network (GAT) can be used to extract features from the relationship graph and determine the first relationship. GAT can be used to obtain the node's local structural characteristics (such as the distribution of neighboring nodes), global position (importance in the entire graph), and node attributes (such as CPU utilization and memory usage). Neighboring nodes are one or more nodes adjacent to the current node.
[0104] In an embodiment of the present application, by adopting a dynamic graph construction method to model each node, it is possible to capture not only the static topological structure, but also the dynamic changes in the relationship between nodes during operation. By dual analysis of the data transmission volume and dependency relationship between each node, the prediction bias is reduced and the semantic expression ability of the edge is enhanced, which helps to more accurately predict the chain failure scenario, thereby more accurately portraying the actual interaction mode between servers, and providing high-quality graph structure input for subsequent feature extraction and load prediction.
[0105] Based on any of the above embodiments, Figure 4 , a method for obtaining the physical layer data, virtual layer data and application layer data of a server cluster ( Figure 1 S101 in the embodiment is described in detail.
[0106] Figure 4 For a flow chart of a method for obtaining various types of data provided in the embodiment of this application, please refer to Figure 4 , the method may include:
[0107] S401: Collect physical layer data, virtual layer data, and application layer data.
[0108] In one possible implementation, physical layer data collected by sensors of the physical hardware of the server cluster can be read through an intelligent platform management interface (IPMI). Exemplarily, the storage format of each piece of collected physical layer data can be a physical node identifier + physical indicator name + physical value + collection timestamp (accurate to milliseconds). The physical indicator name can be the CPU temperature mentioned in the above embodiment, and the corresponding physical value can be 60 degrees Celsius (°C).
[0109] Alternatively, virtual layer data may be periodically collected through a kernel-based virtual machine (KVM). For example, the storage format of each piece of collected virtual layer data may be virtual node identifier + virtual indicator name + virtual value + collection timestamp. The virtual indicator name may be the CPU utilization mentioned in the above embodiment, and the corresponding virtual value may be 80%.
[0110] Application layer data can also be collected through distributed application performance monitoring tools (SkyWalking), distributed tracing tools (Pinpoint), etc., and application layer data can also be collected through embedding points. Exemplarily, the storage format of each collected application layer data can be application node identifier + application indicator name + application value + collection timestamp, where the application indicator name can be the request duration of the microservice call chain mentioned in the above embodiment, and the corresponding application value can be 150 milliseconds (ms);
[0111] S402 : Time-align the physical layer data, the virtual layer data, and the application layer data based on the timestamp in the physical layer data, the timestamp in the virtual layer data, and the timestamp in the application layer data.
[0112] In one possible implementation, the time step grid can be divided into units of preset time steps (for example, 10s), and then the timestamp of each data piece is mapped to the corresponding time step grid, thereby mapping the data corresponding to the timestamp to the corresponding time step grid.
[0113] In one possible implementation, time steps with no data (for example, no virtual layer data within a certain 10 seconds) can be filled with data from the previous time step or a value of 0.
[0114] In one possible implementation, for multiple pieces of similar data within the same time step (for example, the CPU temperature of a physical node is collected three times within 10 seconds), the average or maximum value of the multiple pieces of similar data can be taken as the data of this type within the time step.
[0115] S403: Standardize the aligned data to obtain the physical layer data, virtual layer data, and application layer data of the server cluster.
[0116] In one possible implementation, the aligned data can be normalized by methods such as min-max normalization or z-score standardization.
[0117] In the embodiment of the present application, by collecting multi-level data at the physical layer, virtual layer and application layer and performing timestamp normalization processing on them, the time deviation problem between different data sources can be effectively eliminated, and the accuracy of subsequent modeling and prediction can be improved. In addition, standardization of various types of data can eliminate the dimensional differences of various indicators, making different types of data comparable and fusible in subsequent models, laying the foundation for building unified data.
[0118] Based on the above embodiment, after predicting the load of the server cluster in the future period according to the second feature, the resource configuration of the server cluster can also be adjusted. For example, Figure 5 A method for adjusting resource allocation is provided in the embodiment of this application. For details, see Figure 5 , the method is as follows:
[0119] S501: Determine a plurality of control instructions corresponding to the load based on a first policy generation model, where the plurality of control instructions are used to adjust resource configuration of a server cluster.
[0120] In a possible implementation, the multiple control instructions may include: an instruction to adjust an instance, an instruction to adjust a resource quota, and an instruction to reserve resources.
[0121] Among them, adjusting instances refers to increasing or decreasing the number or migrating the location of virtual machines, containers and / or application layer service instances; adjusting resource quotas refers to adjusting the deployment method of each service or application on different nodes; reserving resources refers to reserving CPU cores and memory for key physical nodes, and / or reserving resource pools for specific virtual instances or application services, etc.
[0122] S502: Sort the multiple control instructions to obtain an execution order of the multiple control instructions.
[0123] In a possible implementation, the multiple control instructions determined in S501 may be processed based on a control instruction priority arbitration method to obtain an execution order of the multiple control instructions. The control instruction priority arbitration method may include control instruction priority rules and conflict detection.
[0124] In one possible implementation, a set of clear control instruction priority rules can be set in advance. For example, when resources are tight, ensuring the stable operation of core business services may be given the highest priority; when resources are sufficient, optimizing costs may become the main consideration.
[0125] For example, the control instruction priority rule may be:
[0126] First priority: instructions for reserving resources; second priority: instructions for adjusting instances; third priority: instructions for adjusting resource quotas.
[0127] Setting the first priority to the instruction for reserving resources can avoid the problem of failure to add instances later due to lack of available resources if the management resources are not locked in advance; setting the second priority to the instruction for adjusting instances can avoid the problem of adjusting resource quotas first but the total number of instances is insufficient to cope with high loads; setting the third priority to the instruction for adjusting resource quotas can adjust the quota according to the load pressure of each instance after the number of instances is determined, avoiding resource waste (for example, if the total number of instances is sufficient, there is no need to expand the resource quota first).
[0128] In one possible implementation, each control instruction can be checked for potential conflicts. For example, one control instruction might increase the number of instances to cope with high load, while another might reduce resource allocation to save costs. In this case, a pre-defined priority rule determines which control instruction should be executed.
[0129] It should be noted that as new data is input (such as the latest load forecast value, real-time collected data, etc.), the electronic equipment can automatically re-evaluate and adjust the control instruction priority rules.
[0130] S503: Execute multiple control instructions according to the execution order.
[0131] In a possible implementation, an ordered control instruction execution sequence may be first determined according to the execution order of the multiple control instructions, and then the multiple control instructions may be executed according to the ordered execution sequence.
[0132] In one possible implementation, based on the determined execution order, an ordered control instruction execution sequence is determined, so that the control instructions are executed safely in order without causing system instability or other negative effects.
[0133] In one possible implementation, if the instruction to reserve resources is executed, the buffer resource configuration can be obtained; if the instruction to adjust the instance is executed, the updated node list can be obtained; if the instruction to adjust the resource quota is executed, a new resource allocation plan can be obtained.
[0134] The updated node list includes the available computing nodes in the server cluster and their current status (such as load status, health status, etc.); the new resource allocation plan refers to the deployment method of each service or application on different nodes.
[0135] In an embodiment of the present application, by sorting multiple control instructions and executing the sorted control instructions in order, conflicts and oscillations in the resource adjustment process are avoided, a reasonable execution order is determined by the control instruction priority rules, and combined with instance adjustment, quota change and resource reservation strategy, refined control of resource allocation is achieved.
[0136] based on Figure 5 In an embodiment, after executing multiple control instructions, the following steps may be performed: determining configuration parameters based on the sorted multiple control instructions, the configuration parameters including the number of instances, resource quota threshold, and resource reservation ratio; and processing resource configuration of the server cluster according to the configuration parameters.
[0137] In one possible implementation, configuration parameters refer to specific settings or policies used to optimize the energy efficiency of a server cluster. Configuration parameters may include controlling operations at the hardware level, such as adjusting memory frequency, adjusting fan speed, etc.
[0138] Energy-saving configuration parameters can determine specific energy-saving measures based on the status of each node. For example, for nodes with lower loads in the updated node list, a more aggressive energy-saving strategy can be adopted, such as setting the nodes with lower loads in the updated node list to deep sleep mode or performing CPU frequency reduction operations on the nodes with lower loads; while for nodes with higher loads, more conservative energy-saving measures can be taken to achieve better performance.
[0139] Energy-saving configuration parameters can determine specific energy-saving measures based on the resource allocation scheme. For example, if certain nodes are designated as standby nodes to handle burst traffic, the nodes can be set to a low-power standby state until they are actually needed.
[0140] The energy-saving configuration parameters can determine specific energy-saving measures based on the buffer resource configuration. For example, reserved resources in the buffer resource configuration that have not been used for a long time (eg, no sudden load occupation within 2 hours) can be set to deep sleep mode.
[0141] In the embodiment of the present application, hardware collaborative tuning is achieved through energy-saving configuration parameters, thereby improving energy efficiency and taking into account both performance and energy-saving requirements.
[0142] based on Figure 5 In an embodiment, before determining multiple control instructions corresponding to the load based on the first strategy generation model, the electronic device may also: process the load according to the second strategy generation model to predict multiple predicted control instructions corresponding to the load; and determine the first strategy generation model based on the multiple predicted control instructions.
[0143] In a possible implementation, the structure of the first policy generation model is the same as that of the second policy generation model, but the parameters of the first policy generation model and the parameters of the second policy generation model may be different.
[0144] In one possible implementation, the load output at the historical moment S104 can be input into the second strategy generation model to determine the corresponding multiple predicted control instructions. It should be understood that the embodiment of the present application adopts a staged training. In stage one, the instruction for outputting a fixed reserved resource remains unchanged, and based on the load at the historical moment, the instruction for adjusting the instance corresponding to the load and the instruction for adjusting the resource quota corresponding to the load are determined; in stage two, based on the load at the historical moment, the instruction for adjusting the instance corresponding to the load, the instruction for adjusting the resource quota corresponding to the load, and the instruction for reserving resources corresponding to the load are determined.
[0145] Among them, when the training of stage one reaches a certain degree of convergence, that is, the second strategy generation model performs relatively stably in determining the instructions for adjusting the instance corresponding to the load and the instructions for adjusting the resource quota corresponding to the load, and can meet certain performance requirements, stage two training can be carried out.
[0146] In an embodiment of the present application, in the process of predicting multiple predicted control instructions corresponding to the load, a multi-stage training method is adopted to gradually increase the complexity of adjusting resources, which is conducive to the rapid convergence of the second strategy generation model in the early stage of training and the establishment of a stable basic strategy.
[0147] In some embodiments, determining a first strategy generation model based on multiple predicted control instructions may include: determining first parameters corresponding to multiple predicted control instructions based on multiple predicted control instructions and a multi-objective function; performing constraint processing on the multiple predicted control instructions, and determining a second parameter based on the first parameter; and determining the first strategy generation model based on multiple predicted control instructions when the second parameter is maximum.
[0148] In one possible implementation, the multi-objective function may include performance indicators and energy efficiency indicators, wherein the performance indicators include, for example, load satisfaction rate and response delay compliance rate, and the energy efficiency indicators include, for example, resource utilization rate and power consumption reduction rate. For example, the multi-objective function may be expressed as follows:
[0149]
[0150] in, Indicates the weight corresponding to the performance indicator, Indicates the weight corresponding to energy efficiency indicators, Represents the mean value of performance indicators, Represents the mean value of energy efficiency indicators, where .
[0151] In one possible implementation, the first parameter is the quantitative score of multiple predicted control instructions under a multi-objective function. The first parameter is used to evaluate the adaptability of the predicted control instructions to the dual objectives of improving performance and optimizing energy efficiency. Corresponding to the above-mentioned stage one, the first parameter can calculate the incremental contribution of instructions for adjusting instances and instructions for adjusting resource quotas to improving performance and optimizing energy efficiency; corresponding to the above-mentioned stage two, the first parameter can calculate the incremental contribution of instructions for adjusting instances, instructions for adjusting resource quotas, and instructions for reserving resources to improving performance and optimizing energy efficiency.
[0152] In one possible implementation, the second parameter is the score obtained by combining the first parameter with multiple constraints. Constraint processing may include, for example, verifying whether the control instruction complies with node dependencies (e.g., when scaling a POD to physical node A, node A must be a non-highly loaded node to avoid violating the rule of prioritizing low-load nodes for deployment); verifying whether the control instruction exceeds resource quotas (e.g., when reserving a 4-core CPU for physical node B, a small reserved ratio must be met to avoid crowding out core business resources); and verifying whether the control instruction meets minimum business requirements (e.g., when scaling a POD to 1, a high number of queries per second (QPS) in the call chain must be met to avoid poor service quality).
[0153] In one possible implementation, the second parameter can be expressed as follows:
[0154]
[0155] in, represents the first parameter, b represents the constraint weight. The more the control instruction conforms to the above constraints, the greater the constraint weight. Conversely, the less the control instruction conforms to the above constraints, the smaller the constraint weight.
[0156] For example, when the first parameter is 85 points but does not meet the above constraints (for example, the POD is deployed to a high-load node), the constraint weight can be 0.8 and the second parameter can be 68 (85x0.8) points; when the first parameter is higher and meets the above constraints, the second parameter can be the same as the first parameter. Furthermore, the second parameter can be slightly larger than the first parameter (for example, the constraint weight can be 1.2).
[0157] In one possible implementation, the parameters in the second strategy generation model can be reversely adjusted using the second parameter. When the second parameter is maximum, the first strategy generation model can be determined based on the parameters in the second strategy generation model and the structure of the second strategy generation model.
[0158] In a possible implementation, after the first strategy generation model is determined, the first strategy generation model may be compressed using a distillation method or the like, thereby saving resources used to store the first strategy generation model.
[0159] In the embodiment of the present application, the strategy optimization direction is guided by a multi-objective function, combined with constraint processing and distillation methods, which not only ensures the security of the strategy, but also improves the deployment efficiency and model lightweight level of the first strategy generation model, achieving a balance between performance and energy efficiency.
[0160] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0161] Figure 6 This is a schematic diagram of the structure of the server load prediction device provided in the embodiment of the present application. Figure 6 As shown, the embodiment of the present application further provides a server load prediction device 60, comprising: an acquisition module 61, a first determination module 62, a second determination module 63, and a prediction module 64, wherein:
[0162] The acquisition module 61 is used to acquire the physical layer data, virtual layer data and application layer data of the server cluster.
[0163] The first determination module 62 is used to determine the first relationship between the servers in the server cluster based on the physical layer data, the virtual layer data and the application layer data. The first relationship includes the nodes corresponding to the first characteristics of each type of data and the edges between the nodes. The edges are used to indicate the association relationship between the data corresponding to the nodes.
[0164] The second determining module 63 is configured to determine a second feature corresponding to the server cluster according to features of the data corresponding to each node in the first relationship, where the second feature includes a spatial relationship and a temporal relationship between the data corresponding to each node.
[0165] The prediction module 64 is configured to predict the load of the server cluster in a future period based on the second feature.
[0166] A server load prediction device provided in an embodiment of the present application can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar and will not be repeated here.
[0167] In a possible implementation, the second determining module 63 is further configured to:
[0168] Determine a third feature based on the first feature corresponding to each node, where the third feature includes spatial and temporal fusion information between the data corresponding to each node in the first relationship;
[0169] Based on the third feature, the second feature is determined.
[0170] In a possible implementation, the second determining module 63 is further configured to:
[0171] Extracting the time feature of the first feature corresponding to each node to obtain the fourth feature corresponding to each node;
[0172] For any node, the first feature corresponding to the node and the fourth feature corresponding to the node are fused to obtain the fused feature corresponding to the node;
[0173] The third feature is determined according to the fusion features corresponding to each node.
[0174] In a possible implementation, the second determining module 63 is further configured to:
[0175] Determining a plurality of first weights between the nodes, the first weights being used to indicate the importance of the nodes to the load forecast at a plurality of time steps;
[0176] The second feature is determined according to the third feature corresponding to each node and the plurality of first weights.
[0177] In a possible implementation, the second determining module 63 is further configured to:
[0178] Performing weighted processing on the third features corresponding to each node according to the multiple first weights to obtain a weighted feature corresponding to each node;
[0179] Determine a first matrix based on the weighted features corresponding to each node, where the first matrix is used to indicate the degree of dependence between the nodes;
[0180] The time series features of the first matrix are extracted to obtain the second features.
[0181] In a possible implementation, the first determining module 62 is further configured to:
[0182] Determine a relationship graph based on the physical layer data, virtual layer data, and application layer data. The relationship graph includes nodes corresponding to each type of data and edges between the nodes.
[0183] According to the relationship diagram, a first relationship is determined.
[0184] In a possible implementation, the first determining module 62 is further configured to:
[0185] According to the data transmission volume between various types of data, the initial edges between nodes are obtained;
[0186] Based on the dependency relationship between various types of data, the initial edges are processed to obtain the edges between nodes;
[0187] The relationship graph is determined based on the edges between the nodes and the nodes corresponding to the characteristics of each type of data.
[0188] In a possible implementation, the acquisition module 61 is configured to:
[0189] Obtain physical layer data, virtual layer data, and application layer data of the server cluster, including:
[0190] Collect physical layer data, virtual layer data and application layer data;
[0191] Time-aligning the physical layer data, the virtual layer data, and the application layer data based on the timestamp in the physical layer data, the timestamp in the virtual layer data, and the timestamp in the application layer data;
[0192] The aligned data is standardized to obtain the physical layer data, virtual layer data, and application layer data of the server cluster.
[0193] In a possible implementation, the server load prediction device 60 further includes an execution module 65, which is configured to:
[0194] Determining, based on the first policy generation model, a plurality of control instructions corresponding to the load, the plurality of control instructions being used to adjust resource configuration of the server cluster;
[0195] Sorting multiple control instructions to obtain the execution order of the multiple control instructions;
[0196] Execute multiple control instructions according to the execution order.
[0197] In a possible implementation, the server load prediction device 60 further includes a processing module 66, wherein the processing module 66 is configured to:
[0198] Based on the sorted control instructions, the configuration parameters are determined, including the number of instances, resource quota threshold, and resource reservation ratio.
[0199] Process the resource configuration of the server cluster according to the configuration parameters.
[0200] In a possible implementation, the server load prediction device 60 further includes a third determination module 67, wherein the third determination module 67 is configured to:
[0201] Generate a model according to the second strategy, process the load, and predict a plurality of predicted control instructions corresponding to the load;
[0202] A first strategy generation model is determined based on the plurality of predicted control instructions.
[0203] In a possible implementation, the third determining module 67 is further configured to:
[0204] Determining first parameters corresponding to the plurality of predicted control instructions according to the plurality of predicted control instructions and the multi-objective function;
[0205] performing constraint processing on the plurality of predicted control instructions and determining a second parameter based on the first parameter;
[0206] When the second parameter is maximized, a first strategy generation model is determined based on the plurality of predicted control instructions.
[0207] A server load prediction device provided in an embodiment of the present application can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar and will not be repeated here.
[0208] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes: at least one processor 71 and a memory 72. Optionally, the electronic device 70 also includes a communication component 73. The processor 71, the memory 72 and the communication component 73 are connected via a bus.
[0209] During the specific implementation process, at least one processor 71 executes the computer-executable instructions stored in the memory 72, so that the at least one processor 71 executes the above-mentioned server load prediction method embodiment.
[0210] The specific implementation process of the processor 71 can be found in the above-mentioned method embodiment. Its implementation principle and technical effects are similar, and will not be repeated here in this embodiment.
[0211] In the above embodiments, it should be understood that the processor may be a CPU, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0212] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0213] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses shown in the drawings of this application are not limited to just one bus or just one type of bus.
[0214] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned server load prediction method embodiments when running.
[0215] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0216] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server load prediction method embodiments are implemented.
[0217] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps of any of the above-mentioned server load prediction method embodiments.
[0218] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0219] The above is a detailed introduction to a method for predicting the load of a physical server and an electronic device provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A server load prediction method, characterized in that: include: Obtain physical layer data, virtual layer data, and application layer data of the server cluster; determining, based on the physical layer data, the virtual layer data, and the application layer data, a first relationship between the servers in the server cluster, the first relationship comprising nodes corresponding to first features of the various types of data, and edges between the nodes, the edges being used to indicate associations between the data corresponding to the nodes; determining a second feature corresponding to the server cluster based on features of the data corresponding to each node in the first relationship, the second feature including a spatial relationship and a temporal relationship between the data corresponding to each node; Based on the second feature, the load of the server cluster in a future period is predicted.
2. The method according to claim 1, characterized in that The determining, based on the characteristics of the data corresponding to each node in the first relationship, the second characteristic corresponding to the server cluster includes: determining a third feature according to the first features corresponding to the nodes, the third feature comprising spatial and temporal fusion information between the data corresponding to the nodes in the first relationship; The second feature is determined based on the third feature.
3. The method according to claim 2, characterized in that The determining the third feature according to the first feature corresponding to each node includes: Extracting a time feature from the first feature corresponding to each node to obtain a fourth feature corresponding to each node; For any node, the first feature corresponding to the node is fused with the fourth feature corresponding to the node to obtain a fused feature corresponding to the node; The third feature is determined according to the fusion features corresponding to each node.
4. The method according to claim 2, characterized in that The determining the second feature according to the third feature includes: Determining a plurality of first weights between the nodes, the first weights being used to indicate the importance of the nodes at a plurality of time steps to the load forecast; The second feature is determined according to the third feature corresponding to each node and a plurality of first weights.
5. The method according to claim 4, characterized in that The determining the second feature according to the third feature corresponding to each node and the plurality of first weights includes: Performing weighted processing on the third features corresponding to the nodes according to the plurality of first weights to obtain weighted features corresponding to the nodes; Determining a first matrix according to the weighted features corresponding to the nodes, where the first matrix is used to indicate the degree of dependence between the nodes; Extracting time series features from the first matrix to obtain the second features.
6. The method according to claim 1, characterized in that Determining a first relationship between servers in the server cluster according to the physical layer data, the virtual layer data, and the application layer data includes: Determining a relationship graph based on the physical layer data, the virtual layer data, and the application layer data, the relationship graph including nodes corresponding to various types of data and edges between the nodes; The first relationship is determined according to the relationship graph.
7. The method according to claim 6, characterized in that The determining of a relationship graph according to the physical layer data, the virtual layer data, and the application layer data includes: According to the data transmission volume between various types of data, the initial edges between nodes are obtained; Based on the dependency relationship between various types of data, the initial edges are processed to obtain the edges between nodes; The relationship graph is determined based on the edges between the nodes and the nodes corresponding to the characteristics of each type of data.
8. The method according to claim 1, characterized in that The obtaining of the physical layer data, virtual layer data and application layer data of the server cluster includes: Collecting the physical layer data, the virtual layer data, and the application layer data; Time-aligning the physical layer data, the virtual layer data, and the application layer data based on a timestamp in the physical layer data, a timestamp in the virtual layer data, and a timestamp in the application layer data; The aligned data is standardized to obtain the physical layer data, virtual layer data, and application layer data of the server cluster.
9. The method according to claim 1, characterized in that After predicting the load of the server cluster in a future period according to the second feature, the method further includes: Determining, based on a first policy generation model, a plurality of control instructions corresponding to the load, wherein the plurality of control instructions are used to adjust resource configuration of the server cluster; Sorting the multiple control instructions to obtain an execution order of the multiple control instructions; The plurality of control instructions are executed according to the execution order.
10. The method according to claim 9, characterized in that After executing the multiple control instructions, the method further includes: Determining configuration parameters based on the sorted multiple control instructions, the configuration parameters including the number of instances, resource quota thresholds, and resource reservation ratios; The resource configuration of the server cluster is processed according to the configuration parameters.
11. The method according to claim 9, characterized in that Before determining the plurality of control instructions corresponding to the load based on the first strategy generation model, the method further includes: Generate a model according to a second strategy, process the load, and predict a plurality of predicted control instructions corresponding to the load; A first strategy generation model is determined based on the plurality of predicted control instructions.
12. The method according to claim 11, characterized in that The determining a first strategy generation model based on the plurality of predicted control instructions includes: determining, according to the plurality of predicted control instructions and a multi-objective function, first parameters corresponding to the plurality of predicted control instructions; performing constraint processing on the plurality of predicted control instructions, and determining a second parameter based on the first parameter; When the second parameter is maximum, the first strategy generation model is determined based on the multiple predicted control instructions.
13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server load prediction method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server load prediction method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the server load prediction method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Service node performance prediction method, device, equipment, readable medium and product
CN119718860A
Load balancing method and device, electronic equipment, storage medium and program product
CN120276873A
Method for deploying an application workload on a cluster
US20200404076A1