Resource allocation method, storage medium, electronic equipment and program product
By preprocessing the operating data of server nodes and analyzing the resource prediction model, dynamic resource allocation strategies are generated, and the problems of inaccurate and static resource allocation in the existing technology are solved, and efficient resource utilization and system performance are achieved.
Patent Information
- Application Number
- CN202510519963.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art only uses a simple statistical model to predict the resource of server nodes, resulting in inaccurate configuration and static resource configuration and cannot be flexibly adjusted according to actual needs, resulting in resource bottlenecks or waste.
The operation data of multiple server nodes in the server system is collected, data preprocessed, and input it into the resource prediction model. The resource usage requirements of multiple server nodes within a predetermined time range are predicted, the dependencies between time nodes are identified, the resource configuration strategy is generated, and the resource configuration is dynamically adjusted.
Through accurate resource demand forecasts and dynamic adjustments, avoid insufficient resources or waste, optimize resource configuration, improve system response speed and user satisfaction, reduce operation costs, and improve overall performance and efficiency of the server system.
Smart Images

Figure CN120066798A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a resource allocation method, a storage medium, an electronic device, and a program product. Background Art
[0002] Server management technology refers to a series of methods and technologies used to monitor, maintain, and optimize server performance. It covers all aspects from hardware status monitoring to software resource allocation, aiming to ensure that the server can operate efficiently and stably while maximizing resource utilization and reducing operating costs.
[0003] Currently, related technologies only perform resource prediction on server nodes through a simple statistical model, resulting in inaccurate configuration based on the resource prediction results. Moreover, the resource allocation is static and cannot be flexibly adjusted according to actual needs, which may also lead to resource bottlenecks or resource waste. Summary of the Invention
[0004] The present disclosure provides a resource allocation method, a storage medium, an electronic device, and a program product. Its main purpose is to solve the problem that related technologies only perform resource prediction on server nodes through a simple statistical model, resulting in inaccurate configuration based on the resource prediction results, and the resource allocation is static and cannot be flexibly adjusted according to actual needs, which may also lead to resource bottlenecks or resource waste.
[0005] In a first aspect, the present application provides a resource allocation method, including: Collecting operation data of multiple server nodes in a server system to form an operation data set; Performing data preprocessing on the operation data set to obtain a to-be-predicted operation data set; Inputting the to-be-predicted operation data set into a resource prediction model to predict the resource usage requirements of multiple server nodes within a predetermined time range, obtaining the resource usage requirements corresponding to multiple time nodes within the predetermined time range respectively, where the resource prediction model is used to identify the dependency relationships between time nodes and perform resource prediction; Generating resource allocation strategies corresponding to multiple server nodes according to the prediction results, and dynamically adjusting the configured resources of multiple server nodes based on the resource allocation strategies.
[0006] In a second aspect, the present application provides a resource allocation device, including: A collection module configured to collect operation data of multiple server nodes in a server system to form an operation data set; A processing module configured to perform data preprocessing on the operation data set to obtain a to-be-predicted operation data set; A prediction module, configured to input a to-be-predicted operation dataset into a resource prediction model to predict the resource usage requirements of multiple server nodes within a predetermined time range, and obtain the resource usage requirements corresponding to multiple time nodes within the predetermined time range respectively, where the resource prediction model is used to identify the dependency relationships between multiple time nodes and perform resource prediction; An adjustment module, configured to generate resource configuration policies corresponding to multiple server nodes according to the prediction results, and dynamically adjust the configured resources of the multiple server nodes based on the resource configuration policies.
[0007] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0008] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the method of the first aspect is implemented.
[0009] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0010] The resource allocation method, storage medium, electronic device, and program product provided by the present disclosure, wherein the method includes: collecting the operation data of multiple server nodes in the server system to form an operation data set; performing data preprocessing on the operation data set to obtain a to-be-predicted operation data set; inputting the to-be-predicted operation data set into a resource prediction model to predict the resource usage requirements of multiple server nodes within a predetermined time range, and obtaining the resource usage requirements corresponding to multiple time nodes within the predetermined time range respectively, where the resource prediction model is used to identify the dependency relationships between multiple time nodes and perform resource prediction; generating resource allocation policies corresponding to multiple server nodes according to the prediction results, and dynamically adjusting the allocated resources of multiple server nodes based on the resource allocation policies. Compared with the related technologies, the present application can collect the operation data of multiple server nodes in the server system to form an operation data set, and perform preprocessing to obtain a to-be-predicted operation data set. By inputting the to-be-predicted data into the resource prediction model, the resource usage requirements of each server node are predicted, and the resource usage requirements of multiple time nodes within a predetermined time period are obtained. Among them, the resource prediction model is used to identify the dependency relationships between time nodes. By using the resource prediction model, the long-term dependency relationships between multiple time nodes can be effectively captured, making the resource demand prediction more accurate. Furthermore, the system can make preparations for resource allocation in advance, avoid the situation of resource shortage or waste, optimize the resource allocation, improve the system response speed and user satisfaction, and at the same time reduce the operation cost; the resource allocation policy generated based on the prediction results dynamically adjusts the resource allocation, which can make the present application flexibly allocate resources according to the actual needs, ensure the balanced resource allocation of each server node, not only improve the resource utilization rate, reduce unnecessary energy consumption, but also enhance the overall performance and efficiency of the server system, and at the same time enhance the scalability and flexibility of the system.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to illustrate the embodiments of the present application more clearly, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 FIG. shows a schematic flowchart of a resource allocation method provided by an embodiment of the present application; Figure 2 FIG. shows a schematic flowchart of another resource allocation method provided by an embodiment of the present application; Figure 3 Shows a schematic diagram of an example provided by an embodiment of the present application; Figure 4 Shows a schematic diagram of an example provided by an embodiment of the present application; Figure 5 Shows a schematic structural diagram of a resource configuration device provided by an embodiment of the present application. Detailed implementation manners
[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.
[0016] Server management technology refers to a series of methods and technologies for monitoring, maintaining, and optimizing server performance. It covers all aspects from hardware status monitoring to software resource allocation, aiming to ensure that the server can operate efficiently and stably while maximizing resource utilization and reducing operating costs.
[0017] In the field of server management, traditional server management systems only monitor some performance metrics, resulting in incomplete data; the data collection methods for different server nodes are not unified, affecting the accuracy of data analysis, and existing resource prediction methods rely on simple statistical models and are difficult to capture complex long-term dependencies, resulting in inaccurate resource demand predictions. At the same time, traditional resource configuration methods are static and cannot be flexibly adjusted according to actual needs, resulting in resource bottlenecks or waste during certain time periods.
[0018] To improve the technical problems in the related art that only simple statistical models are used to predict resources for server nodes, resulting in inaccurate configuration situations based on the resource prediction results, and the resource configuration is static and cannot be flexibly adjusted according to actual needs, which will also lead to resource bottlenecks or resource waste. This embodiment provides a resource configuration method, as Figure 1 shown, this method includes the following steps: Step 101: Collect the operation data of multiple server nodes in the server system to form an operation data set.
[0019] In the embodiments of the present application, the server system including multiple server nodes can be a distributed computing environment or a cluster system. Such an architecture is designed to provide high availability, fault tolerance, and scalability. Each server node can be a physical machine or a virtual machine, and they cooperate with each other to provide services or process tasks.
[0020] In some examples, the operation data of the server node can specifically include but is not limited to the utilization rate of the central processing unit (CPU), the memory usage rate, the storage input / output (I / O), and the network traffic, etc.
[0021] For the embodiments of the present application, the operation data set contains the operation data corresponding to each server node at consecutive different time nodes.
[0022] Step 102: Perform data preprocessing on the operation data set to obtain a to-be-predicted operation data set.
[0023] In the embodiments of the present application, the to-be-predicted operation data set can specifically be a data set that can be used for resource usage demand prediction after performing data preprocessing on the operation data set.
[0024] In some examples, data preprocessing is a key step in data analysis and machine learning projects. It involves a series of operations on the original data to prepare for subsequent analysis or model training. Good data preprocessing can significantly improve the model performance, accelerate the training process, and help improve the prediction accuracy. Specifically, the process of data preprocessing can include but is not limited to data cleaning, data integration, data transformation, feature selection, data dimensionality reduction, etc.
[0025] Step 103: Input the to-be-predicted operation data set into the resource prediction model to predict the resource usage demands of multiple server nodes within a predetermined time range, and obtain the resource usage demands corresponding to multiple time nodes within the predetermined time range respectively.
[0026] Among them, the resource prediction model is used to identify the dependency relationships between multiple time nodes and perform resource prediction.
[0027] In the embodiments of the present application, the resource prediction model can specifically be a model trained based on a long short-term memory network for identifying the dependency relationships between time nodes and performing resource prediction.
[0028] In some examples, Long Short-Term Memory Networks (LSTM) is a special type of Recurrent Neural Network (RNN) specifically designed to learn dependencies over long time intervals. Due to its unique structure, LSTM can effectively address the vanishing gradient or exploding gradient problems encountered by traditional RNNs when dealing with long-term dependencies. Therefore, the model trained based on the long short-term memory network in the embodiments of this application can identify the dependencies between multiple time nodes, making the prediction results of the embodiments of this application more accurate.
[0029] For this embodiment, the predetermined time range can be a time range based on a period after the current time. Correspondingly, the multiple time nodes can be the time nodes within the preset time range.
[0030] Step 104: Generate resource configuration policies corresponding to multiple server nodes according to the prediction results, and dynamically adjust the configured resources of the multiple server nodes based on the resource configuration policies.
[0031] In some examples, adjusting the resource configuration according to the prediction results can be that if the prediction shows that the load will increase significantly in a certain future period, more server nodes can be started in advance or the resource quotas of existing nodes can be increased. Conversely, if the load is expected to decrease, the number of active nodes can be appropriately reduced to save costs.
[0032] Compared with the related technologies, in this embodiment, the running data of multiple server nodes in the server system can be collected to form a running data set, and preprocessed to obtain a running data set to be predicted. By inputting the data to be predicted into the resource prediction model, the resource usage requirements of each server node can be predicted, and the resource usage requirements at multiple time nodes within a predetermined time period can be obtained. Among them, the resource prediction model is used to identify the dependencies between time nodes. By using the resource prediction model, the long-term dependencies between multiple time nodes can be effectively captured, making the resource demand prediction more accurate. Furthermore, the system can make preparations for resource configuration in advance, avoid the situation of resource shortage or waste, optimize the resource configuration, improve the system response speed and user satisfaction, and at the same time reduce the operation cost; the dynamic adjustment of the resource configuration based on the resource configuration policy generated according to the prediction results can enable this embodiment to flexibly allocate resources according to actual needs, ensure the balanced resource configuration of each server node, not only improve the resource utilization rate, reduce unnecessary energy consumption, but also enhance the overall performance and efficiency of the server system, and at the same time enhance the scalability and flexibility of the system.
[0033] Further, as a refinement and extension of the above embodiment, the following methods can be adopted but are not limited to, such as Figure 2 As shown, the method includes: Step 201: Based on the monitoring agent software configured for each server node among multiple server nodes, collect the operation data of multiple performance metrics corresponding to each server node at a predetermined time interval.
[0034] Among them, the multiple performance metrics at least include one or more of processor utilization rate, memory usage rate, storage input / output, and network traffic.
[0035] In the embodiments of the present application, configuring monitoring agent software for each server node is one of the key steps to ensure the healthy and efficient operation of the system. These agent software are responsible for collecting various performance metrics of the server and sending the data to the central monitoring system for analysis and display.
[0036] Exemplarily, in the embodiments of the present application, a uniformly configured monitoring agent software can be installed on each server node. Among them, the monitoring agent software supports the monitoring and data collection of CPU utilization rate, memory usage rate, storage input / output (I / OSS), and network traffic; the monitoring agent regularly obtains the CPU utilization rate, memory usage rate, storage I / O rate, and network traffic performance metrics data (i.e., the multiple performance metrics in the embodiments of the present application) from each server node according to the set time interval, and obtains an initial multi-dimensional operation data set (i.e., the operation data set in the embodiments of the present application).
[0037] In some examples, when monitoring the server performance, the CPU utilization rate, memory usage rate, storage I / O (input / output), and network traffic are four very key performance metrics. They respectively reflect the workload and resource consumption of different aspects of the server. Specifically, the CPU utilization rate refers to the percentage of time that the processor is used to execute non-idle tasks. It shows the system's ability to process tasks within a period of time; the memory usage rate shows the proportion of physical memory allocated to various running processes. It helps us understand the memory pressure situation of the system; the storage I / O includes information such as disk read and write speeds and frequencies, reflecting the load and efficiency of the storage subsystem. This directly affects the data read and write efficiency; the network traffic refers to the amount of data entering and leaving the server, which can be used to evaluate the quality of the network connection and the bandwidth usage.
[0038] Step 202: Based on the operation data corresponding to multiple server nodes respectively, form an operation data set.
[0039] It should be noted that by installing a uniformly configured monitoring agent software on each server node and regularly collecting multi-dimensional data from each node, not only the comprehensiveness and consistency of the data are ensured, but also a solid foundation is provided for subsequent data analysis and processing. The method can capture the changes in the server state in real time, help to discover potential problems in a timely manner, and thus improve the stability and reliability of the system.
[0040] Step 203: Perform data preprocessing on the running data set to obtain the running data set to be predicted.
[0041] Optionally, step 203 may specifically include: performing data cleaning on the running data set to obtain the cleaned running data set; determining, based on the cleaned running data set, the basic statistics corresponding to each of the multiple performance metrics, and the overall health index of each server node; and forming the running data set to be predicted based on the basic statistics and the overall health index.
[0042] Exemplarily, for the data at each time point, the 3σ principle can be applied to identify and remove outliers; the linear interpolation method can be used to fill in the missing data at some time points; based on the cleaned data, the basic statistics of each performance metric can be calculated; based on the cleaned data, the overall health index can be calculated through Formula 1 , which is used to accurately reflect the overall running status of the server at this moment. Formula 1 is specifically as follows: (Formula 1) where is the average value of the CPU utilization rate, is the average value of the memory usage rate, is the variance of the storage I / O rate, is the average value of the network traffic; according to the calculation results, key performance indicators of the server are generated.
[0043] It should be noted that by applying the 3σ principle to identify and remove outliers, using the linear interpolation method to fill in the missing data, and calculating the basic statistics and the overall health index of each performance metric, the quality and usability of the data can be effectively improved, ensuring the authenticity and accuracy of the data used in subsequent analysis, making the key indicators generated based on these data more reliable, and thus improving the accuracy and effectiveness of the system performance evaluation.
[0044] Step 204: Input the running data set to be predicted into the resource prediction model to predict the resource usage requirements of multiple server nodes within a predetermined time range, and obtain the resource usage requirements corresponding to multiple time nodes within the predetermined time range.
[0045] Among them, the resource prediction model is used to identify the dependency relationships between multiple time nodes and perform resource prediction.
[0046] Optionally, the training process of the resource prediction model may specifically include: collecting the historical operation data sets of multiple server nodes; using the historical basic statistics corresponding to each performance metric among the multiple performance metrics of the multiple server nodes in the historical operation data set as input features, and using the historical resource requirements of the multiple server nodes observed in the historical operation data set as output features to form a training data set; establishing the resource prediction model based on a long short-term memory network, training the resource prediction model using historical data, and adjusting the parameters of the resource prediction model until the target accuracy rate is achieved to obtain a trained resource prediction model.
[0047] Exemplarily, a long short-term memory network (LSTM) is selected to establish a risk prediction model. The key server performance metrics are used as input features. At the same time, the actually observed resource demand is used as the output label to form a training data set. The historical data, that is, the historical operation data in the embodiments of the present application, is used to train the LSTM model, and the model parameters are adjusted until the expected accuracy rate is achieved. During the training process, the mean square error (MSE) is used as the loss function, and the optimization goal is to minimize the difference between the predicted value and the true value. The cross-validation method is used to verify the trained LSTM model. Among them, the loss function is specifically shown in Formula 2 below: (Formula 2) In Formula 2, is the true resource demand of the th sample, is the predicted resource demand of the th sample, is the total number of samples.
[0048] As an optional method, the trained LSTM model is used to predict the resource demand in a future period of time. The latest multi-dimensional operation data set is input, and the predicted resource demand at each future time point is output based on Formula 3. Formula 3 is specifically shown as follows: (Formula 3) In Formula 3, is the predicted resource demand at time, , , , are respectively the average CPU utilization rate, the variance of memory usage rate, the peak storage I / O rate, and the improved overall health index at time.
[0049] It should be noted that by using the machine learning algorithm of Long Short-Term Memory Network (LSTM) to establish a risk prediction model, not only can the long-term dependencies in time series data be accurately captured, but also the future resource requirements can be effectively predicted. The accurate prediction ability enables the system to make preparations for resource allocation in advance, avoid resource shortages or waste, thereby optimizing resource allocation, improving the system's response speed and user satisfaction, and reducing operating costs at the same time.
[0050] Step 205: Generate resource configuration strategies corresponding to multiple server nodes according to the prediction results, and dynamically adjust the configured resources of the multiple server nodes based on the resource configuration strategies.
[0051] Optionally, before step 205, the method of this embodiment may include: determining the resource occupancy and comprehensive resource load index corresponding to each of the multiple server nodes based on the dataset of the operation to be predicted; correspondingly, step 205 may specifically include: generating resource configuration strategies corresponding to the multiple server nodes according to the resource occupancy and comprehensive resource load index, and dynamically adjusting the configured resources of the multiple server nodes based on the resource configuration strategies.
[0052] Exemplarily, according to the latest multi-dimensional operation dataset, calculate the resource occupancy of each current server node, and calculate the comprehensive resource load index of each current server node based on Formula 4. Formula 4 is specifically as follows: (Formula 4) In Formula 4, , , , are the corresponding weight coefficients respectively , , , are respectively the average value of CPU utilization rate, the variance of memory usage rate, the peak value of storage I / O rate, and the improved overall health index at time represents the comprehensive resource load index of the server node at time, and the larger its value, the more tense the resource occupancy of the node.
[0053] As an optional method, generating resource configuration strategies corresponding to multiple server nodes according to the resource occupancy and comprehensive resource load index may specifically be: based on compare with the current resource status to formulate specific resource adjustment strategies; when it is predicted that the CPU utilization rate of a certain server node will increase significantly in a certain period of time in the future, and other nodes have sufficient idle CPU resources, consider migrating some virtual machine instances to the nodes with more idle resources.
[0054] In some examples, the dynamic adjustment of the configured resources of multiple server nodes based on the resource configuration policy can specifically be: performing operations according to the formulated resource adjustment policy, where the operations include migrating virtual machine instances and reallocating computing tasks; after the resource adjustment is completed, collecting new operation data and calculating the comprehensive resource load index again to verify whether the adjustment effect reaches the expected goal.
[0055] It should be noted that by calculating the comprehensive resource load index and formulating specific resource adjustment policies according to the prediction results, the method can flexibly respond to the changes in resource requirements in different time periods. This method can not only improve resource utilization rate, reduce unnecessary energy consumption, but also ensure the efficient operation of the system during peak periods, release resources for other tasks during off-peak periods, achieving the goal of improving the overall performance and efficiency of the system, and at the same time enhancing the scalability and flexibility of the system.
[0056] Optionally, before performing the "dynamic adjustment of the configured resources of multiple server nodes based on the resource configuration policy", the method of this embodiment further includes: collecting the historical workload data sets of multiple server nodes; inputting the historical workload data sets into the load prediction model to predict the workload intensity of multiple server nodes within a predetermined time range, obtaining the workload intensity indexes corresponding to multiple time nodes within the predetermined time range respectively; determining the workload statuses corresponding to multiple server nodes respectively based on the workload intensity indexes.
[0057] Among them, the load prediction model is used to identify the dependency relationships between multiple time nodes and perform load prediction.
[0058] In the embodiments of the present application, the workload intensity of a server node refers to the pressure level that the server bears when processing various tasks, which is usually evaluated by monitoring multiple key performance indicators (KPIs).
[0059] In some examples, the load prediction model is a model trained based on a long short-term memory network for identifying the dependency relationships between time nodes and performing load prediction.
[0060] Exemplarily, extracting key workload features from historical load data, the key workload features include the average value, standard deviation, and peak value, defining a comprehensive workload index to quantify the workload intensity at each time point, and the expression is: (Formula Five) Among them, is the average value of the CPU utilization rate at time is the variance of the memory usage rate at time is Store the peak value of the I / O rate at a moment, for the network traffic at a moment, which represents the workload intensity at a moment.
[0061] For this embodiment, based on historical data and the current workload trend, use the machine learning model LSTM to predict the workload pattern in a future period of time. The load prediction model outputs the predicted workload index at each future time point , and judge the future workload status according to the set threshold.
[0062] Optionally, when performing "determine the workload status corresponding to each of the multiple server nodes based on the workload intensity index", the method of this embodiment may include: comparing the workload intensity indexes corresponding to each of the multiple server nodes with a first workload intensity threshold and a second workload intensity threshold respectively, where the first workload intensity threshold is less than the second workload intensity threshold; when it is determined that the workload intensity index of the target server node among the multiple server nodes is less than the first workload intensity threshold, determine the workload status of the target server node as a flat state; when it is determined that the workload intensity index of the target server node among the multiple server nodes is greater than or equal to the first workload intensity threshold and less than or equal to the second workload intensity threshold, determine the workload status of the target server node as a normal state; when it is determined that the workload intensity index of the target server node among the multiple server nodes is greater than the second workload intensity threshold, determine the workload status of the target server node as a peak state.
[0063] Optionally, the process of determining the first workload intensity threshold and the second workload intensity threshold includes: performing clustering processing on the historical workload data set to cluster out a first data set corresponding to the flat workload state and a second data set corresponding to the peak workload state in the historical workload data set; determining the first workload intensity threshold based on the first data set, and determining the second workload intensity threshold based on the second data set.
[0064] In some examples, the clustering algorithm K-means can be used to classify the historical workload data to identify typical peak periods and flat periods; set thresholds and , which represent the determination criteria for peak periods and flat periods respectively; when the comprehensive workload index in a certain time period exceeds , then this time period is marked as a peak period; if it is lower than , then it is marked as a flat period.
[0065] Optionally, when performing "dynamically adjusting the configured resources of multiple server nodes based on a resource configuration policy", the method of this embodiment specifically further includes: adjusting the resource configuration policy based on the workload statuses respectively corresponding to the multiple server nodes; and dynamically adjusting the configured resources of the multiple server nodes according to the adjusted resource configuration policy.
[0066] For this embodiment, according to the prediction results, a specific resource adjustment policy is formulated. For the expected peak periods, the resource configurations of relevant server nodes are increased in advance to cope with the upcoming high-load demands; for the flat periods, the resource configurations can be appropriately reduced to release some resources for other tasks or reduce energy consumption; the resource adjustment policy is executed, including migrating virtual machine instances, reallocating computing tasks, or adjusting the scale of the server cluster; the entire load balancing process is recorded to generate a detailed load balancing policy report.
[0067] It should be noted that by defining a comprehensive workload index and using a clustering algorithm to identify peak periods and flat periods, the future workload patterns can be predicted more accurately. The method not only helps to adjust server resources in advance to cope with the upcoming high-load demands, but also can reasonably release resources during the low-peak periods to reduce energy consumption. In this way, the system can always maintain an efficient load balancing state, ensure that user requests can be quickly responded to at any time, and maximize the resource utilization efficiency.
[0068] Optionally, after performing "generating the resource configuration policies corresponding to the multiple server nodes according to the prediction results and dynamically adjusting the configured resources of the multiple server nodes based on the resource configuration policies", the method of this embodiment further includes: collecting the hardware status information of the multiple server nodes, and monitoring the hardware status scores respectively corresponding to the multiple servers based on the hardware status information; and generating an alarm message for a target server node when it is monitored that the hardware status score of the target server node among the multiple server nodes is greater than a predetermined hardware status threshold.
[0069] Wherein, the target server node is any one of the multiple server nodes.
[0070] In the embodiment of the present application, according to historical data and industry standards, a health index threshold is set for each type of hardware status information; for each index, a warning threshold and a danger threshold are set; based on the collected hardware status information, the comprehensive health score H health (t) of each server node is calculated through Formula 6, and Formula 6 is specifically as follows: (Formula 6) In Formula 6, is the number of different hardware status indicators, is the The weight coefficient of a hardware status indicator is the th hardware status indicator at time, and the standardized score H health at time (t) represents the comprehensive health score of the server node at time t.
[0071] Exemplarily, the comprehensive health score H health of each server node is monitored in real time. Once the score of a certain node exceeds the preset warning threshold, a warning notification is immediately triggered; when a potential failure risk is detected, based on the current multi-dimensional operation data set, the cause is analyzed; a specific resource migration plan is formulated, and the critical tasks on the affected node are preferentially migrated to other healthy server nodes; according to the formulated resource migration plan, the migration operation is executed.
[0072] It should be noted that by setting the health indicator threshold and monitoring the server hardware status in real time, potential failure risks can be detected early, and the warning mechanism can be automatically triggered. This method not only improves the system's fault prevention ability but also can quickly migrate the affected tasks to other healthy server nodes without affecting the business, ensuring the high availability and stability of the system. In addition, it also provides valuable data support for subsequent maintenance and upgrade.
[0073] Optionally, after executing the resource configuration policies corresponding to multiple server nodes according to the prediction results and dynamically adjusting the configured resources of the multiple server nodes based on the resource configuration policies, the method of this embodiment further includes: collecting the adjusted operation data of the multiple server nodes, and determining the adjusted comprehensive resource load indexes of the multiple server nodes based on the adjusted operation data; verifying the adjustment conditions of the multiple server nodes based on the adjusted comprehensive resource load indexes.
[0074] For this embodiment, after the resource adjustment is completed, new operation data is collected, and the comprehensive resource load index is calculated again to verify whether the adjustment effect reaches the expected goal; It should be noted that by calculating the comprehensive resource load index and formulating specific resource adjustment strategies according to the prediction results, the resource requirements changes in different time periods can be flexibly responded to. This method can not only improve the resource utilization rate, reduce unnecessary energy consumption, but also ensure the efficient operation of the system during peak periods and release resources for other tasks during off-peak periods, achieving the goal of improving the overall performance and efficiency of the system, and at the same time enhancing the scalability and flexibility of the system.
[0075] In some examples, based on the resource configuration method in the above embodiment, such as Figure 3As shown in the figure, an embodiment of the present application can also provide a resource allocation system, including: a data collection module, a feature extraction module, a risk prediction model module, a dynamic resource allocation module, an early warning module, and a load balancing module.
[0076] The data collection and initialization module is used to install monitoring agent software on each server node and regularly collect CPU utilization rate, memory usage rate, storage I / O rate, and network traffic from each node to obtain an initial multi-dimensional operation data set.
[0077] The feature extraction module is used to remove outliers and fill in missing data, calculate the basic statistics of each performance metric and the overall health index to obtain the key indicators of server performance.
[0078] The risk prediction model module is used to select the LSTM algorithm to establish a risk prediction model, input historical data for training and predict future resource requirements to obtain a prediction result.
[0079] The dynamic resource allocation module is used to adjust the resource allocation among servers according to the prediction result, calculate the comprehensive resource load index, formulate and execute a resource adjustment strategy to obtain a resource allocation plan; The early warning module is used to monitor the hardware status of the server in real time, calculate the comprehensive health score, automatically trigger an early warning mechanism when detecting potential failure risks, and formulate a resource migration plan to obtain a resource migration strategy.
[0080] The load balancing module is used to identify peak hours and off-peak periods, define the comprehensive workload index, dynamically adjust server resources, ensure that the system can efficiently respond to user requests while avoiding resource waste, and obtain a load balancing strategy.
[0081] Compared with the related technologies, in this embodiment, the running data of multiple server nodes in the server system can be collected to form a running data set, and preprocessed to obtain a to-be-predicted running data set. By inputting the to-be-predicted data into the resource prediction model, the resource usage requirements of each server node can be predicted, and the resource usage requirements at multiple time nodes within a predetermined time period can be obtained. Among them, the resource prediction model is used to identify the dependency relationships between time nodes. By using the resource prediction model, the long-term dependency relationships between multiple time nodes can be effectively captured, making the resource demand prediction more accurate. Furthermore, the system can make preparations for resource allocation in advance, avoid the situation of resource shortage or waste, optimize the resource allocation, improve the system response speed and user satisfaction, and at the same time reduce the operation cost; the resource allocation strategy generated based on the prediction results can dynamically adjust the resource allocation, enabling this embodiment to flexibly allocate resources according to actual needs, ensuring the balanced resource allocation of each server node, not only improving the resource utilization rate, reducing unnecessary energy consumption, but also enhancing the overall performance and efficiency of the server system, and at the same time enhancing the scalability and flexibility of the system.
[0082] To illustrate the specific implementation process of this embodiment, the following specific application examples are given, such as Figure 4 shown, but not limited thereto: The first step: Initialize each server node and collect multi-dimensional data of CPU utilization rate, memory usage rate, storage I / O, and network traffic to obtain an initial multi-dimensional running data set; specifically include: Install uniformly configured monitoring agent software on each server node; The monitoring agent software supports the monitoring and data collection of CPU utilization rate, memory usage rate, storage I / OSS, and network traffic; the monitoring agent regularly obtains the CPU utilization rate, memory usage rate, storage I / O rate, and network traffic performance index data from each server node according to the set time interval, and obtains an initial multi-dimensional running data set.
[0083] The second step: Use data cleaning and feature extraction methods to preprocess the multi-dimensional running data set to obtain the key indicators of server performance; specifically include: For the data at each time point, apply the 3σ principle to identify and remove outliers; use linear interpolation to fill in the missing data at some time points; calculate the basic statistics of each performance index based on the cleaned data; calculate the overall health index to accurately reflect the overall running status of the server at this moment; generate the key indicators of server performance according to the calculation results.
[0084] Step 3: Establish a risk prediction model based on machine learning algorithms, input historical data into the risk prediction model to predict resource requirements for a period of time in the future, and obtain prediction results. Specifically, it includes: Selecting a Long Short-Term Memory (LSTM) network to establish a risk prediction model; Using the key performance indicators of server performance as input features; At the same time, using the actually observed resource demand as the output label to form a training dataset; Training the LSTM model with historical data and adjusting the model parameters until the expected accuracy is achieved; Using the Mean Squared Error (MSE) as the loss function during the training process, and the optimization goal is to minimize the difference between the predicted value and the true value; Using the cross-validation method to verify the trained LSTM model; Using the trained LSTM model to predict the resource requirements for a period of time in the future, input the latest multi-dimensional operation dataset, and output the predicted resource requirements for each time point in the future.
[0085] Step 4: Adopt a method of dynamically adjusting resource allocation, and adjust the resource allocation among servers according to the prediction results to obtain a resource allocation plan. Specifically, it includes: Calculating the resource occupancy of each server node based on the latest multi-dimensional operation dataset; Calculating the comprehensive resource load index of each server node ; Based on Compare with the current resource status to formulate specific resource adjustment strategies; When it is predicted that the CPU utilization rate of a certain server node will increase significantly in a certain period of time in the future, and other nodes have sufficient idle CPU resources, consider migrating some virtual machine instances to the nodes with more idle resources; According to the formulated resource adjustment strategies, perform operations; The operations include migrating virtual machine instances and reallocating computing tasks; After the resource adjustment is completed, collect new operation data and calculate the comprehensive resource load index again to verify whether the adjustment effect meets the expected goal.
[0086] Step 5: Adopt a method of continuously monitoring the hardware status of servers for health monitoring. When a potential failure risk is detected, automatically trigger an early warning mechanism and formulate a resource migration plan to obtain a resource migration strategy. Specifically, it includes: Setting health index thresholds for each type of hardware status information according to historical data and industry standards; For each indicator, setting warning thresholds and danger thresholds; Calculating the comprehensive health score of each server node based on the collected hardware status information ; Real-time monitoring of the comprehensive health score of each server node , once it is found that the score of a certain node exceeds the preset warning threshold, immediately trigger a warning notification; When a potential failure risk is detected, based on the current multi-dimensional operation dataset, analyze the reasons; Formulate a specific resource migration plan, and give priority to migrating key tasks on the affected nodes to other server nodes with good health conditions; According to the formulated resource migration plan, perform the migration operation.
[0087] Step 6: Identify peak periods and off-peak periods based on the workload pattern, and dynamically adjust server resources to obtain a load balancing strategy; specifically including: extracting key workload characteristics from historical data; the key workload characteristics include the average value, standard deviation, and peak value; defining a comprehensive workload index to quantify the workload intensity at each time point; using the K-means clustering algorithm to classify historical workload data to identify typical peak periods and off-peak periods; setting thresholds and , which represent the judgment criteria for peak periods and off-peak periods respectively; when the comprehensive workload index exceeds , then this time period is marked as a peak period; if it is lower than , then it is marked as an off-peak period; based on historical data and the current workload trend, use the machine learning model LSTM to predict the workload pattern in the future for a period of time; the prediction model outputs the predicted workload index for each future time point, and judge the future workload status according to the set thresholds; according to the prediction results, formulate specific resource adjustment strategies; for the expected peak periods, increase the resource configuration of relevant server nodes in advance to cope with the upcoming high-load demand; for off-peak periods, the resource configuration can be appropriately reduced to release some resources for other tasks or reduce energy consumption; execute the resource adjustment strategy, including migrating virtual machine instances, reallocating computing tasks, or adjusting the scale of the server cluster; record the entire load balancing process to generate a detailed load balancing strategy report.
[0088] Step 7: Adopt an effectiveness evaluation method to regularly check the running effect of the system, compare the difference between the actual performance and the prediction target, and update the risk prediction model.
[0089] Compared with the related technologies, in this embodiment, the operation data of multiple server nodes in the server system can be collected to form an operation data set, and preprocessed to obtain a to-be-predicted operation data set. By inputting the to-be-predicted data into the resource prediction model, the resource usage requirements of each server node are predicted, and the resource usage requirements at multiple time nodes within a predetermined time period are obtained. Among them, the resource prediction model is used to identify the dependency relationships between time nodes. By using the resource prediction model, the long-term dependency relationships between multiple time nodes can be effectively captured, making the resource demand prediction more accurate. Furthermore, the system can prepare for resource allocation in advance, avoid resource shortages or waste, optimize resource allocation, improve the system response speed and user satisfaction, and at the same time reduce the operation cost. The resource allocation strategy generated based on the prediction results can dynamically adjust the resource allocation, enabling this embodiment to flexibly allocate resources according to actual needs, ensuring the balanced resource allocation of each server node, not only improving the resource utilization rate, reducing unnecessary energy consumption, but also enhancing the overall performance and efficiency of the server system, and at the same time enhancing the scalability and flexibility of the system.
[0090] An embodiment of the present application further provides a resource allocation device, as Figure 5 shown. The device includes: a collection module 31, a processing module 32, a prediction module 33, and an adjustment module 34.
[0091] The collection module 31 is configured to collect the operation data of multiple server nodes in the server system to form an operation data set; The processing module 32 is configured to perform data preprocessing on the operation data set to obtain a to-be-predicted operation data set; The prediction module 33 is configured to input the to-be-predicted operation data set into the resource prediction model, predict the resource usage requirements of multiple server nodes within a predetermined time range, and obtain the resource usage requirements corresponding to multiple time nodes within the predetermined time range. The resource prediction model is used to identify the dependency relationships between multiple time nodes and perform resource prediction; The adjustment module 34 is configured to generate resource allocation strategies corresponding to multiple server nodes based on the prediction results, and dynamically adjust the configured resources of multiple server nodes based on the resource allocation strategies. In some examples of this embodiment, the collection module 31 is specifically configured to collect the operation data of multiple performance indicators corresponding to each server node at a predetermined time interval based on the monitoring agent software configured for each server node in multiple server nodes. The multiple performance indicators include at least one or more of processor utilization rate, memory usage rate, storage input / output, and network traffic; and form an operation data set based on the operation data corresponding to multiple server nodes respectively.
[0092] In some examples of this embodiment, the processing module 32 is specifically configured to perform data cleaning processing on the running data set to obtain the cleaned running data set; determine the basic statistic corresponding to each performance metric among a plurality of performance metrics and the overall health index of each server node based on the cleaned running data set; and form a to-be-predicted running data set based on the basic statistic and the overall health index.
[0093] In some examples of this embodiment, the training process of the resource prediction model is specifically configured to collect the historical running data sets of a plurality of server nodes; use the historical basic statistic corresponding to each performance metric among the plurality of performance metrics of the plurality of server nodes in the historical running data set as input features, and use the historical resource requirements of the plurality of server nodes observed in the historical running data set as output features to form a training data set; establish a resource prediction model based on a long short-term memory network, train the resource prediction model using historical data, and adjust the parameters of the resource prediction model until the target accuracy rate is reached to obtain a trained resource prediction model.
[0094] In some examples of this embodiment, the adjustment module 34 is further configured to determine the resource occupancy and the comprehensive resource load index corresponding to each of the plurality of server nodes based on the to-be-predicted running data set; correspondingly, the adjustment module 34 is specifically configured to generate resource configuration policies corresponding to the plurality of server nodes according to the resource occupancy and the comprehensive resource load index, and dynamically adjust the configured resources of the plurality of server nodes based on the resource configuration policies.
[0095] In some examples of this embodiment, the adjustment module 34 is further configured to collect the historical workload data sets of a plurality of server nodes; input the historical workload data sets into a workload prediction model to predict the workload intensity of the plurality of server nodes within a predetermined time range, and obtain the workload intensity index corresponding to each of a plurality of time nodes within the predetermined time range, where the workload prediction model is used to identify the dependency relationship between the plurality of time nodes and perform workload prediction; and determine the workload status corresponding to each of the plurality of server nodes based on the workload intensity index.
[0096] In some examples of this embodiment, the adjustment module 34 is specifically further configured to compare the workload intensity indexes corresponding to multiple server nodes with a first workload intensity threshold and a second workload intensity threshold respectively, where the first workload intensity threshold is less than the second workload intensity threshold; when it is determined that the workload intensity index of the target server node among the multiple server nodes is less than the first workload intensity threshold, determine the workload status of the target server node as a gentle state; when it is determined that the workload intensity index of the target server node among the multiple server nodes is greater than or equal to the first workload intensity threshold and less than or equal to the second workload intensity threshold, determine the workload status of the target server node as a normal state; when it is determined that the workload intensity index of the target server node among the multiple server nodes is greater than the second workload intensity threshold, determine the workload status of the target server node as a peak state.
[0097] In some examples of this embodiment, the determination process of the first workload intensity threshold and the second workload intensity threshold is configured to perform clustering processing on the historical workload data set to cluster out a first data set corresponding to the gentle workload state and a second data set corresponding to the peak workload state in the historical workload data set; determine the first workload intensity threshold based on the first data set, and determine the second workload intensity threshold based on the second data set.
[0098] In some examples of this embodiment, the adjustment module 34 is specifically further configured to adjust the resource configuration policy based on the workload statuses corresponding to multiple server nodes respectively; dynamically adjust the configured resources of multiple server nodes according to the adjusted resource configuration policy.
[0099] In some examples of this embodiment, the adjustment module 34 is further configured to collect the hardware status information of multiple server nodes, and monitor the hardware status scores corresponding to multiple servers respectively based on the hardware status information; when it is monitored that the hardware status score of the target server node among the multiple server nodes is greater than the predetermined hardware status threshold, generate an alarm message for the target server node, where the target server node is any one of the multiple server nodes.
[0100] In some examples of this embodiment, the adjustment module 34 is specifically further configured to, when it is monitored that the hardware status score of the target server node among the multiple server nodes is greater than the predetermined hardware status threshold, migrate the key tasks in the target server node to other server nodes except the target server node among the multiple server nodes.
[0101] In some examples of this embodiment, the adjustment module 34 is further configured to collect the operation data of multiple server nodes after adjustment, and determine the adjusted comprehensive resource load index of the multiple server nodes based on the operation data after adjustment; verify the adjustment conditions of the multiple server nodes based on the adjusted comprehensive resource load index.
[0102] It should be noted that for other corresponding descriptions of each functional unit involved in a resource configuration device provided in this embodiment, reference can be made to Figure 1 the corresponding description in, which will not be elaborated here.
[0103] Based on the method as described above such as Figure 1 shown, correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above such as Figure 1 shown is implemented.
[0104] Based on the method as described above such as Figure 1 shown, correspondingly, this embodiment further provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above such as Figure 1 shown is implemented.
[0105] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of this application.
[0106] Based on the method as described above such as Figure 1 shown, and Figure 5 the virtual device embodiment as shown, in order to achieve the above object, this embodiment of the application further provides an electronic device, such as a personal computer or a server, and the device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as described above such as Figure 1 shown.
[0107] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc., and optionally the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.
[0108] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0109] The storage medium may also include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication between other hardware and software in the information processing physical device.
[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the related art, this embodiment can collect the operation data of multiple server nodes in the server system to form an operation data set, and perform preprocessing to obtain a to-be-predicted operation data set. By inputting the to-be-predicted data into the resource prediction model to predict the resource usage requirements of each server node, the resource usage requirements at multiple time nodes within a predetermined time period are obtained. Among them, the resource prediction model is used to identify the dependency relationships between time nodes. By using the resource prediction model, the long-term dependency relationships between multiple time nodes can be effectively captured, making the resource demand prediction more accurate. Furthermore, the system can make preparations for resource allocation in advance, avoid the situation of resource shortage or waste, optimize the resource allocation, improve the system response speed and user satisfaction, and at the same time reduce the operation cost; the resource allocation strategy generated based on the prediction results dynamically adjusts the resource allocation, enabling this embodiment to flexibly allocate resources according to actual needs, ensuring the balanced resource allocation of each server node, not only improving the resource utilization rate, reducing unnecessary energy consumption, but also enhancing the overall performance and efficiency of the server system, and at the same time enhancing the scalability and flexibility of the system.
[0111] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0112] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A resource configuration method, characterized in that: include: Collecting the operation data of multiple server nodes in the server system to form an operation data set; Performing data preprocessing on the operating data set to obtain an operating data set to be predicted; Inputting the to-be-predicted running data set into a resource prediction model, predicting the resource usage requirements of the multiple server nodes within a predetermined time range, and obtaining the resource usage requirements corresponding to the multiple time nodes within the predetermined time range, wherein the resource prediction model is used to identify the dependency relationship between the multiple time nodes and perform resource prediction; Resource configuration policies corresponding to the plurality of server nodes are generated according to the prediction results, and configuration resources of the plurality of server nodes are dynamically adjusted based on the resource configuration policies.
2. The method according to claim 1, characterized in that: The operation data set is formed by collecting the operation data of multiple server nodes in the server system, including: Based on the monitoring agent software configured on each of the multiple server nodes, collecting the operation data of multiple performance indicators corresponding to each of the multiple server nodes at a predetermined time interval, wherein the multiple performance indicators include at least one or more of processor utilization, memory usage, storage input and output, and network traffic; The operation data set is composed based on the operation data respectively corresponding to the multiple server nodes.
3. The method according to claim 2, characterized in that The step of performing data preprocessing on the operating data set to obtain the operating data set to be predicted includes: Performing data cleaning on the operating data set to obtain a cleaned operating data set; Determining, based on the cleaned operating data set, a basic statistic corresponding to each of the multiple performance indicators and an overall health index of each server node; The to-be-predicted operating data set is composed based on the basic statistics and the overall health index.
4. The method according to claim 3, characterized in that: The training process of the resource prediction model includes: Collecting historical operation data sets of the plurality of server nodes; Taking the historical basic statistics corresponding to each performance indicator of the multiple performance indicators of the multiple server nodes in the historical operation data set as input features, and taking the historical resource requirements of the multiple server nodes observed in the historical operation data set as output features to form a training data set; The resource prediction model is established based on a long short-term memory network, the resource prediction model is trained using historical data, and the parameters of the resource prediction model are adjusted until the target accuracy is reached to obtain the trained resource prediction model.
5. The method according to claim 4, characterized in that Before generating resource configuration policies corresponding to the plurality of server nodes according to the prediction results and dynamically adjusting the configuration resources of the plurality of server nodes based on the resource configuration policies, the method further includes: Based on the to-be-predicted running data set, determining resource occupancy and comprehensive resource load indexes respectively corresponding to the plurality of server nodes; Generating resource configuration policies corresponding to the plurality of server nodes according to the prediction results, and dynamically adjusting the configuration resources of the plurality of server nodes based on the resource configuration policies, includes: Generate resource configuration policies corresponding to the multiple server nodes according to the resource occupancy status and the comprehensive resource load index, and dynamically adjust the configuration resources of the multiple server nodes based on the resource configuration policies.
6. The method according to claim 1, characterized in that Before dynamically adjusting the configuration resources of the multiple server nodes based on the resource configuration strategy, the method further includes: collecting historical workload data sets of the plurality of server nodes; Inputting the historical workload data set into a load prediction model, predicting the workload intensity of the multiple server nodes within a predetermined time range, and obtaining workload intensity indexes corresponding to the multiple time nodes within the predetermined time range, wherein the load prediction model is used to identify the dependency relationship between the multiple time nodes and perform load prediction; Based on the workload intensity index, workload states corresponding to the multiple server nodes are determined respectively.
7. The method according to claim 6, characterized in that The determining, based on the workload intensity index, workload states corresponding to the plurality of server nodes respectively includes: Comparing the workload intensity indexes corresponding to the multiple server nodes respectively with a first workload intensity threshold and a second workload intensity threshold, respectively, the first workload intensity threshold being less than the second workload intensity threshold; In the case where it is determined that the workload intensity index of a target server node among the plurality of server nodes is less than the first workload intensity threshold, determining the workload state of the target server node as a flat state; In the case where it is determined that the workload intensity index of a target server node among the plurality of server nodes is greater than or equal to the first workload intensity threshold and less than or equal to the second workload intensity threshold, determining the workload state of the target server node as a normal state; When it is determined that the workload intensity index of the target server node among the plurality of server nodes is greater than the second workload intensity threshold, the workload state of the target server node is determined to be a peak state.
8. The method according to claim 7, characterized in that The process of determining the first workload intensity threshold and the second workload intensity threshold comprises: Performing clustering processing on the historical workload data set to cluster out a first data set corresponding to a flat working state and a second data set corresponding to a peak working state in the historical workload data set; The first workload intensity threshold is determined based on the first data set, and the second workload intensity threshold is determined based on the second data set.
9. The method according to claim 6, characterized in that Dynamically adjusting the configuration resources of the multiple server nodes based on the resource configuration strategy also includes: Adjusting the resource configuration strategy based on the workload states respectively corresponding to the multiple server nodes; The configuration resources of the multiple server nodes are dynamically adjusted according to the adjusted resource configuration strategy.
10. The method according to claim 1, characterized in that After generating resource configuration policies corresponding to the plurality of server nodes according to the prediction results, and dynamically adjusting the configuration resources of the plurality of server nodes based on the resource configuration policies, the method further includes: Collecting hardware status information of the plurality of server nodes, and monitoring hardware status scores corresponding to the plurality of servers respectively based on the hardware status information; When it is monitored that the hardware status score of a target server node among the multiple server nodes is greater than a predetermined hardware status threshold, alarm information of the target server node is generated, and the target server node is any one of the multiple server nodes.
11. The method according to claim 10, characterized in that The method further comprises: When it is detected that the hardware status score of a target server node among the multiple server nodes is greater than a predetermined hardware status threshold, critical tasks in the target server node are migrated to other server nodes among the multiple server nodes except the target server node.
12. The method according to claim 1, characterized in that After generating resource configuration policies corresponding to the plurality of server nodes according to the prediction results, and dynamically adjusting the configuration resources of the plurality of server nodes based on the resource configuration policies, the method further includes: Collecting the adjusted operation data of the multiple server nodes, and determining the adjusted comprehensive resource load indexes of the multiple server nodes based on the adjusted operation data; The adjustment status of the multiple server nodes is verified based on the adjusted comprehensive resource load index.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
14. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.
15. A computer program product having a computer program stored thereon, characterized in that: When the computer program product is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Load-aware scheduling method based on deep learning
CN119065835A
Server resource utilization rate improving method, system, device, medium and product
CN119292776A
Intelligent management method and system based on server cluster
CN119473803A
Cited By
PKS-based resource integration optimization method and apparatus, and storage medium
CN120315901A
Server resource allocation method and device, storage medium and electronic equipment
CN120407197A