Server resource allocation method and device, storage medium and electronic equipment

By training the data processing volume prediction model and dynamically adjusting the server resource allocation, the problem that the server resource allocation cannot match the business load requirements is solved, and the rational allocation of resources and the efficient operation of the business system is achieved.

CN120407197AInactive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510895965.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, server resource allocation cannot match the service load requirements, resulting in the problem of wasted or insufficient resources.

Method used

By training the data processing volume prediction model, predicting future data processing volume based on historical data processing volume, and dynamically adjusting server resource allocation.

Benefits of technology

It realizes dynamic and reasonable allocation of server resources, avoids resource waste and insufficient resources, and ensures the high reliability and efficiency of the business system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407197A_ABST
    Figure CN120407197A_ABST
Patent Text Reader

Abstract

The invention discloses a server resource allocation method and device, a storage medium and electronic equipment, and relates to the technical field of servers, the method is applied to a server cluster, the server cluster at least comprises a first server, and the method comprises the following steps: training a data processing amount prediction model through historical data processing amount of a first task, through the data processing amount prediction model, a second data processing amount generated by processing a first task by a first server in a second preset time period can be predicted according to a first data processing amount generated by processing the first task by the first server in a first preset time period; and adjusting server resources allocated to the first task by the first server according to the second data processing amount. The technical problem that server resource allocation cannot be matched with service load requirements is solved, and the technical effect of ensuring dynamic and reasonable allocation of server resources is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and in particular, to a method and device for allocating server resources, a storage medium, and an electronic device. Background Art

[0002] In the digital age, servers are responsible for massive data processing and business support tasks. With the diversification and complexity of various business scenarios, the requirements for the efficiency of server resource allocation are also getting higher and higher. The traditional server resource allocation methods are mostly static allocation or dynamic allocation based on simple rules. The static allocation method cannot adapt to the real-time changes of business loads, and it is easy to cause resource waste or resource shortage. The dynamic allocation method based on simple rules can adjust the allocation of server resources to a certain extent according to preset conditions, but it is difficult to accurately meet the complex, changeable, and uncertain business requirements.

[0003] It can be seen that there is a problem in the related art that the allocation of server resources cannot match the business load requirements. Summary of the Invention

[0004] This application provides a method and device for allocating server resources, a storage medium, and an electronic device to at least solve the problem that system anomalies are likely to occur during the concurrent upgrade of a distributed storage system in the related art.

[0005] This application provides a method for allocating server resources, which is applied to a server cluster. The server cluster includes at least a first server, and the method includes: obtaining a first data processing volume generated by the first server for processing a first task, where the first data processing volume represents the data processing volume within a first preset time period, and the end time of the first preset time period is the current time; inputting the first data processing volume into a data processing volume prediction model to obtain a prediction result output by the data processing volume prediction model, where the data processing volume prediction model is trained based on the historical data processing volumes generated by the first task for an initial prediction model, and the prediction result is used to indicate a second data processing volume generated by the first server for processing the first task, and the second data processing volume represents the data processing volume within a second preset time period, and the start time of the second preset time period is the current time; determining a first allocated resource allocated to the first task by the first server at the current time, and adjusting the first allocated resource to a second allocated resource based on the prediction result.

[0006] The present application also provides an allocation device for server resources, including: an acquisition module, configured to acquire a first data processing volume generated by the first server for processing a first task, where the first data processing volume represents the data processing volume within a first preset time period, and the end time of the first preset time period is the current time; a prediction module, configured to input the first data processing volume into a data processing volume prediction model to obtain a prediction result output by the data processing volume prediction model, where the data processing volume prediction model is trained based on the historical data processing volumes generated by the first task for an initial prediction model, and the prediction result is used to indicate a second data processing volume generated by the first server for processing the first task, the second data processing volume represents the data processing volume within a second preset time period, and the start time of the second preset time period is the current time; an allocation module, configured to determine a first allocation resource allocated by the first server to the first task at the current time, and adjust the first allocation resource to a second allocation resource based on the prediction result.

[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above server resource allocation methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program, when executed by a processor, implements the steps of any one of the above server resource allocation methods.

[0009] The present application also provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements the steps of any one of the above server resource allocation methods.

[0010] Through the present application, a data processing volume prediction model is trained based on the historical data processing volumes of the first task, and then the data processing volume prediction model can be used to predict the second data processing volume generated by the first server for processing the first task within a second preset time period according to the first data processing volume generated by the first server for processing the first task within a first preset time period, so as to adjust the server resources allocated by the first server to the first task according to the second data processing volume. The technical problem that the server resource allocation cannot match the service load requirements is solved, and the technical effect of ensuring the dynamic and reasonable allocation of server resources is achieved. Description of the Drawings

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1 It is a schematic diagram of an application scenario of a method for allocating server resources according to an embodiment of the present application;

[0013] Figure 2 It is a schematic flowchart of an optional method for allocating server resources according to an embodiment of the present application;

[0014] Figure 3 It is a schematic diagram of an optional system for allocating server resources according to an embodiment of the present application;

[0015] Figure 4 It is a block diagram of the structure of an intelligent prediction module of an optional system for allocating server resources according to an embodiment of the present application;

[0016] Figure 5 It is a block diagram of the structure of a resource scheduling execution module of an optional system for allocating server resources according to an embodiment of the present application;

[0017] Figure 6 It is a block diagram of the structure of an optional device for allocating server resources according to an embodiment of the present application. Detailed implementation manners

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0019] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following will further describe the present application in detail with reference to the drawings and specific implementation manners.

[0021] According to one aspect of the embodiments of the present application, a method for allocating server resources is provided. Optionally, in this embodiment, the above method for allocating server resources may be, but is not limited to, applied to a hardware environment including a terminal device 102 and a server 104 as shown in Figure 1 Figure. The server 104 can be connected to the terminal device 102 through a network and can be used to provide services (such as application services, etc.) for the terminal device 102 or the client installed on the terminal device 102. A database can be set up on the server 104 or independently of the server 104 to provide data storage services for the server 104.

[0022] The above network may include, but is not limited to, at least one of the following: wired network, wireless network. The above wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be, but is not limited to, a PC (Personal Computer), mobile phone, tablet computer, etc. The server 104 may be, but is not limited to, a cloud server, a server cluster or other server types.

[0023] The method for allocating server resources in the embodiments of the present application may be executed by the server 104, or may be executed by the terminal device 102, or may be jointly executed by the server 104 and the terminal device 102. Among them, the execution of the method for allocating server resources in the embodiments of the present application by the terminal device 102 may also be executed by the client installed thereon.

[0024] Taking the allocation method of server resources in this embodiment executed by the terminal device 102 as an example, here, the terminal device 102 can be a physical host. The allocation method of server resources in this embodiment is applied to the physical host, and the memory that the physical host can call is divided into multiple memory levels. One memory level of the multiple memory levels includes at least one type of memory, and the multiple memory levels include a first memory level corresponding to the physical memory of the physical host. Here, the physical host can be an enterprise-level server, a cluster server, an office computer, an embedded device, etc., which are entity devices that can serve as the underlying hardware support in a virtualized environment; the physical memory is a physical memory module directly and tightly connected to the host hardware, and is the core and foundation of the memory architecture system. The physical memory is usually composed of dynamic random access memory (abbreviated as DRAM), which has extremely fast read and write speeds and can respond to the memory access requests of the processor with a nanosecond-level response time, making the physical memory suitable for carrying the core code of the virtual machine operating system, frequently called system function libraries, and key process data that is running at high speed. For example, in the initial stage of virtual machine startup, the operating system kernel needs to quickly load and initialize various hardware drivers and establish a basic system running environment. At this time, the physical memory can complete data reading and writing operations with extremely high efficiency, ensuring that the virtual machine can start quickly and stably. During the operation of the virtual machine, application program parts with extremely high requirements for memory read and write performance, such as the transaction processing module of the database management system and the real-time rendering engine, also rely on the physical memory to ensure their efficient operation, thereby maintaining the fluency and response timeliness of the entire virtual machine system.

[0025] Figure 2 It is a schematic flowchart of an optional method for allocating server resources according to an embodiment of the present application. As Figure 2 shown, the process of this method may include the following steps:

[0026] Step S202, obtain the first data processing amount generated by the first server for processing the first task, where the first data processing amount represents the data processing amount within a first preset time period, and the end moment of the first preset time period is the current moment;

[0027] Optionally, in the above step S202, for example, the first server is a server of an e-commerce platform, the first task is a transaction order service, and the first preset time period can be the week before a promotion activity. The first data processing amount is the data processing amount generated by the server for processing the order service in the week before the promotion activity.

[0028] Step S204: Input the first data processing volume into the data processing volume prediction model to obtain the prediction result output by the data processing volume prediction model. The data processing volume prediction model is trained based on the historical data processing volumes generated by the first task for the initial prediction model. The prediction result is used to indicate the second data processing volume generated by the first server for processing the first task. The second data processing volume represents the data processing volume within the second preset time period, and the start time of the second preset time period is the current time;

[0029] Optionally, in the above step S204, for example, the second preset time period may be within one hour after the start of the promotion activity.

[0030] Step S206: Determine the first allocated resource assigned by the first server to the first task at the current time, and adjust the first allocated resource to the second allocated resource based on the prediction result.

[0031] Through the above steps, the data processing volume prediction model is trained according to the historical data processing volumes of the first task. Furthermore, the data processing volume prediction model can predict the second data processing volume generated by the first server for processing the first task within the second preset time period based on the first data processing volume generated by the first server for processing the first task within the first preset time period, so as to adjust the server resources allocated by the first server for the first task according to the second data processing volume. This solves the technical problem that the server resource allocation cannot match the business load requirements and achieves the technical effect of ensuring the dynamic and reasonable allocation of server resources.

[0032] In an exemplary embodiment, determining the first allocated resource assigned by the first server to the first task at the current time and adjusting the first allocated resource to the second allocated resource based on the prediction result includes: obtaining the first resource configuration value configured by the first server for the first task; in the case where it is determined that the second data processing volume is less than the first preset value, adjusting the first resource configuration value to the second resource configuration value, and allocating the second allocated resource to the first task based on the second resource configuration value, where the second resource configuration value is less than the first resource configuration value; in the case where it is determined that the second data processing volume is greater than the second preset value, adjusting the first resource configuration value to the third resource configuration value, and allocating the second allocated resource to the first task based on the third resource configuration value, where the second preset value is greater than the first preset value and the third resource configuration value is greater than the first resource configuration value.

[0033] Optionally, in the above embodiments, when the second data processing volume is greater than the first preset value, it indicates that the service load corresponding to the first task will be at a relatively high level. At this time, more server resources need to be allocated to the first task to ensure the normal operation of the service. When the second data processing volume is less than the second preset value, it indicates that the service load corresponding to the first task will be at a relatively low level. At this time, the server resources allocated to the first task need to be reduced to avoid the idle of server resources during the business trough period.

[0034] Through the above embodiments, the server resources allocated to the service can be dynamically adjusted according to the real-time change of the service load, thus avoiding the slow system response or even system crash caused by insufficient resources during the sudden increase of business volume such as e-commerce promotion, and also solving the problem of idle server resources during the business trough period.

[0035] In an exemplary embodiment, the method further includes: when it is determined that the second data processing volume is greater than the third preset value, determining a second server from the server cluster, where the third preset value is greater than the second preset value, and the resource occupancy rate of the second server is less than the preset resource occupancy rate; allocating resources for the first task in the second server so that the second server starts to process the first task.

[0036] Optionally, in the above embodiments, when the second data processing volume is greater than the third preset value, it indicates that the service load corresponding to the first task is about to reach the upper limit of the first server's capacity, and the first server can no longer continue to process the data generated by the first task. At this time, an idle second server can be selected from the server cluster, and the first task can be migrated from the first server to the second server.

[0037] Optionally, in the above embodiments, migrating the first task from the first server to the second server can be to migrate all of the first task to the second server, or to migrate some data to the second server, or the first server can continue to process the existing data, and the second server can process the new data of the first task. In addition, when the first server fails, the tasks on the first server can also be migrated to other servers.

[0038] Through the above embodiments, the service can be quickly migrated to a normal server when the server is overloaded or fails, avoiding service interruption and ensuring the high reliability of the service system.

[0039] In an exemplary embodiment, determining a second server from a server cluster includes: sending a heartbeat detection signal to the servers in the server cluster at a preset time interval, where the heartbeat detection signal includes the signal sending time; receiving a heartbeat response signal for the server to respond to the heartbeat detection signal, where the heartbeat response signal includes the signal response time, the resource occupancy rate of the server, and the operating state of the server; determining the servers in the server cluster that meet the preset conditions according to the heartbeat response signal, where the preset conditions include: the signal response time period is less than the preset time period, the resource occupancy rate is less than the preset resource occupancy rate, the operating state is normal, and the signal response time period represents the time period between the signal response time and the signal sending time; and determining the server with the smallest resource occupancy rate among the servers that meet the preset conditions as the second server.

[0040] Optionally, in the above embodiment, for example, the control center of the server cluster sends a heartbeat detection signal to each server every 30 seconds. After receiving the heartbeat detection signal, the servers in the cluster immediately return a response signal, and the response signal includes the status information of the server, such as the resource occupancy rate, etc. If the control center does not receive the response signal from the server or the response time exceeds 60 seconds, it is considered that the server may have a problem. According to the response signals of each server, the resource usage of each server can be determined, and then the server with the smallest resource occupancy rate can be selected to undertake the processing tasks of the high-load server.

[0041] Through the above embodiment, the server status can be monitored in a timely manner in a distributed server cluster environment, the server resources can be integrated, the overall efficiency of the server cluster can be improved, and the dynamic and intelligent allocation of server resources is realized.

[0042] In an exemplary embodiment, there is a main network link and a standby network link between the first server and the second server. After allocating resources for the first task in the second server, the method further includes: performing data transmission between the first server and the second server through the main network link to synchronize the data processing progress of the first task; and performing data transmission between the first server and the second server through the standby network link when it is determined that the main network link fails.

[0043] Optionally, in the above embodiment, the servers in the server cluster ensure the consistency and accuracy of the resource information in the cluster through a distributed consistency algorithm. The distributed resource management network constructed in the server cluster adopts a redundant network link design. When the main link fails, it automatically switches to the standby link, ensuring the uninterrupted transmission of the resource status information between the servers in the cluster and increasing the fault tolerance rate of the server cluster.

[0044] In an exemplary embodiment, the first server is further configured to process a second task. The method further includes: when it is determined that both the first task and the second task require additional allocated resources, comparing the first task priority of the first task and the second task priority of the second task; when it is determined that the first task priority is greater than the second priority, preferentially adding allocated resources to the first task; and when it is determined that the first task priority is less than the second priority, preferentially adding allocated resources to the second task.

[0045] Optionally, in the above embodiment, for example, in an e-commerce platform, the payment service priority is higher than the product browsing service. When the traffic of the e-commerce platform surges, the server preferentially schedules server resources for the payment service to ensure the reliability of the core service.

[0046] In an exemplary embodiment, before inputting the first data processing volume into the data processing volume prediction model, the method further includes: obtaining the first historical data processing volume generated by the first server in processing the first task during a first historical period, and the second historical data processing volume generated by the first server in processing the first task during a second historical period, where the first historical period is before the second historical period; using the first historical data processing volume as input data and the second historical data processing volume as output data to train an initial prediction model to obtain a data processing volume prediction model. The data processing volume prediction model includes a convolutional neural network, a long short-term memory network, and a fully connected layer. The convolutional neural network is used to extract local features in the input data, the long short-term memory network is used to capture the long-term dependencies of the local features in the time series, and the fully connected layer is used to integrate the local features and the long-term dependencies to obtain the output data.

[0047] Optionally, in the above embodiment, for example, the first task is the order service of an e-commerce platform. The first historical data processing volume includes the past three years' promotion period and daily sales data of the e-commerce platform. A data processing volume prediction model is constructed through a convolutional neural network (CNN, Convolutional Neural Network) and a long short-term memory network (LSTM, Long Short-Term Memory), making full use of the local feature extraction ability of the CNN and the time series processing ability of the LSTM to improve the accuracy of resource demand prediction, and combining real-time business data to predict the resource demand of the business volume at different times on the day of the promotion activity.

[0048] Optionally, in the above embodiments, the process of inputting the first data processing volume into the data processing volume prediction model to obtain the prediction result output by the data processing volume prediction model includes: dividing the first data processing volume according to a preset time window to generate a plurality of time window data; extracting local features from the plurality of time window data through a convolutional neural network, where the local features represent the correlation between adjacent time window data in the plurality of time window data; performing standardized data processing on the local features to obtain a three-dimensional data matrix, where the dimensions of the three-dimensional data matrix include the number of samples, the time step, and the feature dimension, and the standardized processing is used to convert the data format of the local features into a preset data format; extracting the long-term dependence relationship of the three-dimensional data matrix through a long short-term memory network, where the long short-term memory network includes an input gate, a forget gate, and an output gate, the input gate is used to receive the three-dimensional data matrix, the forget gate is used to filter the short-term data fluctuation features of the three-dimensional data matrix to obtain long-term data trend features, and the output gate is used to output the long-term dependence relationship based on the long-term data trend features; integrating the local features and the long-term dependence relationship through a fully connected layer to obtain the prediction result.

[0049] In this embodiment, the CNN is used to extract local features from the input data. For example, during the data surge period, the CNN scans the server resource occupancy data within a preset duration window through a convolutional kernel of a preset size to extract local correlation features. After strengthening the local correlation features using the ReLU activation function, the dimension is compressed through a pooling layer to generate local features. The local correlation features are, for example, features such as the sudden increase in the co-occupancy of multiple server resources and batch data processing.

[0050] Optionally, for example, the above data surge period can be the payment peak period of an e-commerce platform. The above CNN can scan the disk I / O data within a 5-minute time window through a convolutional kernel of 3×1 specification. The local feature of the sudden increase in the co-occupancy of multiple server resources can be set to the sudden increase in the co-occupancy of the CPU and memory within 10 seconds, and the local feature of batch data processing can be set to three consecutive burst reads and writes.

[0051] The LSTM is used to capture the long-term dependence relationship of local features in the time series. For example, the LSTM can extract the data trend of a specific time period from historical data. First, the local features corresponding to the historical data are organized into a three-dimensional matrix [number of samples, time step, feature dimension]. The weights are calculated through the hidden layer neurons. The three-dimensional matrix is received through the input gate, the irrelevant fluctuation features are discarded through the forget gate to obtain the long-term data trend features, and the long-term dependence relationship is output through the output gate based on the long-term data trend features.

[0052] Optionally, for example, the above historical data is data during the large promotion period in the past three years on an e-commerce platform (for example, the large promotion duration is 7 days), and the specific time period is 1 hour after the start of the large promotion. The local features corresponding to the historical data are organized into a three-dimensional matrix. For example, the CPU utilization data for the past 7 days can be organized into a three-dimensional matrix [number of samples, 1008 time steps (7 days × 144 minutes), feature dimension]. The above irrelevant fluctuation features include, for example, data fluctuations during low-load periods at night.

[0053] The fully connected layer is used to integrate local features and long-term dependencies to obtain output data. By integrating the local features extracted by the CNN and the long-term dependencies captured by the LSTM, a CNN-LSTM hybrid model is constructed, making full use of the local feature extraction ability of the CNN and the time series processing ability of the LSTM to improve the accuracy of resource demand prediction.

[0054] Exemplarily, the above CNN processes the 5-minute window data during the large promotion period on the e-commerce platform and outputs local features such as "the CPU occupancy rate increases pulse-wise by 15% - 20% within 10 minutes after the release of each wave of promotion previews". The LSTM processes the CPU utilization time series during the large promotion periods in the past three years on the e-commerce platform and outputs long-term trend features such as "the overall resource growth demand during the large promotion period is 5 times". The fully connected layer integrates the above two types of features to generate the resource demand curve for the next 4 hours.

[0055] Through the above method, the data processing volume prediction model can more accurately predict the second data processing volume generated by the first server when processing the first task within the second preset time period, thereby providing a more accurate basis for the dynamic adjustment of server resources.

[0056] Through the above embodiments, by predicting the real-time resource requirements of the service for the server and combining the real-time load of the server and the urgency of the service, an efficient and reasonable resource allocation strategy can be determined.

[0057] Next, an optional method for allocating server resources in an embodiment of the present application will be described in conjunction with optional embodiments. In an optional embodiment, as Figure 3 shown, the method for allocating server resources can be implemented in cooperation with the following server resource allocation system. In an optional embodiment, as Figure 3 shown, the above server resource allocation system includes the following modules:

[0058] 1. Data Acquisition Module: It is used to collect business data, resource usage data, and network status data during the operation of the server in real time. In the server of an e-commerce platform, the data acquisition module is connected to each hardware interface of the server through a standardized data transmission protocol, and can collect business data such as the number of product views, the number of order generations, and the number of payment transactions in real time, as well as resource usage data such as CPU usage rate, memory occupancy, and disk I / O read / write speed, and can also collect network status data such as real-time traffic and latency of network bandwidth. For example, in the week before a promotion event, the data acquisition module continuously collects data to provide a basis for subsequent prediction and decision-making.

[0059] 2. Intelligent Prediction Module: It is used to predict the future demand for server resources by the business based on the data collected by the data acquisition module, using machine learning algorithms and deep learning models.

[0060] As Figure 4 shown, the intelligent prediction module also includes the following units:

[0061] Data Preprocessing Unit 42: It is used to perform preprocessing operations such as cleaning and normalization on the collected data.

[0062] Model Training Unit 44: It is used to train and optimize machine learning models and deep learning models with the preprocessed data.

[0063] Prediction Result Generation Unit 46: It is used to output the prediction result of future resource demand using the trained model.

[0064] In an optional embodiment, for example, the specific steps for the data preprocessing unit to clean a large amount of messy data collected are as follows:

[0065] Remove duplicate values: Traverse the data set, and find and delete exactly the same duplicate records by comparing the unique identifier or all fields of each record.

[0066] Handle missing values: Identify the missing values in the data set, and you can use relevant functions or tools to mark the positions of the missing values. For numerical data, the mean, median, or mode can be used for filling; for non-numerical data, the most common value (mode) can be selected for filling according to the specific situation, or more complex machine learning algorithms can be used for predictive filling.

[0067] Clean error values: Check the value range of the data. For example, if a negative number appears in the age field, it is regarded as an error, and the error data is corrected or deleted; verify the data format. For example, the date field should conform to a specific format, and the data that does not conform to the format needs to be corrected to ensure data logic consistency. For example, the quantity of goods in the order data cannot be negative, etc.

[0068] The data preprocessing unit can normalize the cleaned data in the following ways:

[0069] Max - Min normalization: Calculate the maximum and minimum values in the dataset, and then use the formula to map the data to the interval [0, 1], where, represents the normalized value, represents the original data, represents the minimum value of the dataset, represents the maximum value of the dataset.

[0070] Z - score normalization: Calculate the mean and standard deviation of the dataset, and use the formula to normalize the data so that the data follows a normal distribution with a mean of 0 and a standard deviation of 1;

[0071] Log - transformation normalization: For data with a large data span and a right - skewed distribution, take the logarithm of non - zero data, that is to compress the data range and make the data distribution more uniform.

[0072] The training process of the model training unit is as follows:

[0073] For example, using the promotion period and daily sales data in the past few years, combined with machine learning models and deep learning models to train a prediction model. In this embodiment, the machine learning model selects the combination of Long Short - Term Memory (LSTM) and Support Vector Regression (SVR, Support Vector Regression) algorithms; LSTM is good at dealing with long - term dependencies in time series and can effectively capture the trend of server resource usage changing over time; SVR performs well in small - sample and non - linear regression problems and can accurately model complex resource demand patterns. The deep learning model selects Convolutional Neural Network (CNN). CNN has an advantage in processing data with local correlations and can effectively extract local patterns in resource usage data. The fully - connected layer is mainly responsible for integrating the features extracted by CNN and LSTM to build a CNN - LSTM hybrid prediction model, making full use of the local feature extraction ability of CNN and the time - series processing ability of LSTM to improve the accuracy of resource demand prediction.

[0074] The CNN is used to extract local features from the input data. For example, during the peak payment period, the CNN scans the disk I / O data within a 5-minute time window through a 3×1 convolutional kernel, identifies the local feature of "sudden increase in the collaborative occupancy of the CPU and memory within 10 seconds", and after strengthening the feature through the ReLU activation function, compresses the dimension through the pooling layer to generate a local pattern feature vector. In specific applications, the disk I / O read and write rate is divided according to time windows, and the data of each window forms a two-dimensional matrix. The CNN extracts local correlation features through a sliding window, such as the local feature of "data batch processing" corresponding to three consecutive burst reads and writes.

[0075] The LSTM is used to capture the long-term dependencies of local features in the time series. For example, in the scenario of a major promotion on an e-commerce platform, the LSTM can extract the time trend of "the resource demand shows an exponential growth within 1 hour at the start of the event" from the data of the promotion period in the past 3 years, discard the irrelevant fluctuations during the low-load period at night through the forget gate, and output a feature vector containing the long-term trend. The specific process is as follows: Organize the local features corresponding to the CPU utilization data of the past 7 days into a three-dimensional matrix [number of samples, 1008 time steps (7 days × 144 minutes), feature dimension], calculate the weights through the neurons in the hidden layer, receive the three-dimensional matrix through the input gate, discard the irrelevant fluctuation features through the forget gate to obtain the long-term data trend features, and output the long-term dependencies based on the long-term data trend features through the output gate.

[0076] The fully connected layer is used to integrate local features and long-term dependencies to obtain the output data. By integrating the local features extracted by the CNN and the long-term dependencies captured by the LSTM, a CNN-LSTM hybrid model is constructed, which makes full use of the local feature extraction ability of the CNN and the time series processing ability of the LSTM to improve the accuracy of resource demand prediction.

[0077] The prediction process of the hybrid model is as follows: The LSTM branch processes the CPU utilization time series of the past 7 days and outputs long-term trend features (such as the overall resource growth demand during the major promotion is 5 times); the CNN branch processes the 5-minute window data of the recent preheating activities and outputs local pattern features (such as the CPU occupancy rate increases pulsatively by 15% - 20% within 10 minutes after the release of each wave of promotion previews); the two types of features are concatenated through the fully connected layer to generate the resource demand curve for the next 4 hours.

[0078] Based on the trained hybrid prediction model, the prediction result generation unit predicts the demand for server resources by the business volume at different time periods on the event day. For example, it is predicted that within 1 hour after the start of the promotion activity, the order generation volume will surge, and the demand for CPU and memory resources will reach 5 times that of daily, and the network bandwidth demand will increase by 3 times.

[0079] 3. Resource Allocation Decision Module: It is used to generate the optimal server resource allocation strategy based on the prediction results of the intelligent prediction module. Considering the predicted resource requirements, business priorities (such as the priority of payment services is higher than that of product browsing), and the real-time load of the server, it quickly traverses and filters multiple resource allocation schemes, evaluates and sorts the schemes to determine the optimal one. For example, when determining the start of an event, 80% of the CPU resources and 70% of the memory resources are preferentially allocated to payment services to ensure the smooth payment process, and sufficient bandwidth is allocated to product browsing services at the same time.

[0080] 4. Resource Scheduling Execution Module: It is used to dynamically adjust and allocate server resources according to the strategy generated by the resource allocation decision module.

[0081] As Figure 5 shown, the resource scheduling execution module also includes the following units:

[0082] Resource Allocation and Adjustment Unit 52: It is responsible for real-time adjustment of the server's local resources according to the resource allocation strategy;

[0083] Service Migration Unit 54: In a distributed server cluster environment, when service migration is required, it is responsible for migrating services from overloaded or faulty servers to target servers, and ensuring service continuity and data integrity. For example, the resource allocation and adjustment unit reduces the allocation of local CPU, memory and other resources of the server during the off-peak period of business to avoid resource waste; increases resource allocation during the peak period of business to ensure the efficient operation of the business, and maintains stable performance according to the changes in business requirements. For example, during an e-commerce promotion, resources are allocated to core services.

[0084] Service Migration Unit: When the server resources are overloaded or faulty, quickly migrate the service to a normal server to avoid service interruption and ensure system availability. For example, during a promotion event, if a server for processing orders is overloaded with resources, automatically migrate some order processing services to a server with idle resources to ensure the continuity of order processing services and data integrity.

[0085] 5. Distributed Cluster Management Module: It is used to monitor the resource status of each server in the distributed server cluster environment in real time, and realize the dynamic integration of cluster resources and service migration. For example, it uses the heartbeat detection mechanism to monitor the running status and resource status of each server in the cluster in real time, and uses the distributed consistency algorithm to ensure the consistency and accuracy of resource information. When the user access volume in a certain region suddenly increases, resulting in a high load on the server cluster in that region, automatically divert some service requests to the server cluster with idle resources in other regions to realize the dynamic integration of cluster resources.

[0086] 6. Security Protection Module: It is used to ensure data security and system stability. The security protection module establishes a secure data interaction channel with other modules, and encrypts the commodity information, user data, and order data transmitted between modules. It monitors network attack behaviors in real time. During the event, if a network attack on the payment system is detected, it immediately activates the emergency protection mechanism, such as restricting access from abnormal IPs and adjusting network traffic strategies, to ensure data security and the stable operation of the system.

[0087] The above server resource allocation system adopts a modular design, and each module has a standardized interface, which is convenient for upgrading, replacing, or expanding new functions of a single module, and flexibly adapting to changes in business scenarios and technological developments. Through the coordinated operation of each module of the above server resource allocation system, the dynamic and reasonable allocation of server resources is achieved.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (for example, Read-Only Memory (ROM) / Random Access Memory (RAM), magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present application.

[0089] According to another aspect of the embodiments of the present application, there is also provided an allocation device for server resources. This allocation device for server resources can be used to implement the allocation method of server resources provided in the above embodiments, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0090] Figure 6 is a structural block diagram of an optional allocation device for server resources according to the embodiments of the present application. As shown in Figure 6 the allocation device for server resources includes:

[0091] An acquisition module 62, configured to acquire a first data processing amount generated by a first server when processing a first task, where the first data processing amount represents the data processing amount within a first preset time period, and the end time of the first preset time period is the current time;

[0092] A prediction module 64 is configured to input a first data processing volume into a data processing volume prediction model to obtain a prediction result output by the data processing volume prediction model. The data processing volume prediction model is trained based on historical data processing volumes generated by a first task for an initial prediction model. The prediction result is used to indicate a second data processing volume generated by the first server for processing the first task. The second data processing volume represents the data processing volume within a second preset time period, and the start time of the second preset time period is the current time.

[0093] An allocation module 66 is configured to determine a first allocation resource allocated by the first server to the first task at the current time, and adjust the first allocation resource to a second allocation resource based on the prediction result.

[0094] Through the server resource allocation device provided by this application, a data processing volume prediction model is trained based on the historical data processing volume of the first task. Furthermore, the data processing volume prediction model can be used to predict the second data processing volume generated by the first server for processing the first task within a second preset time period according to the first data processing volume generated by the first server for processing the first task within a first preset time period. Thus, the server resources allocated by the first server for the first task are adjusted according to the second data processing volume. This solves the technical problem that server resource allocation cannot match the business load requirements, and achieves the technical effect of ensuring dynamic and reasonable allocation of server resources.

[0095] In an exemplary embodiment, the allocation module 66 is further configured to determine a first allocation resource allocated by the first server to the first task at the current time, and adjust the first allocation resource to a second allocation resource based on the prediction result, including: obtaining a first resource configuration value configured by the first server for the first task; in the case where it is determined that the second data processing volume is less than a first preset value, adjusting the first resource configuration value to a second resource configuration value, and allocating a second allocation resource for the first task based on the second resource configuration value, where the second resource configuration value is less than the first resource configuration value; in the case where it is determined that the second data processing volume is greater than a second preset value, adjusting the first resource configuration value to a third resource configuration value, and allocating a second allocation resource for the first task based on the third resource configuration value, where the second preset value is greater than the first preset value, and the third resource configuration value is greater than the first resource configuration value.

[0096] In an exemplary embodiment, the allocation module 66 is further configured to, in the case where it is determined that the second data processing volume is greater than a third preset value, determine a second server from a server cluster, where the third preset value is greater than the second preset value, and the resource occupancy rate of the second server is less than a preset resource occupancy rate; allocate resources for the first task in the second server so that the second server starts to process the first task.

[0097] In an exemplary embodiment, the allocation module 66 is further configured to send a heartbeat detection signal to the servers in the server cluster at a preset time interval, where the heartbeat detection signal includes the signal sending time; receive a heartbeat response signal for the server to respond to the heartbeat detection signal, where the heartbeat response signal includes the signal response time, the resource occupancy rate of the server, and the operating state of the server; determine the servers in the server cluster that meet the preset conditions according to the heartbeat response signal, where the preset conditions include: the signal response time period is less than the preset time period, the resource occupancy rate is less than the preset resource occupancy rate, and the operating state is normal, and the signal response time period represents the time period between the signal response time and the signal sending time; and determine the server with the smallest resource occupancy rate among the servers that meet the preset conditions as the second server.

[0098] In an exemplary embodiment, there is a primary network link and a standby network link between the first server and the second server. The allocation module 66 is further configured to perform data transmission between the first server and the second server through the primary network link to synchronize the data processing progress of the first task; and perform data transmission between the first server and the second server through the standby network link when it is determined that the primary network link fails.

[0099] In an exemplary embodiment, the first server is further configured to process a second task. The server resource allocation device is further configured to compare the first task priority of the first task and the second task priority of the second task when it is determined that both the first task and the second task need to increase the allocated resources; preferentially allocate additional resources to the first task when it is determined that the first task priority is greater than the second priority; and preferentially allocate additional resources to the second task when it is determined that the first task priority is less than the second priority.

[0100] In an exemplary embodiment, the prediction module 64 is further configured to obtain the first historical data processing volume generated by the first server for processing the first task in the first historical time period, and the second historical data processing volume generated by the first server for processing the first task in the second historical time period, where the first historical time period is before the second historical time period; use the first historical data processing volume as input data and the second historical data processing volume as output data to train an initial prediction model to obtain a data processing volume prediction model, where the data processing volume prediction model includes a convolutional neural network, a long short-term memory network, and a fully connected layer. The convolutional neural network is used to extract local features from the input data, the long short-term memory network is used to capture the long-term dependencies of the local features in the time series, and the fully connected layer is used to integrate the local features and the long-term dependencies to obtain the output data.

[0101] For the description of the features in the corresponding embodiments of the above server resource allocation device, reference may be made to the relevant descriptions in the corresponding embodiments of the server resource allocation method, which will not be elaborated here one by one.

[0102] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the server resource allocation method.

[0103] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the server resource allocation method when running.

[0104] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0105] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server resource allocation method.

[0106] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server resource allocation method.

[0107] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0108] The above has introduced in detail a method and apparatus, a storage medium, and an electronic device of a distributed storage system provided in this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for allocating server resources, characterized in that, It is applied to a server cluster, and the server cluster includes at least a first server, including: Obtain the first data processing volume generated by the first server for processing the first task, where the first data processing volume represents the data processing volume within a first preset time period, and the end time of the first preset time period is the current time; Input the first data processing volume into a data processing volume prediction model to obtain a prediction result output by the data processing volume prediction model. The data processing volume prediction model is trained based on the historical data processing volume generated by the first task for an initial prediction model. The prediction result is used to indicate the second data processing volume generated by the first server for processing the first task. The second data processing volume represents the data processing volume within a second preset time period, and the start time of the second preset time period is the current time; Determine the first allocation resource allocated to the first task by the first server at the current time, and adjust the first allocation resource to a second allocation resource based on the prediction result.

2. The method for allocating server resources according to claim 1, characterized in that, Determine the first allocation resource allocated to the first task by the first server at the current time, and adjust the first allocation resource to a second allocation resource based on the prediction result, including: Obtain the first resource configuration value configured by the first server for the first task; In the case where it is determined that the second data processing volume is less than a first preset value, adjust the first resource configuration value to a second resource configuration value, and allocate the second allocation resource to the first task based on the second resource configuration value, where the second resource configuration value is less than the first resource configuration value; In the case where it is determined that the second data processing volume is greater than a second preset value, adjust the first resource configuration value to a third resource configuration value, and allocate the second allocation resource to the first task based on the third resource configuration value, where the second preset value is greater than the first preset value, and the third resource configuration value is greater than the first resource configuration value.

3. The method for allocating server resources according to claim 2, characterized in that, The method further includes: In the case where it is determined that the second data processing volume is greater than a third preset value, determine a second server from the server cluster, where the third preset value is greater than the second preset value, and the resource occupancy rate of the second server is less than a preset resource occupancy rate; Allocate resources for the first task in the second server so that the second server starts to process the first task.

4. The method for allocating server resources according to claim 3, characterized in that, Determine a second server from the server cluster, including: Send a heartbeat detection signal to the servers in the server cluster at a preset time interval, where the heartbeat detection signal includes the signal sending time; Receiving a heartbeat response signal in response to the heartbeat detection signal by the server, where the heartbeat response signal includes a signal response time, a resource occupancy rate of the server, and an operating state of the server; Determining a server in the server cluster that meets a preset condition according to the heartbeat response signal, where the preset condition includes: a signal response time period is less than a preset time period, a resource occupancy rate is less than the preset resource occupancy rate, and the operating state is normal, and the signal response time period represents a time period between the signal response time and the signal sending time; Determining the server with the smallest resource occupancy rate among the servers that meet the preset condition as the second server.

5. The method for allocating server resources according to claim 3, wherein There is a main network link and a standby network link between the first server and the second server. After allocating resources for the first task in the second server, the method further includes: Performing data transmission between the first server and the second server through the main network link to synchronize the data processing progress of the first task; In the case of determining that the main network link fails, performing data transmission between the first server and the second server through the standby network link.

6. The method for allocating server resources according to claim 1, wherein The first server is further configured to process a second task, and the method further includes: In the case of determining that both the first task and the second task need to increase allocated resources, comparing a first task priority of the first task and a second task priority of the second task; In the case of determining that the first task priority is greater than the second priority, preferentially increasing the allocated resources for the first task; In the case of determining that the first task priority is less than the second priority, preferentially increasing the allocated resources for the second task.

7. The method for allocating server resources according to claim 1, wherein Before inputting the first data processing amount into the data processing amount prediction model, the method further includes: Obtaining a first historical data processing amount generated by the first server for processing the first task in a first historical time period, and a second historical data processing amount generated by the first server for processing the first task in a second historical time period, where the first historical time period is before the second historical time period; Training an initial prediction model with the first historical data processing amount as input data and the second historical data processing amount as output data to obtain the data processing amount prediction model, where the data processing amount prediction model includes a convolutional neural network, a long short-term memory network, and a fully connected layer. The convolutional neural network is used to extract local features in the input data, the long short-term memory network is used to capture long-term dependencies of the local features in a time series, and the fully connected layer is used to integrate the local features and the long-term dependencies to obtain the output data.

8. An allocation device for server resources, characterized in that, Including: An acquisition module, configured to acquire a first data processing volume generated by the first server in processing a first task, where the first data processing volume represents the data processing volume within a first preset time period, and the end time of the first preset time period is the current time; A prediction module, configured to input the first data processing volume into a data processing volume prediction model to obtain a prediction result output by the data processing volume prediction model, where the data processing volume prediction model is obtained by training an initial prediction model based on historical data processing volumes generated by the first task, and the prediction result is used to indicate a second data processing volume generated by the first server in processing the first task, and the second data processing volume represents the data processing volume within a second preset time period, and the start time of the second preset time period is the current time; An allocation module, configured to determine a first allocation resource allocated by the first server to the first task at the current time, and adjust the first allocation resource to a second allocation resource based on the prediction result.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the method for allocating server resources according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the method for allocating server resources according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Resource adjustment method and device and storage medium

    CN115756812A

  • Load prediction method and device based on hybrid model in container cloud environment

    CN119440977A

  • Server resource scheduling system integrating AI and edge computing

    CN119718682A

  • Resource prediction method and device based on NeuralProphet, and storage medium

    CN120066678A

  • Resource allocation method, storage medium, electronic equipment and program product

    CN120066798A

Cited By

  • Service distribution method based on machine learning and related device

    CN120782215A