Running state optimization method and device of server and medium
By combining sliding time windows with models, the server operating status is dynamically evaluated, solving the problem that static thresholds cannot adapt to load fluctuations, achieving efficient resource scheduling and load prediction, and improving the server's operating stability and resource utilization efficiency.
Patent Information
- Application Number
- CN202511172323.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Traditional server operation status optimization methods rely on static thresholds and cannot dynamically match the fluctuating characteristics of server load, resulting in low anomaly detection accuracy and high false positive and false negative rates.
By obtaining the server's operating parameter data, using a sliding time window to extract feature vectors, and combining the isolation forest anomaly detection model and the LSTM prediction model, the server's operating status is dynamically evaluated, and future parameter prediction values and dynamic thresholds are constructed to achieve dynamic monitoring of server load and resource scheduling.
It reduces the false alarm rate and missed alarm rate, identifies performance degradation or failure risks in advance, avoids service interruption or performance degradation, and improves resource utilization efficiency.
Smart Images

Figure CN120704993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of server optimization, and in particular to a method, device and medium for optimizing the operating status of a server. Background Art
[0002] As the core node for data processing and business carrying, the server's operational stability and resource utilization efficiency directly affect the service quality of the entire information system. Currently, the business load faced by servers is becoming increasingly complex, placing stringent demands on the server's real-time response capabilities and continuous operation capabilities. Traditional server operation status optimization methods mostly rely on static threshold monitoring, that is, pre-setting fixed thresholds for parameters such as CPU usage, and triggering alarms or resource scheduling when real-time parameters exceed the threshold. However, this approach has significant limitations: server load has dynamic characteristics, and fixed thresholds are difficult to adapt to load fluctuations in different business scenarios, which can easily lead to false alarms (such as short-term peaks triggering unnecessary scheduling) or missed alarms (such as the load slowly climbing but not reaching the threshold but approaching the critical point of failure). Summary of the Invention
[0003] The purpose of the present invention is to provide a method, device and medium for optimizing the operating status of a server to solve the problem that traditional static thresholds cannot dynamically match the fluctuating characteristics of server load, resulting in low anomaly detection accuracy and high false alarm and missed alarm rates.
[0004] According to a first aspect of the present invention, a method for optimizing the operating status of a server is provided, the method comprising the following steps: Obtain server operating parameter data; the operating parameters include CPU usage, disk IOPS, disk latency, network packet loss rate and TCP retransmission rate.
[0005] The server's operating parameter data is slid according to a preset sliding time window to obtain a server operation feature vector; the server operation feature vector includes the CPU usage entropy value, disk pressure index, network retransmission gradient, and packet loss rate variation coefficient.
[0006] The server operation feature vector is input into the first model and the second model in parallel, the server operation status abnormal value is obtained through the first model, and the target parameter prediction value of the server in each collection cycle in the future preset time period is obtained through the second model; the target parameter prediction value includes the CPU usage prediction value, the disk IOPS prediction value and the network packet loss rate prediction value.
[0007] The monitoring parameter threshold of each collection cycle of the server in the future preset time period is obtained based on the abnormal value of the server operation status and the target parameter prediction value of each collection cycle of the server in the future preset time period; the monitoring parameter includes CPU usage.
[0008] If the real-time monitoring parameter data of the server in a certain collection period in the future time period exceeds the corresponding monitoring parameter threshold, the preset resource scheduling optimization action is executed.
[0009] According to a second aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for optimizing the operating status of the server when executing the computer program.
[0010] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the above-mentioned method for optimizing the operating status of the server.
[0011] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention extracts the characteristic vectors of multiple parameters (such as CPU usage entropy, disk pressure index, network retransmission gradient, packet loss rate variation coefficient, etc.) through a sliding time window, and combines it with the outliers output by the first model to realize dynamic evaluation of the server operation status, get rid of the limitations of static thresholds, and reduce false alarm rate and missed alarm rate; the future values of key parameters such as CPU usage, disk IOPS, network packet loss rate, etc. are predicted by the second model, and dynamic thresholds are constructed in combination with outliers, comprehensively considering the correlation and future trends of multiple parameters to capture the potential risks of server operation status; based on the target parameter prediction values and dynamic thresholds in the future preset time period, performance degradation or failure risks can be identified in advance, and optimization actions can be transformed from post-response to pre-prevention, effectively avoiding service interruption or performance degradation, and ensuring that resource scheduling actions are triggered when necessary (such as when future parameters exceed dynamic thresholds), thereby reducing resource waste caused by invalid scheduling and improving the overall resource utilization efficiency of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0013] Figure 1 This is a flowchart of a method for optimizing the operating status of a server provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0015] Example 1: According to this embodiment, Figure 1 As shown, a method for optimizing the running state of a server is provided, the method comprising the following steps: S100, obtaining operating parameter data of the server; the operating parameters include CPU usage, disk IOPS, disk latency, network packet loss rate and TCP retransmission rate.
[0016] As a specific implementation method, the following operating parameters are collected periodically (e.g., every 10 seconds) through the server's built-in monitoring tools (such as top, iostat, netstat of the Linux system, or distributed monitoring systems such as Prometheus and Zabbix): CPU usage: refers to the busyness of the server CPU per unit time (range 0-100%), reflecting the usage of computing resources. Disk IOPS: refers to the number of disk input / output operations per second, reflecting the frequency of disk reads and writes (unit: times / second). Disk latency: refers to the average response time for the disk to complete a read and write operation (unit: milliseconds), reflecting the efficiency of disk reads and writes. Network packet loss rate: refers to the proportion of data packets lost in network transmission to the total number of data packets sent (range 0-100%), reflecting the stability of network transmission. TCP retransmission rate: refers to the proportion of data packets that need to be retransmitted due to data packet loss in the TCP protocol to the total number of data packets sent (range 0-100%), which can further reflect the quality of network transmission.
[0017] Based on the above collected parameters, comprehensive and real-time raw data can be provided for subsequent feature extraction and status assessment.
[0018] S200, sliding the server's operating parameter data according to a preset sliding time window to obtain a server operating feature vector; the server operating feature vector includes a CPU usage entropy value, a disk pressure index, a network retransmission gradient, and a packet loss rate variation coefficient.
[0019] In this embodiment, considering that the parameter data of a single cycle is easily affected by instantaneous fluctuations (such as sudden CPU peaks) and cannot reflect the overall operation trend of the server, this embodiment uses a preset sliding time window (the size of the sliding time window and the sliding step are empirical values, optionally, the window size is 5 minutes and the sliding step is 1 minute) to continuously intercept parameter sequences of fixed lengths, which can smooth instantaneous fluctuations and extract more stable feature vectors.
[0020] In this embodiment, entropy is an indicator of data disorder. The higher the CPU usage entropy, the greater the CPU load fluctuation within the window (such as frequent switching between low and high loads), and the more unstable the server operation. As a specific embodiment, the probability distribution of CPU usage within the window is statistically analyzed, and then the CPU usage entropy is calculated using the entropy formula. Compared with the simple average CPU usage, the entropy can more comprehensively reflect the stability of the CPU load. Those skilled in the art will appreciate that the entropy formula is prior art and will not be repeated here.
[0021] In this embodiment, disk performance is affected by both read and write frequency (IOPS) and response speed (latency); as a specific implementation, the disk pressure index is p, p=(b1 / b max )×(a1 / a0), b1 is the average disk IOPS in the current window, b max is the maximum IOPS of the disk, a1 is the average disk latency in the current window, a0 is the preset disk latency threshold, a0>0. max >0, b1 reflects the current read and write frequency; b max This value reflects the disk's maximum capacity and is preset by the manufacturer. a1 reflects the current response speed, and a0 is the upper limit of normal operating latency, for example, 50ms. A larger p value indicates that the disk experiences more severe latency violations during high-frequency read and write operations, indicating increased pressure.
[0022] In this embodiment, the packet loss rate variation coefficient can measure the degree of network fluctuation. As a specific implementation, the packet loss rate variation coefficient is L. l 2 is less than or equal to the preset packet loss rate threshold, L=0; when l 2 is greater than the preset packet loss rate threshold, L= l 1 / l 2; l 1 is the standard deviation of the packet loss rate in the current window, l 2 is the average packet loss rate within the current window. A larger L indicates more drastic fluctuations in the packet loss rate and more unstable network transmission. Optionally, the preset packet loss rate threshold is an empirical value greater than 0, such as 0.5%.
[0023] In this embodiment, the changing trend of the TCP retransmission rate (such as whether it has increased rapidly recently) can better predict the risk of network deterioration than the current value. As a specific implementation method, the network retransmission gradient is d, d=∑ n i=0 (e -λ×i / (∑ n i=0 e -λ×i )×(R0-R i )) / max(R0,r), R0 is the TCP retransmission rate of the current cycle, R i is the TCP retransmission rate of the previous i cycles of the current cycle, r is the preset minimum TCP retransmission rate threshold, λ is the preset attenuation coefficient, i ranges from 1 to n, and n is the preset number of cycles. Among them, n is an empirical value, for example, n=5, e -λ×i is the attenuation factor, which makes the recent data (i is small) have a higher weight; r is the preset value, r>0, for example, 0.1%; the numerator ∑ n i=0 (e -λ×i / (∑ n i=0 e -λ×i )×(R0-R i )) is the weighted difference between the current and historical retransmission rates, reflecting the changing trend of the retransmission rate. The denominator, max(R0,r), is used for normalization. A positive value of d indicates a rapid increase in the TCP retransmission rate and deteriorating network quality. A negative value of d indicates a decrease in the TCP retransmission rate and improving network quality.
[0024] S300, input the server operation feature vector into the first model and the second model in parallel, obtain the server operation status abnormal value through the first model, and obtain the target parameter prediction value of the server in each collection cycle in the future preset time period through the second model; the target parameter prediction value includes the CPU usage prediction value, the disk IOPS prediction value and the network packet loss rate prediction value.
[0025] In this embodiment, the evaluation of whether the current server operation status is abnormal is realized based on the first model. As a specific implementation method, the first model is an isolation forest anomaly detection model, and the server operation status abnormality value is y, y=2 -e / c, e is the average path length of the server running feature vector in q trees, q is the number of abnormal trees constructed by the isolation forest anomaly detection model, and c is the correction factor of the sample size based on the isolation forest anomaly detection model. Wherein, c>0. The larger the value of y, the more abnormal the server running state is; optionally, q=100. Those skilled in the art are aware that the training process of the isolation forest is a prior art, wherein, during the splitting process, the isolation forest algorithm randomly selects a feature from the multidimensional features and randomly selects a segmentation value of the feature, which will not be repeated here. This embodiment is based on the isolation forest to automatically learn the normal mode to achieve anomaly detection with a lower false alarm rate.
[0026] In this embodiment, an assessment of the future operating status trend of the server is implemented based on the second model. Server parameters (such as CPU utilization, disk IOPS, and network packet loss rate) have time series characteristics. LSTM (Long Short-Term Memory Network) can capture long-term dependencies and is suitable for predicting the future operating status trend of the server. As a specific embodiment, the second model is an LSTM prediction model. The input of the second model also includes a target parameter sequence of the most recent first time period; the target parameter sequence includes a CPU utilization sequence, a disk IOPS sequence, and a network packet loss rate sequence; the second model is used to obtain the target parameter prediction value of the server in a preset time period in the future based on the input server operation feature vector and the target parameter sequence of the most recent first time period.
[0027] As a specific implementation method, the input of the LSTM prediction model includes two types of data: one is the server operation feature vector extracted through a sliding time window (including CPU usage entropy, disk pressure index, network retransmission gradient, and packet loss rate variation coefficient); the other is the target parameter sequence (including CPU usage sequence, disk IOPS sequence, and network packet loss rate sequence) collected at a preset sampling frequency (for example, collected every 10 seconds) within the most recent first time period (which can be an empirical value, such as 10 minutes). The input data needs to be Min-Max normalized and processed into a tensor form of [number of samples, time step (if the first time period is 10 minutes and collected every 10 seconds, then the time step is 60), feature dimension (7 dimensions, i.e., 4-dimensional feature vector + 3-dimensional parameter sequence)]. The LSTM prediction model uses a multi-layer LSTM stacking design. Specifically, the input layer is followed by two LSTM layers (the first layer has 128 hidden units and returns a sequence, and the second layer has 64 hidden units and returns a sequence). A dropout layer (dropout rate 0.2) is inserted in between to prevent overfitting. The LSTM layer output passes through an attention layer (dynamically assigning weights to each time step) before being connected to a fully connected layer (32 units, ReLU activation). The output is finally output through an output layer (3 units, linear activation). The model output is the predicted value of the target parameter for a preset time period in the future (for example, for the next 5 minutes, a 30-time-step prediction sequence is generated for CPU utilization, disk IOPS, and network packet loss rate, each corresponding to 10 seconds. The output is denormalized to restore the actual physical value). During training, historical operational data was used to divide the training, validation, and test sets into a ratio of 7:1.5:1.5. A weighted mean squared error (MSE) was used as the loss function (weights for CPU usage, disk IOPS, and network packet loss were set to 0.5, 0.3, and 0.2, respectively). The Adam optimizer (with an initial learning rate of 0.001 and decay) was used, and training was iteratively performed with a batch size of 64. An early stopping mechanism (stopping if the validation set loss did not decrease after five consecutive rounds) was used to determine the optimal model. Retraining was performed regularly (e.g., weekly) using newly collected data to adapt to server state changes. Furthermore, linear interpolation was used for missing values during training, and outliers exceeding 3σ were smoothed to ensure model robustness. This design enabled the LSTM prediction model to fully capture the temporal dependencies of server parameters and accurately predict parameter trends for a preset time period, providing a reliable basis for future state predictions for subsequent dynamic threshold calculation and resource scheduling optimization.
[0028] S400, obtaining a monitoring parameter threshold of each collection cycle of the server in a future preset time period according to the abnormal value of the server operation status and the target parameter prediction value of each collection cycle of the server in a future preset time period; the monitoring parameter includes CPU usage.
[0029] As a specific implementation method, the CPU usage threshold of a certain collection cycle in a preset time period in the future is h, h=min(h max ,h min ×(1+α×k)), h min is the lower limit of the preset CPU usage threshold, min() is the minimum value, h max is the upper limit of the preset CPU usage threshold, k is the relative change rate, k=max(0,β1×(f1-u1) / (u1+ε)+β2×(f2-b1) / (b1+ε)+β3×(f3-u2) / (u2+ε)), β1, β2 and β3 are the weights corresponding to CPU usage, disk IOPS and network packet loss rate respectively, max( ) is the maximum value, β1, β2 and β3 are all greater than 0 and less than 1, f1, f2 and f3 are respectively the predicted values of CPU usage, disk IOPS and network packet loss rate of the collection period in the future preset time period, u1, b1 and u2 are respectively the CPU usage of the current window, the disk IOPS of the current window and the network packet loss rate of the current window, u1≠0, b1≠0, u2≠0, ε is a preset constant, ε>0, α is the adjustment coefficient, α=α0×(1+η×tanh(g×y)), y is the abnormal value of the server operation status, η is the preset coupling coefficient, η>0, α0 is the baseline coefficient, g is the preset amplification coefficient, g>1, tanh( ) is the hyperbolic tangent function. Optional, h min 、h max , β1, β2, β3, η, α0 and g are empirical values, and ε is a preset minimum value to prevent the denominator from being 0, for example, h min =50%,h max =85%, β1=0.5, β2=0.3, β3=0.2, h 1,0 =50%, α0=0.2, η=0.4, g=2.
[0030] Among them, h min Ensure that the server always retains basic resource redundancy (such as burst requests and process scheduling) to avoid excessive relaxation of thresholds under long-term low load, which may lead to a decrease in risk resistance; maxThe maximum increase in the threshold is limited to prevent it from approaching 100% without limit due to prediction errors or anomaly amplification (a CPU overload can cause process blocking and dramatically increased response delays). The weighting relationship is: β1 > β2 > β3 (sum = 1). This adapts to the characteristics of general servers, where computing is the primary load, followed by storage and network. CPU utilization is the core metric that directly reflects computing resources and has the highest weight (0.5). Disk IOPS (such as database reads and writes and log writes) has the second-highest indirect impact on CPU (0.3, which can be increased to 0.4-0.5 in I / O-intensive scenarios). Network packet loss rate, which primarily consumes CPU indirectly through TCP retransmissions, has the lowest weight (0.2, which can be increased to 0.3 in network-intensive scenarios). α0 controls the basic adjustment range of the threshold under normal conditions to prevent excessive fluctuations in the absence of anomalies. η controls the influence of the outlier value y on the adjustment coefficient α to prevent over-amplification of anomaly signals. g enhances the nonlinear response sensitivity of the outlier value y, ensuring slow adjustment for small anomalies and rapid adjustment for large anomalies.
[0031] Therefore, this embodiment will h min ~h max As a safety interval of the threshold, it can ensure resource utilization (avoid h min Too high will cause waste), and prevent overload (such as h max ≤85%); load change trends are determined by combining predicted CPU usage, disk IOPS, and network packet loss rates. Weight allocation focuses on computation and multi-resource linkage to adapt to the load characteristics of most servers. Abnormal response uses g and η to achieve gradient perception, avoiding oversensitivity to small abnormalities and slow response to large abnormalities. Safety threshold processing prevents mathematical calculation anomalies and ensures that the formula still runs stably in extreme scenarios (such as low load and zero packet loss).
[0032] S500: If the real-time monitoring parameter data of a certain collection period of the server in a future time period exceeds the corresponding monitoring parameter threshold, a preset resource scheduling optimization action is executed.
[0033] In this embodiment, S500 includes: if the CPU usage data of the server in a certain collection period in a future time period exceeds the corresponding CPU usage threshold, then executing a preset resource scheduling optimization action; wherein the corresponding CPU usage threshold refers to the CPU usage threshold of the same collection period. For example, if the CPU usage data of the server in the first collection period in the future time period exceeds the CPU usage threshold of the server in the first collection period in the future time period, then executing the preset resource scheduling optimization action. For another example, if the CPU usage data of the server in the fifth collection period in the future time period exceeds the CPU usage threshold of the server in the fifth collection period in the future time period, then executing the preset resource scheduling optimization action.
[0034] As a specific embodiment, the preset resource scheduling optimization actions include at least one of the following resource scheduling optimization actions: load balancing and resource expansion. Load balancing involves migrating some tasks to idle servers to reduce CPU pressure on the current server; resource expansion involves temporarily increasing the number of CPU cores to improve server processing power. Those skilled in the art will appreciate that the process of executing resource scheduling optimization actions is well known in the art and will not be further described here.
[0035] This embodiment extracts feature vectors of multiple parameters (such as CPU usage entropy, disk pressure index, network retransmission gradient, packet loss rate variation coefficient, etc.) through a sliding time window, and combines them with the outliers output by the first model to achieve dynamic assessment of the server operating status, breaking away from the limitations of static thresholds and reducing false alarm and missed alarm rates. The second model predicts the future values of key parameters such as CPU usage, disk IOPS, and network packet loss rate, and constructs dynamic thresholds based on the outliers. This comprehensively considers the correlation and future trends of multiple parameters to capture potential risks in the server operating status. Based on the predicted values of the target parameters and dynamic thresholds for a preset time period in the future, performance degradation or failure risks can be identified in advance, and optimization actions can be transformed from post-response to pre-prevention, effectively avoiding service interruption or performance degradation. It can also ensure that resource scheduling actions are triggered when necessary (such as when future parameters exceed dynamic thresholds). This can reduce resource waste caused by ineffective scheduling and improve the overall resource utilization efficiency of the server cluster.
[0036] Example 2: This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed: Obtain server operating parameter data; the operating parameters include CPU usage, disk IOPS, disk latency, network packet loss rate and TCP retransmission rate.
[0037] The server's operating parameter data is slid according to a preset sliding time window to obtain a server operation feature vector; the server operation feature vector includes the CPU usage entropy value, disk pressure index, network retransmission gradient, and packet loss rate variation coefficient.
[0038] The server operation feature vector is input into the first model and the second model in parallel, the server operation status abnormal value is obtained through the first model, and the target parameter prediction value of the server in each collection cycle in the future preset time period is obtained through the second model; the target parameter prediction value includes the CPU usage prediction value, the disk IOPS prediction value and the network packet loss rate prediction value.
[0039] The monitoring parameter threshold of each collection cycle of the server in the future preset time period is obtained based on the abnormal value of the server operation status and the target parameter prediction value of each collection cycle of the server in the future preset time period; the monitoring parameter includes CPU usage.
[0040] If the real-time monitoring parameter data of the server in a certain collection period in the future time period exceeds the corresponding monitoring parameter threshold, the preset resource scheduling optimization action is executed.
[0041] Example 3: This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the following steps are implemented: Obtain server operating parameter data; the operating parameters include CPU usage, disk IOPS, disk latency, network packet loss rate and TCP retransmission rate.
[0042] The server's operating parameter data is slid according to a preset sliding time window to obtain a server operation feature vector; the server operation feature vector includes the CPU usage entropy value, disk pressure index, network retransmission gradient, and packet loss rate variation coefficient.
[0043] The server operation feature vector is input into the first model and the second model in parallel, the server operation status abnormal value is obtained through the first model, and the target parameter prediction value of the server in each collection cycle in the future preset time period is obtained through the second model; the target parameter prediction value includes the CPU usage prediction value, the disk IOPS prediction value and the network packet loss rate prediction value.
[0044] The monitoring parameter threshold of each collection cycle of the server in the future preset time period is obtained based on the abnormal value of the server operation status and the target parameter prediction value of each collection cycle of the server in the future preset time period; the monitoring parameter includes CPU usage.
[0045] If the real-time monitoring parameter data of the server in a certain collection period in the future time period exceeds the corresponding monitoring parameter threshold, the preset resource scheduling optimization action is executed.
[0046] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0047] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for optimizing the operating status of a server, characterized in that: The method comprises the following steps: Obtaining server operating parameter data; the operating parameters include CPU usage, disk IOPS, disk latency, network packet loss rate, and TCP retransmission rate; Sliding the server's operating parameter data according to a preset sliding time window to obtain a server operating feature vector; the server operating feature vector includes the CPU usage entropy value, disk pressure index, network retransmission gradient, and packet loss rate variation coefficient; Inputting the server operation feature vector into the first model and the second model in parallel, obtaining the server operation status abnormal value through the first model, and obtaining the target parameter prediction value of the server in each collection cycle in a future preset time period through the second model; the target parameter prediction value includes the CPU usage prediction value, the disk IOPS prediction value, and the network packet loss rate prediction value; Obtaining a monitoring parameter threshold value of each collection period of the server in a future preset time period based on the abnormal value of the server operation status and the target parameter prediction value of each collection period of the server in a future preset time period; the monitoring parameter includes CPU usage; If the real-time monitoring parameter data of the server in a certain collection period in the future time period exceeds the corresponding monitoring parameter threshold, the preset resource scheduling optimization action is executed.
2. The method for optimizing the operating status of a server according to claim 1, wherein: The CPU usage threshold of a certain collection cycle in the future preset time period is h, h=min(h max ,h min ×(1+α×k)), h min is the lower limit of the preset CPU usage threshold, min() is the minimum value, h max is the upper limit of the preset CPU usage threshold, k is the relative change rate, k=max(0,β1×(f1-u1) / (u1+ε)+β2×(f2-b1) / (b1+ε)+β3×(f3-u2) / (u2+ε)), β1, β2 and β3 are the weights corresponding to CPU usage, disk IOPS and network packet loss rate respectively, max( ) is the maximum value, β1, β2 and β3 are all greater than 0 and less than 1, f1, f2 and f3 are respectively the predicted values of CPU usage, disk IOPS and network packet loss rate of the collection period in the future preset time period, u1, b1 and u2 are respectively the CPU usage, disk IOPS and network packet loss rate of the current window, ε is the preset constant, ε>0, α is the adjustment coefficient, α=α0×(1+η×tanh(g×y)), y is the abnormal value of the server operation status, η is the preset coupling coefficient, η>0, α0 is the baseline coefficient, g is the preset amplification coefficient, g>1, and tanh( ) is the hyperbolic tangent function.
3. The method for optimizing the operating status of a server according to claim 1, wherein: The disk pressure index is p, p=(b1 / b max )×(a1 / a0), b1 is the average disk IOPS in the current window, b max is the maximum disk IOPS, a1 is the average disk latency in the current window, a0 is the preset disk latency threshold, and a0>0.
4. The method for optimizing the operating status of a server according to claim 1, wherein: The packet loss rate variation coefficient is L, when l 2 is less than or equal to the preset packet loss rate threshold, L=0; when l 2 is greater than the preset packet loss rate threshold, L= l 1 / l 2; l 1 is the standard deviation of the packet loss rate in the current window, l 2 is the average packet loss rate in the current window.
5. The method for optimizing the running state of a server according to claim 1, wherein: The network retransmission gradient is d, d=∑ n i=0 (e -λ×i / (∑ n i=0 e -λ×i )×(R0-R i )) / max(R0,r), R0 is the TCP retransmission rate of the current cycle, R i is the TCP retransmission rate of the i cycles before the current cycle, r is the preset minimum TCP retransmission rate threshold, λ is the preset attenuation coefficient, i ranges from 1 to n, and n is the preset number of cycles.
6. The method for optimizing the operating status of a server according to claim 1, wherein: The first model is the Isolation Forest Anomaly Detection Model, where the server operation status anomaly value is y, y=2 -e / c , e is the average path length of the server running feature vector in q trees, q is the number of anomaly trees constructed by the isolation forest anomaly detection model, and c is the correction factor of the sample size based on the isolation forest anomaly detection model.
7. The method for optimizing the running state of a server according to claim 1, wherein: The second model is an LSTM prediction model. The input of the second model also includes the target parameter sequence of the most recent first period; the target parameter sequence includes the CPU usage sequence, the disk IOPS sequence and the network packet loss rate sequence; the second model is used to obtain the target parameter prediction value of the server in the future preset time period based on the input server operation feature vector and the target parameter sequence of the most recent first period.
8. The method for optimizing the running state of a server according to claim 1, wherein: The preset resource scheduling optimization action includes at least one of the following resource scheduling optimization actions: load balancing and resource expansion.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the server operation status optimization method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for optimizing the operating status of a server according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
DCS system autonomous controllable upgrading and system fault prediction method and system
CN119536177A
Server load state evaluation method based on dynamic evaluation algorithm
CN119902905A
Machine room automation equipment abnormal state monitoring method and device
CN120378329A
System for detecting malicious nodes in a wireless sensor network and a method thereof
US20250234201A1
Service abnormality prediction method and device, storage medium, and electronic device
WO2023045829A1