Data monitoring method of a server and electronic device

By employing dynamic reference data and fault prediction models in enterprise-level data centers, server resources are dynamically adjusted, solving the problems of inaccurate fault identification and resource waste under static rules, and achieving more efficient operation and maintenance management.

CN120849222BActive Publication Date: 2025-11-28LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511350294.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-28
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

In existing technologies, the operation and maintenance management of enterprise-level data centers relies on manual on-duty personnel and static rules, which leads to inaccurate fault identification, frequent false alarms affecting business operations, and unreasonable resource allocation, resulting in resource waste.

Method used

By employing dynamic reference data and fault prediction models, server hardware performance data is collected, historical data from the resource pool is used to determine current reference data, and analysis is conducted in conjunction with the fault prediction model to dynamically adjust server resources.

Benefits of technology

It improved the accuracy of server indicator data identification, ensured the normal operation of business, improved resource utilization, and reduced false alarms and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849222B_ABST
    Figure CN120849222B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data supervision method of server and electronic equipment, it is related to server technical field, including the acquisition various current hardware index data in server;For each current hardware index data, using the current reference data corresponding to current hardware index data to analyze current hardware index data, determine whether current hardware index data is abnormal index data;Current reference data is determined based on each corresponding historical index data stored in the history data of resource pool;In the case where it is determined that current hardware index data is abnormal index data, using the fault prediction model established in advance to analyze current hardware index data, obtain corresponding prediction index data;Server resource is regulated based on prediction index data.It solves the technical problem of low accuracy of abnormal index identification and low resource utilization, achieves the technical effect of improving index data identification accuracy and resource utilization, conducive to normal operation of business.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and particularly relates to a data monitoring method of a server and an electronic device. BACKGROUND

[0002] In related technologies, when performing operation and maintenance management on an enterprise-level data center, a mode of combining manual on-duty with static rules is mainly used for operation and maintenance management, fault recognition is performed based on a static threshold, and an alarm system is triggered, but this mode cannot accurately recognize abnormal indicators, leading to frequent false positives of key indicators during a business peak period, and affecting normal operation of the business. In addition, resources in related technologies are allocated by using fixed quotas, leading to insufficient utilization of resources.

[0003] Therefore, how to improve the accuracy of indicator data recognition of a server, guarantee normal operation of a business, and improve resource utilization has become a problem to be solved by those skilled in the art. SUMMARY

[0004] The present application provides a data monitoring method of a server and an electronic device, which can improve the accuracy of indicator data recognition of a server, better guarantee normal operation of a business, and improve resource utilization.

[0005] The present application provides a data monitoring method of a server, comprising:

[0006] collecting various types of current hardware indicator data in the server;

[0007] For each type of current hardware indicator data, current reference data corresponding to the current hardware indicator data is used to analyze the current hardware indicator data, to determine whether the current hardware indicator data is abnormal indicator data; the current reference data is determined based on each corresponding historical indicator data stored in historical data of a resource pool;

[0008] In a case where it is determined that the current hardware indicator data is abnormal indicator data, a pre-established fault prediction model is used to analyze the current hardware indicator data, to obtain corresponding predicted indicator data;

[0009] Based on the predicted indicator data, the server resources are regulated.

[0010] The present application also provides a data monitoring device of a server, comprising:

[0011] a collection module, configured to collect various types of current hardware indicator data in the server;

[0012] The first analysis module is configured to analyze the current hardware indicator data by using the current reference data corresponding to the current hardware indicator data for each type of current hardware indicator data, and determine whether the current hardware indicator data is abnormal indicator data; the current reference data is determined based on each corresponding historical indicator data stored in the historical data of the resource pool;

[0013] The second analysis module is configured to analyze the current hardware indicator data by using the pre-established fault prediction model in a case where it is determined that the current hardware indicator data is abnormal indicator data, and obtain corresponding prediction indicator data;

[0014] The regulation module is configured to regulate the server resources based on the prediction indicator data.

[0015] The application further provides an electronic device, including a memory configured to store a computer program, and a processor configured to execute the computer program to implement the steps of the data monitoring method of the server.

[0016] The application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the data monitoring method of the server.

[0017] The application further provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps of the data monitoring method of the server.

[0018] From the above technical solution, the beneficial effects of the application are as follows:

[0019] In the data monitoring method of the server, the current reference data for identifying whether the current hardware indicator data is abnormal indicator data is obtained according to each corresponding historical indicator data in the historical data of the resource pool, that is, the reference data in the application is obtained according to all current historical data, and is not a fixed value. Therefore, after collecting each type of current hardware indicator data in the server, the current hardware indicator data is analyzed according to the current reference data corresponding to the type, so that whether the current hardware indicator data is abnormal indicator data can be more accurately determined. If the current hardware indicator data is abnormal indicator data, the current hardware indicator data is further analyzed by using the pre-established fault analysis model to obtain corresponding prediction indicator data, and the server resources are further regulated according to the prediction indicator data, so that the accuracy of the indicator data identification of the server can be improved, the normal operation of the business can be better ensured, and the utilization rate of the server resources can be improved.

[0020] In addition, the server data supervision method also provides corresponding implementation devices, electronic equipment and computer readable storage media, which further make the method more practical, and the devices, electronic equipment and computer readable storage media have corresponding advantages. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 A flow chart of a server data supervision method provided by the embodiments of the present application;

[0023] Figure 2 An architecture diagram of a server data supervision system provided by the embodiments of the present application;

[0024] Figure 3 A structural diagram of a server data supervision device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0026] It should be noted that, in the description of the present application, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.

[0027] It should be noted that the current enterprise-level data center generally adopts a combination of manual duty and static rules for operation and maintenance management mode, and when dealing with the dynamic complexity of modern cloud native architecture, the following problems may exist.

[0028] First, alarm systems based on static threshold triggers suffer from significant response delays. Taking the log analysis of a large e-commerce platform as an example, the average time from the appearance of signs of disk I / O anomalies to triggering an alarm is 127 seconds, during which time cascading failures may have already occurred. This demonstrates that static thresholds are ill-suited to the cyclical fluctuations in business load, leading to frequent false alarms for key indicators during peak business periods and impacting normal business operations.

[0029] Secondly, traditional resource allocation uses a fixed quota model. For example, data from a financial institution's test environment showed that the average CPU utilization rate was less than 12% during non-production periods, but the memory reservation was as high as 85% of the configuration limit, resulting in an annual waste of more than $2 million in resources.

[0030] Therefore, this application provides a server data monitoring method that can improve the accuracy of indicator data identification and resource utilization. To enable those skilled in the art to better understand this application, the following detailed description is provided in conjunction with the accompanying drawings and specific embodiments.

[0031] The specific application environment architecture or specific hardware architecture on which the execution of the server's data monitoring method depends is described here.

[0032] Embodiments of this application provide a data monitoring method for a server, combined with, for example... Figure 1 The flowchart shown illustrates the execution of the server's data monitoring method, providing a detailed description of the method. The method includes the following steps, S110 to S140.

[0033] S110: Collects various current hardware indicator data from the server.

[0034] It should be noted that in this embodiment of the application, various current hardware indicator data of the server can be collected through the BMC (Baseboard Management Controller). For example, the current hardware indicator data may include current CPU data, current power consumption, current fan speed, current throughput, voltage parameters, etc.

[0035] S120: For each type of current hardware indicator data, the current hardware indicator data is analyzed using the current reference data corresponding to the current hardware indicator data to determine whether the current hardware indicator data is abnormal indicator data; the current reference data is determined based on the corresponding historical indicator data stored in the historical data of the resource pool.

[0036] In actual application, each type of hardware index data collected each time can be stored into the historical data of the resource pool. For each type of hardware index data, the corresponding current reference data can be determined according to each historical hardware index data corresponding to the type and stored in the historical data of the resource pool. When identifying each type of current hardware index data, the current hardware index data can be analyzed according to the current reference data corresponding to the current hardware index data, so as to determine whether the current hardware index data is abnormal index data. Since the current reference data in the present application is obtained according to all historical hardware index data corresponding to the type in the historical data of the resource pool, the current reference data is dynamically changed, and thus the current hardware index data can be more accurately determined whether it is abnormal index data.

[0037] S130: in the case where it is determined that the current hardware index data is abnormal index data, the pre-established fault prediction model is used to analyze the current hardware index data to obtain corresponding prediction index data.

[0038] It can be understood that in the case where it is determined that the current hardware index data is abnormal index data, the pre-established fault prediction model can be used to analyze the current hardware index data, and the prediction index data corresponding to the current hardware index data can be obtained.

[0039] In actual application, the steps of S120 and S130 can be realized by the resource decision engine.

[0040] S140: based on the prediction index data, the server resource is regulated.

[0041] After obtaining the prediction index data corresponding to the current hardware index data, the BMC can further regulate the resource of the server according to the prediction index data, so as to realize the dynamic allocation of the server resource, which is beneficial to improve the resource utilization rate of the server.

[0042] Of course, in actual application, the related hardware of the server can also be regulated according to the prediction index data, so as to restore the abnormal index data to normal. For example, the BMC is used to regulate the related hardware of the server according to the prediction index data.

[0043] On the basis of the above embodiment, the technical solution can be described in detail with reference to the architecture diagram shown in Figure 2

[0044] In an implementation manner, the method can further include:

[0045] Each type of current hardware index data except the abnormal index data is stored into the historical data of the resource pool. ​

[0046] It should be noted that in actual application, after collecting various types of current hardware indicator data, the various types of current hardware indicator data can be cached in the resource pool. After analyzing the various types of current hardware indicator data and determining the abnormal indicator data, the abnormal indicator data is deleted, and the remaining each other current hardware indicator data is stored in the historical data of the resource pool to constantly improve the historical data, so as to dynamically adjust the current reference data to improve the recognition accuracy. For example, after the resource decision engine analyzes the various types of current hardware indicator data and determines the abnormal indicator data, the resource decision engine can notify the resource pool which is the abnormal indicator data. After receiving the notification, the resource pool can store the other current hardware indicator data except the abnormal indicator data in the historical data of the resource pool.

[0047] In an embodiment, after storing the other current hardware indicator data except the abnormal indicator data in the historical data of the resource pool, the method can further include:

[0048] Assigning a corresponding weight value to each other current hardware indicator data stored in the historical data of the resource pool.

[0049] In the embodiment of the application, after storing each other current hardware indicator data in the historical data of the resource pool, a corresponding weight value can be assigned to each newly stored other current hardware indicator data. Wherein the corresponding weight values of each other current hardware indicator data are added to 1, and for each type of other current hardware indicator data, according to the weight values corresponding to each historical hardware indicator data corresponding to the type in the historical data, a weight value corresponding to the other current hardware indicator data of the type is determined according to a preset rule. The preset rule can be that all historical data is divided into n groups, the weight value corresponding to the group with the earliest timestamp of the same type of historical hardware indicator data is the smallest, the weight value corresponding to the group with the latest timestamp is the largest, and the weight values of each group are added to 1. For the newly added other current hardware indicator data, the last group is determined, and the weight value of the last group is taken as the weight value of the other current hardware indicator data. If the number of historical hardware indicator data in the last group reaches a preset number, the newly added other current hardware indicator data can be taken as the first historical hardware indicator data of a new group, and the weight values of each group are re-divided to meet the equal difference series from the first group with the earliest timestamp to the current group with the latest timestamp, and the sum of the weight values of each group is 1.

[0050] In the embodiments of the present application, when the current reference data is calculated, the current reference data can be calculated according to the historical hardware indicator data of the category and the weight value corresponding to each historical hardware indicator data, so as to improve the calculation accuracy of the current reference data, thereby improving the recognition accuracy of the hardware indicator data.

[0051] In an embodiment, the current reference data in the embodiments of the present application can include a current indicator empirical value and a current indicator variance.

[0052] Therefore, the process of analyzing the current hardware indicator data by using the current reference data corresponding to the current hardware indicator data in S120 to determine whether the current hardware indicator data is abnormal indicator data can include:

[0053] comparing the value of the current hardware indicator data with the current indicator empirical value to obtain a current indicator difference value;

[0054] comparing the current indicator difference value with the current indicator variance to obtain a current indicator variance difference value;

[0055] determining whether the current indicator variance difference value is greater than a preset multiple of the current indicator variance;

[0056] in the case that the multiple is greater than the preset multiple, determining that the current hardware indicator data is abnormal indicator data;

[0057] in the case that the multiple is less than or equal to the preset multiple, determining that the current hardware indicator data is normal indicator data.

[0058] It should be noted that in the embodiments of the present application, after obtaining each type of current hardware indicator data, the current indicator empirical value and the current indicator variance difference value corresponding to the current hardware indicator data can be obtained for each type of current hardware indicator data. The current indicator empirical value and the current indicator variance are obtained by calculating (for example, weighted average calculation) according to the historical indicator data corresponding to each historical indicator data of the category and the corresponding weight value. The value of the current hardware indicator data can be compared with the current indicator empirical value to obtain a current indicator difference value, and the current indicator difference value is further compared with the current indicator variance to obtain a current indicator variance difference value. If the current indicator variance difference value is greater than the current indicator variance value by a preset multiple (for example, 5 times), it indicates that the current hardware indicator data is abnormal indicator data, otherwise, it is normal indicator data.

[0059] Taking temperature as an example, for the current temperature data, the current temperature data can be compared with the current temperature experience value to obtain a current temperature difference value, and then the current temperature difference value is further compared with the current temperature variance to obtain a current temperature variance difference value. If the current temperature variance difference value is greater than the current temperature variance value by a preset multiple (for example, 5 times), it indicates that the current temperature index data is abnormal temperature data, otherwise, it is normal temperature data. Other types of hardware index data are identified according to this method, which will not be illustrated one by one.

[0060] In an embodiment, the pre-established fault prediction model is used to analyze the current hardware index data in S130 to obtain corresponding prediction index data, including:

[0061] Obtaining the first preset number of abnormal index data of the same type as the current hardware index data;

[0062] Using the pre-established fault prediction model to perform linear regression calculation on the current hardware index data and the first preset number of abnormal index data to obtain a current index calculation result;

[0063] Determining the hardware index change trend according to the current index calculation result;

[0064] In the case where the corresponding hardware index deviates from the current index experience value according to the hardware index change trend, predicting the third preset number of prediction index data in the future according to the current index calculation result and the second preset number of index calculation results.

[0065] It should be noted that in the embodiments of the present application, the fault prediction model can be triggered after the resource decision engine identifies the abnormal index data. The fault prediction model can perform linear regression calculation on the current hardware index data of the abnormal index data and the first preset number of each abnormal index data of the same type previously determined to obtain the current index calculation result. That is, the fault prediction model can perform linear regression calculation on the current hardware index data and the first preset number of abnormal index data of the same type to determine the current index calculation result each time an abnormal current hardware index data is obtained. Then, according to the current index calculation result and the previous multiple current index calculation results, the hardware index change trend is determined. If the hardware index change trend is more and more deviating from the current index experience value, it indicates that the current hardware index data has no numbered trend, but only a deteriorating trend. At this time, the third preset number (3 or 5) of prediction index data in the future can be predicted according to the current index calculation result and the second preset number of index calculation results, so that the BMC can control the state of the corresponding hardware and the resource of the server according to each prediction index data.

[0066] In an embodiment, the method can further include:

[0067] For each type of hardware index data, the next reference data is determined according to all historical index data corresponding to the type of hardware index data in the historical data of the resource pool.

[0068] It should be further noted that, in the embodiments of the present application, after the current hardware index data excluding the abnormal index data is stored in the historical data of the resource pool, and each other current hardware index data is assigned a corresponding weight value, the next reference data corresponding to the type can be further calculated according to each historical index data belonging to the same type in the historical data of the resource pool and the corresponding weight value, that is, the next reference experience value and the next reference variance value are calculated, so as to identify the abnormal data based on the next reference data after the next time the hardware index data of the server is obtained, thereby constantly adjusting the reference data, which is beneficial to improve the accuracy of index data identification.

[0069] In an embodiment, the method can further include:

[0070] In the case where the current hardware index data is determined to be abnormal index data and the next reference data is determined, the next reference data is subjected to drift detection callback.

[0071] In actual application, in order to further improve the accuracy of the reference data, the next reference data calculated can be subjected to drift detection callback in the case where the current hardware index data is abnormal index data, so as to adjust the next reference data to be more accurate.

[0072] In an embodiment, the process of subjecting the next reference data to drift detection callback can include:

[0073] According to the service type input by the user in advance, the target reference value is determined in combination with the corresponding relationship between the service type and the reference value established in advance;

[0074] The next reference data is adjusted according to the target reference value.

[0075] It can be understood that, in the embodiments of the present application, the corresponding reference value can be determined in advance according to different service types, and the reference value can include experience value reference value and variance value reference value. In order to further meet the user demand, the user can select the corresponding service demand according to the actual demand, and the system can determine the corresponding target reference value from the corresponding relationship between the service type and the reference value according to the service demand, and then adjust the next reference data to the target reference value, thereby realizing the drift detection callback of the next reference data.

[0076] In an implementation, the process of performing drift detection callback on the next reference data can include:

[0077] The third preset number of index calculation results calculated by linear regression are obtained;

[0078] A first reference value is determined according to the third preset number of index calculation results;

[0079] The adjusted next reference data is determined according to the first reference value, the first proportion, the reference value corresponding to the next reference data, and the second proportion.

[0080] It can be understood that when performing drift detection callback on the next reference data, a first reference value can also be calculated according to the third preset number (for example, 50) of index calculation results calculated by linear regression, and then the first reference value is multiplied by the first proportion (for example, 30%), the reference value of the calculated next reference data is multiplied by the second proportion (for example, 70%), and the two are added to obtain the adjusted next reference data.

[0081] It should be further pointed out that in actual application, after obtaining each prediction index data, expansion and contraction adjustment can be performed according to each prediction index data, for example, adjusting the number of replicas corresponding to each hardware index data, and dividing a preset size of computing resource from the current remaining computing resource for data monitoring and use.

[0082] In an implementation, the fault prediction model in the embodiments of the present application is trained based on historical index data corresponding to each hardware index data, combined with an attention mechanism and a first double-layer long short-term memory network.

[0083] The hardware index data can also include disk health status, service heartbeat interval, etc. In the model training process, 60 time steps and multiple features (such as CPU, memory, network, log error code, service heartbeat, etc.) can be used as input, combined with an attention mechanism and a first double-layer long short-term memory network (Long Short-Term Memory, LSTM) for model training. The code is as follows:

[0084] inputs = Input(shape=(60, 5));

[0085] lstm1 = LSTM(64, return_sequences=True)(inputs);

[0086] lstm2 = LSTM(32, return_sequences=True)(lstm1);

[0087] attention = Attention()([lstm2, lstm2]) # Self-attention mechanism captures key moments;

[0088] output = Dense(1, activation='sigmoid')(attention);

[0089] model = Model(inputs, output).

[0090] In addition, during data collection in this embodiment, for example, the system can collect CPU and memory usage data once per second, recording corresponding metrics (such as CPU utilization percentage and used memory), and can also collect network bandwidth data once per second, recording metrics including receive rate per second and send rate per second; it can also obtain business load information by real-time parsing of access logs and record data such as the number of requests per second and average response time. The implementation code is as follows:

[0091] CPU / memory utilization in Prometheus+;

[0092] NodeExporter1 second / time cpu_usage_percent, memory_used_gb;

[0093] Network bandwidth: iftop+Telegraf 1 second / time network_rx_mbps, network_tx_mbps;

[0094] Business load is measured in Nginx / AccessLog real-time requests_per_second, avg_response_time.

[0095] In one implementation, the fault prediction model is trained based on historical indicator data corresponding to each hardware indicator, combined with an attention mechanism and a first two-layer long short-term memory network, and includes:

[0096] Data augmentation is performed on the historical data corresponding to each hardware indicator, and the model is trained based on the augmented historical data, combined with the attention mechanism and the first two-layer long short-term memory network, and the training process is adjusted based on the dynamic learning rate.

[0097] If the loss value does not decrease in any of the preset number of training rounds, the training ends, and the trained fault prediction model is obtained.

[0098] It should be noted that in actual application, the first bidirectional long and short term memory network can be used to capture the dependence of the front and back. For example:

[0099] model.add(Bidirectional(LSTM(64, return_sequences=True),

[0100] input_shape=(window_size, 4))) # input 4 features;

[0101] model.add(Bidirectional(LSTM(32)))。

[0102] Multiple output regression can also be performed to predict various indicator data, such as predicting CPU, memory, or network, etc. For example:

[0103] model.add(Dense(3));

[0104] model.compile(loss='mae', optimizer='adam') # use mean absolute error.

[0105] During the training process, data augmentation can be performed on the historical indicator data corresponding to each hardware indicator data, such as adding high-speed noise (σ=0.01) to prevent overfitting. ReduceLROnPlateau callback can also be used to automatically adjust the learning rate. Early stopping mechanism is used during the training process, that is, a loss value is obtained for each training round, and if the loss value does not decrease for a predetermined number (for example, 3) of consecutive loss values, the training can be stopped to prevent overfitting. For example, the following code can be used to implement:

[0106]

[0107] In an embodiment, the process of server resource regulation based on predicted indicator data in S140 described above can include:

[0108] Regulating the server resources based on the predicted indicator data through a pre-established resource dynamic allocation model; the resource dynamic allocation model is trained based on the attention mechanism and the second bidirectional long and short term memory network.

[0109] It should be noted that in actual application, the resource dynamic allocation model trained can be integrated in the BMC, and the resource demand in a future predetermined time period (such as 5 minutes) can be predicted through the resource dynamic allocation model, as follows:

[0110] prediction = model.predict(last_hour_data) # Input shape (1, 60, 4);

[0111] # Output format: [cpu%, memory_gb, network_mbps];

[0112] pred_cpu = prediction[0][0] # For example, 78.5%;

[0113] pred_mem = prediction[0][1] # For example, 12.3 GB;

[0114] pred_net = prediction[0][2] # For example, 245 Mbps.

[0115] Based on this output, scaling decisions can be made. For example, if the predicted future CPU utilization is higher than 85%, it indicates that the load will continue to surge. In this case, two compute nodes can be added immediately to alleviate the pressure. If the predicted future CPU utilization is lower than 30%, it indicates that resources are surplus, and one compute node can be reduced to save costs. The implementation code is as follows:

[0116]

[0117] Additionally, regarding memory, if the predicted memory usage `pred_mem` is more than 20% higher than the currently allocated memory `current_mem`, then memory is re-allocated by a total amount equal to "the predicted value plus an additional 20%", thus reserving a 20% safety buffer to prevent frequent expansions due to insufficient memory. The implementation code is as follows:

[0118] if pred_mem > current_mem * 1.2:

[0119] allocate_memory(pred_mem * 1.2) # Reserve 20% buffer.

[0120] For the network, if the predicted network traffic `pred_net` reaches more than a preset multiple (e.g., 3) of the daily baseline value `baseline_net`, it indicates a traffic surge. In this case, `enable_traffic_shaping()` can be called to activate the traffic shaping mechanism, limiting and scheduling ingress or egress bandwidth to prevent sudden traffic from congesting links, causing latency spikes, or packet loss. The implementation code is as follows:

[0121] if pred_net > baseline_net * 3:

[0122] enable_traffic_shaping() # Enable traffic shaping.

[0123] It should also be noted that, in another implementation, scaling up or down can be triggered via a Kubernetes Python client. For example, the `client` and `config` modules can be imported from the `kubernetes` package for cluster operations. A `scale_up` function can also be defined for subsequent scaling up of a specific resource. Internally, the `scale_up` function first calls `config.load_kube_config()` to read the local default kube-config file (`~ / .kube / config`) and obtain connection credentials to the target cluster. Then, it creates an API client instance using `client.AppsV1Api()`. This client instance is then used to adjust the replica count of application-level resources such as Deployments and StatefulSets, thus achieving scaling up. The implementation code is as follows:

[0124] from kubernetes import client, config;

[0125] def scale_up(resource_type, nodes=1):

[0126] config.load_kube_config();

[0127] api = client.AppsV1Api().

[0128] In practical applications, the current number of replicas can be obtained and updated. Obtaining the current number of replicas can be achieved using the following code:

[0129] deployment= api.read_namespaced_deployment("web-server", "default");

[0130] current_replicas = deployment.spec.replicas.

[0131] Updating the replica count can be achieved using the following code:

[0132] deployment.spec.replicas = current_replicas + nodes

[0133] api.patch_namespaced_deployment("web-server","default", deployment).

[0134] In the process of scaling, Prometheus can push real-time collected metrics data to Kafka, and then Kafka provides the latest 60 data points to Predictor. Predictor calls the LSTM model for inference to generate future resource demand prediction. Then, Predictor sends scaling instructions to Kubernetes through an HTTP POST request, and Kubernetes updates the corresponding Deployment replica number after receiving the scaling instruction, and completes the elastic scaling smoothly. The code implementation is as follows:

[0135] Prometheus->>Kafka: Real-time push of metrics data;

[0136] Kafka->>Predictor: Consume the latest 60 data points;

[0137] Predictor->>Predictor: LSTM model inference;

[0138] Predictor->>Predictor: Send scaling instructions (HTTP POST);

[0139] K8s->>K8s: Adjust the number of Deployment replicas.

[0140] It can be seen that, in the data supervision method of the server, the current reference data used to identify whether the current hardware indicator data is abnormal indicator data is obtained according to each historical indicator data corresponding to the type in the historical data of the resource pool, that is, the reference data in the present application is obtained according to all current historical data, and is not a fixed value. Therefore, after collecting each type of current hardware indicator data in the server, for each type of current hardware indicator data, the current hardware indicator data is analyzed according to the current reference data corresponding to the type, which can more accurately determine whether the current hardware indicator data is abnormal indicator data. If it is abnormal indicator data, the current hardware indicator data is analyzed by the pre-established fault analysis model to obtain corresponding predicted indicator data, and the server resources are further regulated according to the predicted indicator data, so that the accuracy of the indicator data identification of the server can be improved, the normal operation of the business can be better guaranteed, and the utilization rate of the server resources can be improved.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.

[0142] The embodiment of the present application also provides a server data supervision device, please refer to Figure 3 The server data supervision device comprises:

[0143] The acquisition module 11 is configured to acquire each type of current hardware indicator data in the server.

[0144] The first analysis module 12 is configured to analyze the current hardware indicator data by using the current reference data corresponding to the current hardware indicator data for each type of current hardware indicator data, and determine whether the current hardware indicator data is abnormal indicator data; the current reference data is determined based on each corresponding historical indicator data stored in the historical data of the resource pool.

[0145] The second analysis module 13 is configured to analyze the current hardware indicator data by using the pre-established fault prediction model in the case that the current hardware indicator data is determined to be abnormal indicator data, and obtain corresponding predicted indicator data.

[0146] The regulation module 14 is configured to regulate the server resources based on the predicted indicator data.

[0147] In an embodiment, the device can further comprise:

[0148] The storage module is configured to store other current hardware indicator data except the abnormal indicator data in each type of current hardware indicator data into the historical data of the resource pool.

[0149] In an embodiment, the apparatus can further comprise:

[0150] an assigning module, configured to assign a corresponding weight value to each of the other current hardware index data stored in the historical data of the resource pool.

[0151] In an embodiment, the current reference data comprises a current index mean value and a current index variance value.

[0152] the first analysis module 12 comprises:

[0153] a first comparison unit, configured to compare the value of the current hardware index data with the current index mean value to obtain a current index difference value;

[0154] a second comparison unit, configured to compare the current index difference value with the current index variance value to obtain a current index variance difference value;

[0155] a judgment unit, configured to judge whether the multiple of the current index variance difference value and the current index variance value is greater than a preset multiple;

[0156] a first determination unit, configured to determine that the current hardware index data is abnormal index data when the multiple is greater than the preset multiple;

[0157] a second determination unit, configured to determine that the current hardware index data is normal index data when the multiple is less than or equal to the preset multiple.

[0158] In an embodiment, the second analysis module 13 comprises:

[0159] an obtaining unit, configured to obtain a first preset number of abnormal index data of the same type as the current hardware index data;

[0160] a calculation unit, configured to perform linear regression calculation on the current hardware index data and the first preset number of abnormal index data by using a pre-established fault prediction model to obtain a current index calculation result;

[0161] a third determination unit, configured to determine a hardware index change trend according to the current index calculation result;

[0162] a prediction unit, configured to predict a third preset number of prediction index data in the future according to the current index calculation result and a second preset number of index calculation results when it is determined that the corresponding hardware index deviates from the current index mean value according to the hardware index change trend.

[0163] In an embodiment, the apparatus can further comprise:

[0164] The fourth determining unit is configured to determine the next reference data according to all historical index data corresponding to the hardware index data of each type in the historical data of the resource pool.

[0165] In an embodiment, the apparatus can further include:

[0166] The callback unit is configured to perform drift detection callback on the next reference data in a case where it is determined that the current hardware index data is abnormal index data and the next reference data is determined.

[0167] In an embodiment, the callback unit includes:

[0168] The first determining sub-unit is configured to determine the target reference value according to the service type input by the user in advance and in combination with the pre-established corresponding relationship between the service type and the reference value.

[0169] The adjusting sub-unit is configured to adjust the next reference data according to the target reference value.

[0170] In an embodiment, the callback unit includes:

[0171] The obtaining sub-unit is configured to obtain the third preset number of index calculation results calculated by linear regression calculation.

[0172] The second determining sub-unit is configured to determine the first reference value according to the third preset number of index calculation results.

[0173] The third determining sub-unit is configured to determine the adjusted next reference data according to the first reference value, the first proportion, the reference value corresponding to the next reference data, and the second proportion.

[0174] In an embodiment, the fault prediction model is trained based on the historical index data corresponding to each type of hardware index data in combination with the attention mechanism and the first double-layer long short-term memory network.

[0175] In an embodiment, the fault prediction model is trained based on the historical index data corresponding to each type of hardware index data in combination with the attention mechanism and the first double-layer long short-term memory network, and includes:

[0176] The historical index data corresponding to each type of hardware index data is subjected to data enhancement, and the model training is performed based on the data-enhanced historical index data in combination with the attention mechanism and the first double-layer long short-term memory network, and the training process is adjusted based on a dynamic learning rate.

[0177] In a case where the loss values corresponding to the continuous preset number of training rounds are not reduced, the training is ended to obtain the trained fault prediction model.

[0178] In an embodiment, the regulation module 14 is configured to:

[0179] The resource regulation of the server is performed based on the predicted index data by using a pre-established resource dynamic allocation model, and the resource dynamic allocation model is obtained by training based on an attention mechanism and a second double-layer long short-term memory network.

[0180] It should be noted that the server data supervision apparatus provided in the embodiments of the present application has the same beneficial effects as the server data supervision method provided in the above embodiments. For the features of the embodiments corresponding to the server data supervision apparatus in the embodiments of the present application, reference can be made to the related descriptions of the embodiments corresponding to the server data supervision method, which will not be repeated here.

[0181] The embodiments of the present application also provide an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-mentioned server data supervision method embodiments.

[0182] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned server data supervision method embodiments when running.

[0183] In an exemplary embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0184] The embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned server data supervision method embodiments.

[0185] The embodiments of the present application also provide another computer program product, which includes a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned server data supervision method embodiments.

[0186] Those skilled in the art will further realize that the mere conception of the examples described herein is sufficient to enable practitioners to practice the examples as further described below. Therefore, many modifications and adaptations will be apparent to those skilled in the art (e.g., variations in sizes, dimensions, structures, materials, and / or use of the example's components, changes in the arrangement of components, methods, and / or functions, etc.). For example, the dimensions and / or types of the components can be varied, and the size of the components can be either increased to accommodate larger wireless devices or decreased to accommodate smaller wireless devices. Therefore, the examples are not limited to the specific examples described herein, but instead have numerous applications, modifications and adaptations from the various, and / or obvious, combinations where indicated, of the components described herein, and / or variations of the methods described herein. Therefore, the scope of the application should be determined not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. The disclosures of each patent, patent application, and publication cited above are hereby incorporated herein by reference.

[0187] The above provides a kind of server's data monitoring method and electronic equipment provided by the present application in detail.The principle and implementation of the present application are described in the specific examples in this paper, the above example is only applicable to help understanding the method of the present application and its core idea.It should be pointed out that, for the ordinary skilled in the art, under the premise of not departing from the principle of the present application, the present application can be improved and modified in several ways, and these improvements and modifications also fall within the scope of the claims of the present application.

Claims

1. A method for monitoring server data, characterized in that, include: Collect various current hardware performance data from the server; For each type of current hardware indicator data, the current hardware indicator data is analyzed using the current reference data corresponding to the current hardware indicator data to determine whether the current hardware indicator data is abnormal indicator data; the current reference data is determined based on the corresponding historical indicator data stored in the historical data of the resource pool; If the current hardware indicator data is determined to be abnormal, a pre-established fault prediction model is used to analyze the current hardware indicator data to obtain the corresponding predicted indicator data. Server resources are adjusted based on the predicted indicator data; wherein: The step of analyzing the current hardware indicator data using a pre-established fault prediction model to obtain corresponding predicted indicator data includes: acquiring a first preset number of abnormal indicator data with the same data type as the current hardware indicator; performing linear regression calculations on the current hardware indicator data and the first preset number of abnormal indicator data using the pre-established fault prediction model to obtain the current indicator calculation result; determining the hardware indicator change trend based on the current indicator calculation result and the results of previous calculations of the current indicator; and predicting a third preset number of predicted indicator data based on the current indicator calculation result and the results of the second preset number of indicator calculations, when the corresponding hardware indicator deviates from the empirical value of the current indicator based on the hardware indicator change trend. It also includes: for each type of hardware indicator data, determining the next reference data based on all historical indicator data corresponding to the type of hardware indicator data in the historical data of the resource pool; and when it is determined that the current hardware indicator data is abnormal indicator data and the next reference data is determined, performing a drift detection callback on the next reference data. For each type of hardware indicator data, determining the next reference data based on all historical indicator data corresponding to that type of hardware indicator data in the historical data of the resource pool includes: storing all current hardware indicator data (excluding abnormal indicator data) in the historical data of each type of current hardware indicator data into the historical data of the resource pool, and assigning a corresponding weight value to each other current hardware indicator data, then calculating the next reference data corresponding to that type based on each historical indicator data belonging to the same type in the historical data of the resource pool and its corresponding weight value; the next reference data includes the next reference empirical value and the next reference variance value; The drift detection callback for the next reference data includes: obtaining the calculation results of a third preset number of indicators calculated by linear regression; determining a first reference value based on the calculation results of the third preset number of indicators; and determining the adjusted next reference data based on the first reference value, the first ratio, the reference value corresponding to the next reference data, and the second ratio.

2. The server data monitoring method according to claim 1, characterized in that, Also includes: All current hardware indicator data, excluding the abnormal indicator data, are stored in the historical data of the resource pool. Assign corresponding weight values ​​to each of the other current hardware metric data in the historical data stored in the resource pool.

3. The server data monitoring method according to claim 1, characterized in that, The current reference data includes the current empirical value of the indicator and the current variance of the indicator; The step of analyzing the current hardware indicator data using current reference data corresponding to the current hardware indicator data to determine whether the current hardware indicator data is abnormal indicator data includes: The current hardware indicator data value is compared with the current indicator empirical value to obtain the current indicator difference; The current indicator difference is compared with the current indicator variance to obtain the current indicator variance difference. Determine whether the multiple of the variance difference of the current indicator to the variance of the current indicator is greater than a preset multiple; If the multiple is greater than the preset multiple, the current hardware indicator data is determined to be abnormal indicator data; If the multiple is less than or equal to the preset multiple, the current hardware indicator data is determined to be normal indicator data.

4. The server data monitoring method according to any one of claims 1 to 3, characterized in that, The fault prediction model is trained based on historical indicator data corresponding to each type of hardware indicator data, combined with an attention mechanism and a first two-layer long short-term memory network.

5. The server data monitoring method according to claim 4, characterized in that, The fault prediction model is trained based on historical indicator data corresponding to each type of hardware indicator data, combined with an attention mechanism and a first two-layer long short-term memory network, and includes: Data augmentation is performed on the historical data corresponding to each hardware indicator, and the model is trained based on the augmented historical data, combined with the attention mechanism and the first two-layer long short-term memory network, and the training process is adjusted based on the dynamic learning rate. If the loss value does not decrease in any of the preset number of training rounds, the training ends, and the trained fault prediction model is obtained.

6. The server data monitoring method according to claim 4, characterized in that, The regulation of server resources based on the predicted indicator data includes: The server's resources are adjusted based on the predicted index data using a pre-established dynamic resource allocation model; the dynamic resource allocation model is trained based on an attention mechanism and a second two-layer long short-term memory network.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the data monitoring method for the server as described in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Traffic anti-cheating model evaluation method and device, equipment and storage medium

    CN116418581A

  • Container migration method and apparatus for container cluster management system, device, and medium

    WO2025025492A1