Management method and device for monitoring server

By using the pre-trained server status detection model to analyze the device status characteristics of the monitoring server, detect and generate alarm information in real time, the problem of failure of the monitoring server on the dynamic ring platform cannot be detected in time, and the monitoring efficiency and system reliability are improved.

CN119938449APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510052457.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The monitoring server of the dynamic ring platform cannot be discovered in time when a failure occurs, resulting in reduced monitoring efficiency and impact on data accuracy. The existing technology relies on manual regular inspections or user reports, which is inefficient and prone to omissions or misjudgments.

Method used

A management method for monitoring servers is adopted, and the device status characteristics of the monitoring server are analyzed based on the pre-trained server status detection model, and the server status abnormality is detected in real time and alarm information is generated. The model includes a long and short-term memory network layer, an attention mechanism layer, and a residual connection layer, which can capture long-term dependencies in the time series and identify abnormal states of the server.

Benefits of technology

Real-time monitoring and abnormal detection of monitoring server status is realized, shortening the time interval from abnormal detection to operation and maintenance processing, improving the reliability and efficiency of the system, reducing labor costs, and ensuring the optimal monitoring status of key facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938449A_ABST
    Figure CN119938449A_ABST
Patent Text Reader

Abstract

The invention discloses a management method and device for a monitoring server. The method comprises the steps that multiple sets of first device state data of multiple target devices monitored by a target monitoring server in a first time period are acquired based on a first data acquisition frequency, and each target device corresponds to one set of first device state data at each data acquisition moment; determining a first equipment state feature corresponding to each group of first equipment state data; a pre-trained server state detection model is utilized to analyze the plurality of first device state features to obtain a first state analysis result of the target monitoring server, and the server state detection model at least comprises a long and short term memory network layer, an attention mechanism layer and a residual connection layer; and under the condition that the first state analysis result indicates that the state of the target monitoring server is abnormal, server abnormity alarm information is generated. The technical problem that the fault of the monitoring server cannot be found in time in a dynamic environment platform monitoring scene is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computer room dynamic environment management, and in particular, to a management method and device for a monitoring server. Background Art

[0002] The power environment monitoring platform is a real-time monitoring system designed for key facilities such as data centers, machine buildings, and access network rooms. The platform focuses on monitoring environmental parameters and equipment status, such as core indicators such as temperature, humidity, power supply status, and air conditioning operation, to ensure that these facilities are maintained in optimal working conditions to ensure the stability and reliability of the internal IT (Information Technology) system. However, in actual operation, the monitoring server of the dynamic environment platform, the Supervisory Unit (SU), occasionally "hangs", that is, the SU stops responding or cannot perform its duties normally due to certain factors. This problem is often only discovered when the failure has already occurred. With the expansion of the scale of data centers and the increase in the complexity of IT infrastructure, the number of SU servers managed by the dynamic environment platform has increased significantly. Each SU may be responsible for monitoring multiple sensors or detection nodes. Therefore, once a SU hangs, it will directly lead to a lag in the update of information in its monitoring area, affecting the monitoring efficiency and data accuracy of the entire system. Currently, the dynamic environment platform mainly relies on manual regular inspections or user reports to detect the status of SU. This traditional method not only consumes a lot of human resources, but is also prone to omissions or misjudgments when processing massive data. In addition, when there is a problem with the SU, the data distribution mechanism may cause record confusion or data loss, which undoubtedly increases the complexity of the maintenance team's work. In view of the above challenges, there is an urgent need to develop a more intelligent and automated solution to improve the reliability and efficiency of the dynamic environment platform, reduce labor costs, and ensure that all key facilities are always under optimal monitoring.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide a management method and device for a monitoring server, so as to at least solve the technical problem that a monitoring server failure cannot be discovered in time in a dynamic environment platform monitoring scenario.

[0005] According to one aspect of an embodiment of the present application, a management method for a monitoring server is provided, comprising: obtaining multiple groups of first device status data of each of multiple target devices monitored by a target monitoring server within a first time period based on a first data collection frequency, wherein each target device corresponds to a group of first device status data at each data collection moment; determining first device status features corresponding to each group of first device status data; analyzing the multiple first device status features using a pre-trained server status detection model to obtain a first status analysis result of the target monitoring server, wherein the server status detection model includes at least: a long short-term memory network layer, an attention mechanism layer, and a residual connection layer; generating server abnormality alarm information when the first status analysis result indicates that the target monitoring server status is abnormal.

[0006] Optionally, after obtaining the first status analysis result of the target monitoring server, the method also includes: when the first status analysis result indicates that the status of the target monitoring server is abnormal, obtaining multiple groups of second device status data of multiple target devices in a second time period based on a second data acquisition frequency, wherein the second time period is a sub-time period within the first time period, the second data acquisition frequency is greater than the first data acquisition frequency, and each target device corresponds to a group of second device status data at each data acquisition moment; determining the second device status characteristics corresponding to each group of second device status data; analyzing the multiple second device status characteristics using a server status detection model to obtain the second status analysis result of the target monitoring server; when the second status analysis result indicates that the status of the target monitoring server is normal, determining that the status of the target monitoring server is normal; when the second status analysis result indicates that the status of the target monitoring server is abnormal, generating server abnormality alarm information.

[0007] Optionally, the method also includes: using a server status detection model to analyze multiple first device status characteristics to obtain third device status data for each target device within a third time period, wherein the third time period is a time period after the first time period; for each target device, when the third device status data corresponding to the target device is greater than a preset status threshold, generating device abnormality alarm information.

[0008] Optionally, based on the first data collection frequency, multiple groups of first device status data of each of the multiple target devices monitored by the target monitoring server within the first time period are obtained, including: obtaining cookie information of the power environment monitoring platform to which the target monitoring server belongs, and accessing the monitoring device tree of the power environment monitoring platform based on the cookie information, wherein the monitoring device tree includes multiple monitoring servers under the power environment monitoring platform and structural relationships and device information between multiple devices monitored by each monitoring server; determining multiple data collection moments in the first time period based on the first data collection frequency, and obtaining the device status coding data corresponding to each target device at each data collection moment from the monitoring device tree; determining a data decoding method corresponding to the data coding method in the monitoring device tree to decode the multiple groups of device status coding data to obtain multiple groups of first device status data.

[0009] Optionally, determining the first device status feature corresponding to each group of first device status data includes: for each group of first device status data, preprocessing the first device status data, wherein the preprocessing includes at least one of the following: data cleaning, data normalization; extracting at least one of the following target features in the preprocessed first device status data: data acquisition timestamp, device identification, ambient temperature, ambient humidity, ambient pressure, device voltage, device current; and using a multi-dimensional fusion feature vector formed by combining the various target features as the first device status feature corresponding to the first device status data.

[0010] Optionally, the training process of the server status detection model includes: constructing an initial prediction model, wherein the initial prediction model includes at least: an input layer, an embedding layer, a long short-term memory network layer and an attention mechanism layer for feature analysis and weight determination, a first fully connected layer, a first residual connection layer and a first output layer for predicting the status of the monitoring server based on the feature analysis and weight determination results, and a second fully connected layer, a second residual connection layer and a second output layer for predicting the future status data of each device based on the feature analysis and weight determination results; obtaining multiple groups of fourth device status data of multiple target devices in multiple continuous historical time periods, and pre-processing each group of fourth device status data Processing, wherein the preprocessing includes at least one of the following: data cleaning, data normalization; for each historical time period corresponding to a plurality of groups of fourth device status data after preprocessing, data enhancement processing is performed on the fourth device status data, and the fourth device status features corresponding to each group of fourth device status data after the data enhancement processing are determined, the plurality of fourth device status features are used as a training sample, and the status information of the target monitoring server in the historical time period and the plurality of groups of fourth device status data in the next historical time period of the historical time period are used as sample labels of the training samples; the initial prediction model is iteratively trained using the plurality of training samples and the sample labels to obtain a server status detection model.

[0011] Optionally, the initial prediction model is iteratively trained using multiple training samples and sample labels to obtain a server status detection model, including: dividing multiple training samples into a training set, a validation set and a test set, and determining the hyperparameters of the training model, wherein the hyperparameters include at least a learning rate and a decay rate; in each training cycle, based on the hyperparameters, the initial prediction model is iteratively trained using the training set, wherein in each training batch, a target loss function is constructed based on the model output and the corresponding sample labels, and the model parameters are updated based on the target loss function and a back propagation algorithm, and the target loss function includes a cross entropy loss function; at the end of each training cycle, the performance of the trained model is verified using the validation set, and the hyperparameters are adjusted based on the verification results; after the training is completed, the performance of the obtained server status detection model is tested using the test set.

[0012] Optionally, the method also includes: taking multiple first device state features as a new training sample, and storing the second state analysis results as corresponding sample labels; taking multiple second device state features as a new training sample, and storing the second state analysis results as corresponding sample labels; periodically using all stored new training samples and sample labels to train the server state detection model and update the model parameters.

[0013] According to another aspect of an embodiment of the present application, a management device for a monitoring server is also provided, including: an acquisition module, used to acquire multiple groups of first device status data of multiple target devices monitored by a target monitoring server within a first time period based on a first data collection frequency, wherein each target device corresponds to a group of first device status data at each data collection moment; a feature determination module, used to determine the first device status feature corresponding to each group of first device status data; a state analysis module, used to analyze the multiple first device status features using a pre-trained server state detection model to obtain a first state analysis result of the target monitoring server, wherein the server state detection model includes at least: a long short-term memory network layer, an attention mechanism layer and a residual connection layer; a management module, used to generate server abnormality alarm information when the first state analysis result indicates that the state of the target monitoring server is abnormal.

[0014] According to another aspect of an embodiment of the present application, a computer program product is further provided, the computer program product comprising: a computer program, wherein when the computer program is executed by a processor, the above-mentioned management method of the monitoring server is implemented.

[0015] According to another aspect of an embodiment of the present application, an electronic device is further provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned management method of the monitoring server through the computer program.

[0016] In an embodiment of the present application, the state data of multiple target devices monitored by the target monitoring server within a first time period are collected according to a preset first data collection frequency, thereby ensuring the real-time and sequence integrity of the data and providing a solid foundation for subsequent anomaly detection; the collected original device state data are processed to extract the state features corresponding to each group of data. This step involves data preprocessing, including but not limited to normalization and multimodal fusion, which improves data quality and optimizes model input, so that the model can more accurately identify key signals of device state changes; a pre-trained server state detection model is used to deeply analyze the extracted device state features. The model includes at least a long short-term memory network layer, an attention mechanism layer, and a residual memory layer. Differential connection layer, these components work together to enable the model to capture long-term dependencies in time series, focus on key information in the data, and alleviate the gradient vanishing problem in the training process, so as to more accurately identify abnormal states of the server; based on the analysis results of the model, the first state analysis results of the target monitoring server are obtained. When the analysis results show that the monitoring server state is abnormal, a server abnormality alarm message will be generated immediately to notify maintenance personnel. This immediate response mechanism shortens the time interval from abnormality detection to operation and maintenance processing, ensuring that the problem can be quickly identified and resolved, thereby avoiding the impact of abnormal states on system operation, and thus solving the technical problem of not being able to detect monitoring server failures in a timely manner in the dynamic environment platform monitoring scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 It is a flowchart of an optional monitoring server management method according to an embodiment of the present application;

[0019] Figure 2 is a schematic diagram of the structure of an optional management device for monitoring a server according to an embodiment of the present application;

[0020] Figure 3 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0022] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] Example 1

[0024] According to an embodiment of the present application, a management method for a monitoring server is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0025] Figure 1 is a flow chart of a management method for a monitoring server provided according to an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:

[0026] Step S102, acquiring multiple groups of first device status data of each of multiple target devices monitored by the target monitoring server within a first time period based on the first data collection frequency, wherein each target device corresponds to a group of first device status data at each data collection moment;

[0027] Step S104, determining the first device status feature corresponding to each group of first device status data;

[0028] Step S106, using a pre-trained server status detection model to analyze multiple first device status features to obtain a first status analysis result of the target monitoring server, wherein the server status detection model at least includes: a long short-term memory network layer, an attention mechanism layer, and a residual connection layer;

[0029] Step S108: When the first status analysis result indicates that the target monitoring server is in an abnormal state, server abnormality alarm information is generated.

[0030] The following describes each step of the management method for the monitoring server in conjunction with a specific implementation process.

[0031] First, based on the first data collection frequency, multiple groups of first device status data of each of the multiple target devices monitored by the target monitoring server within a first time period are obtained, wherein each target device corresponds to a group of first device status data at each data collection moment. The process may include the following steps:

[0032] Obtain the cookie information of the power environment monitoring platform to which the target monitoring server belongs, and access the monitoring device tree of the power environment monitoring platform based on the cookie information, wherein the monitoring device tree includes multiple monitoring servers under the power environment monitoring platform and the structural relationship and device information between multiple devices monitored by each monitoring server;

[0033] Determine multiple data collection moments in the first time period based on the first data collection frequency, and obtain device state code data corresponding to each target device at each data collection moment from the monitoring device tree;

[0034] A data decoding method corresponding to a data encoding method in a monitoring device tree is determined to decode a plurality of sets of device status encoding data to obtain a plurality of sets of first device status data.

[0035] It should be noted that obtaining the first device status data based on the first data collection frequency can be performed in real time or periodically, for example, every hour, every 10 minutes, etc. The selection of this first data collection frequency (time interval) is usually based on the urgency of monitoring needs and the real-time requirements of the data.

[0036] After obtaining multiple sets of first device status data, determining the first device status feature corresponding to each set of first device status data may include the following steps:

[0037] For each set of first device status data, preprocessing the first device status data, wherein the preprocessing includes at least one of the following: data cleaning and data normalization processing;

[0038] Extracting at least one of the following target features from the preprocessed first device status data: data acquisition timestamp, device identification, ambient temperature, ambient humidity, ambient pressure, device voltage, and device current;

[0039] The multi-dimensional fusion feature vector formed by combining various target features is used as the first device state feature corresponding to the first device state data.

[0040] After obtaining the above-mentioned first device state characteristics, a plurality of first device state characteristics are analyzed using a pre-trained server state detection model to obtain a first state analysis result of the target monitoring server, wherein the server state detection model includes at least: a long short-term memory network layer, an attention mechanism layer and a residual connection layer.

[0041] As an optional implementation, the training process of the server status detection model may include the following steps:

[0042] S1, constructing an initial prediction model, wherein the initial prediction model at least includes: an input layer, an embedding layer, a long short-term memory network layer, and an attention mechanism layer for feature analysis and weight determination, a first fully connected layer, a first residual connection layer, and a first output layer for predicting the state of the monitoring server based on the feature analysis and weight determination results, and a second fully connected layer, a second residual connection layer, and a second output layer for predicting the future state data of each device based on the feature analysis and weight determination results;

[0043] Among them, the input layer is used to receive input data, where the data is multimodal, including environmental parameters such as temperature, voltage, and current; the embedding layer maps the input raw data to a high-dimensional space to enhance the expression ability of the model. For classification tasks, the embedding layer can convert category data into dense vectors; the long short-term memory network layer is used to process time series data and capture long-term dependencies in the data. In server status detection, it can learn the law of device status changes over time, so as to more accurately predict whether the server is hanging; after the long short-term memory network layer, the introduction of the attention mechanism can improve the model's attention to key information in the input sequence. For server status detection, it helps the model focus on data points that are more important for judging server status abnormalities; a residual connection layer is introduced between the long short-term memory network layer and the fully connected layer to alleviate the gradient vanishing problem and ensure that the model can still learn effectively in the deep network; the first fully connected layer and the first output layer are used for feature extraction and classification, and output the probability prediction of whether the server is hanging; the second fully connected layer, the second residual connection layer, and the second output layer are used for multi-task learning, which can predict the future status data of each device and assist in improving the accuracy of the main task (i.e., hanging detection).

[0044] Among them, the main task can be understood as detecting whether the server is hanging, that is, by analyzing the state characteristics of the first device (such as temperature, voltage, current, etc.), predicting whether the server is running normally at a given time point. The secondary task can be to predict the future state data of each device, such as temperature prediction, voltage prediction, etc. By introducing secondary tasks, the model can learn the laws of environmental parameter changes while also learning the impact of these changes on the server status, thereby improving the model's sensitivity and accuracy in detecting server hanging.

[0045] S2, acquiring multiple groups of fourth device status data of multiple target devices in multiple continuous historical time periods, and preprocessing each group of fourth device status data, wherein the preprocessing includes at least one of the following: data cleaning and data normalization processing;

[0046] It should be noted that the multiple sets of fourth device status data used for training are collected in the same manner as the first device status data.

[0047] S3, for the multiple groups of pre-processed fourth device status data corresponding to each historical time period, perform data enhancement processing on the fourth device status data, and determine the fourth device status features corresponding to each group of fourth device status data after the data enhancement processing, use the multiple fourth device status features as a training sample, and use the status information of the target monitoring server in the historical time period and the multiple groups of fourth device status data in the next historical time period of the historical time period as sample labels of the training samples;

[0048] S4, iteratively train the initial prediction model using multiple training samples and sample labels to obtain a server status detection model. This process can be carried out in the following steps:

[0049] Divide multiple training samples into a training set, a validation set, and a test set, and determine hyperparameters of the training model, wherein the hyperparameters include at least: a learning rate and a decay rate;

[0050] In each training cycle, the initial prediction model is iteratively trained using the training set based on the hyperparameters. In each training batch, the target loss function is constructed based on the model output and the corresponding sample labels, and the model parameters are updated based on the target loss function and the back propagation algorithm. The target loss function includes: cross entropy loss function;

[0051] At the end of each training cycle, the performance of the trained model is verified using the validation set, and the hyperparameters are adjusted based on the verification results.

[0052] After the training is completed, the performance of the obtained server status detection model is tested using the test set.

[0053] After obtaining the trained server status detection model, the model is used to analyze multiple first device status features to obtain a first status analysis result. When the first status analysis result indicates that the target monitored server status is abnormal, server abnormality alarm information is generated.

[0054] As an optional implementation, after obtaining the first status analysis result of the target monitoring server, a secondary confirmation can be performed by increasing the data collection frequency to improve the accuracy and reliability of the alarm information. The process can be carried out in the following steps: when the first status analysis result indicates that the status of the target monitoring server is abnormal, multiple groups of second device status data of multiple target devices in a second time period are obtained based on the second data collection frequency, wherein the second time period is a sub-time period within the first time period, the second data collection frequency is greater than the first data collection frequency, and each target device corresponds to a group of second device status data at each data collection moment, thereby being able to capture in more detail possible changes in the server status in a short period of time. It should be noted that the multiple groups of second device status data are collected in the same way as the above-mentioned first device status data; determining each group of second device status data According to the corresponding second device status characteristics; use the server status detection model to analyze multiple second device status characteristics to obtain the second status analysis result of the target monitoring server; when the second status analysis result indicates that the target monitoring server status is normal, determine that the target monitoring server status is normal, and through secondary analysis, it can be further confirmed whether the server is really in an abnormal state; when the second status analysis result indicates that the target monitoring server status is abnormal, generate server abnormality alarm information. If the server status is determined to be normal in the second status analysis result, then the previous abnormal alarm may be a false alarm, and the abnormal alarm will no longer be triggered, and the server status will be marked as normal, avoiding unnecessary waste of resources and manpower intervention. This alarm mechanism based on secondary confirmation improves the accuracy of alarm information and reduces the time and energy of operation and maintenance personnel in handling false alarms.

[0055] According to the above process, the abnormal conditions detected by the model and their subsequent secondary confirmation results can also be used to continuously optimize and update the server status detection model. The process can be carried out according to the following steps: taking multiple first device status features as a new training sample, and storing the second status analysis results as corresponding sample labels; taking multiple second device status features as a new training sample, and storing the second status analysis results as corresponding sample labels; periodically using all the stored new training samples and sample labels to train the server status detection model, and update the model parameters. By continuously collecting and using abnormal detection results (whether confirming anomalies or eliminating false alarms) to train the model, the performance of the model can be continuously optimized.

[0056] In addition to the above, the server status detection model can also be used to analyze multiple first device status characteristics to obtain third device status data for each target device within a third time period, wherein the third time period is the time period after the first time period; for each target device, when the third device status data corresponding to the target device is greater than a preset status threshold, a device abnormality alarm information is generated.

[0057] It can be understood that the model can be used to monitor the future status of equipment and warn of problems in advance. The model will analyze based on the patterns it has learned and predict the status data of the equipment in the third time period. These data may include but are not limited to environmental parameters such as temperature, voltage, and current. The third time period here refers to any time interval after the first time period. The system sets a preset status threshold to determine whether the equipment status data exceeds the normal range. The status thresholds of different equipment and different dimensions are different. For example, the temperature threshold of a certain equipment may be set to 30°C. If the temperature value predicted by the model exceeds this threshold, it may indicate that the equipment is overheated and needs attention. When the third equipment status data predicted by the model exceeds the preset status threshold, the system will generate equipment abnormality alarm information. These alarm messages can be alarms sent immediately to operation and maintenance personnel through emails, text messages, platform notifications, etc., prompting them to pay attention to and handle the abnormal status of specific equipment.

[0058] By extending the model application to monitor other abnormal conditions of the equipment, it is not only possible to detect server hangs in real time, but also to provide early warnings for other problems that may affect server stability, such as overheating and voltage anomalies, thereby enhancing the system's preventive and responsive capabilities. For example, excessively high temperatures may be a sign that a server is about to hang. By monitoring and providing early warnings, cooling measures can be taken in advance to avoid server hangs. This approach further ensures the stable operation of data centers and computer rooms, and improves the efficiency and accuracy of operation and maintenance work.

[0059] In an embodiment of the present application, the state data of multiple target devices monitored by the target monitoring server within a first time period are collected according to a preset first data collection frequency, thereby ensuring the real-time and sequence integrity of the data and providing a solid foundation for subsequent anomaly detection; the collected original device state data are processed to extract the state features corresponding to each group of data. This step involves data preprocessing, including but not limited to normalization and multimodal fusion, which improves data quality and optimizes model input, so that the model can more accurately identify key signals of device state changes; a pre-trained server state detection model is used to deeply analyze the extracted device state features. The model includes at least a long short-term memory network layer, an attention mechanism layer, and a residual memory layer. Differential connection layer, these components work together to enable the model to capture long-term dependencies in time series, focus on key information in the data, and alleviate the gradient vanishing problem in the training process, so as to more accurately identify abnormal states of the server; based on the analysis results of the model, the first state analysis results of the target monitoring server are obtained. When the analysis results show that the monitoring server state is abnormal, a server abnormality alarm message will be generated immediately to notify maintenance personnel. This immediate response mechanism shortens the time interval from abnormality detection to operation and maintenance processing, ensuring that the problem can be quickly identified and resolved, thereby avoiding the impact of abnormal states on system operation, and thus solving the technical problem of not being able to detect monitoring server failures in a timely manner in the dynamic environment platform monitoring scenario.

[0060] Example 2

[0061] According to an embodiment of the present application, a management device for a monitoring server for implementing the management method of the monitoring server in Embodiment 1 is also provided. Figure 2 As shown, the management device of the monitoring server at least includes: an acquisition module 21, a feature determination module 22, a state analysis module 23 and a management module 24, wherein:

[0062] The acquisition module 21 is used to acquire multiple groups of first device status data of each of the multiple target devices monitored by the target monitoring server within a first time period based on the first data acquisition frequency, wherein each target device corresponds to a group of first device status data at each data acquisition moment;

[0063] A feature determination module 22, configured to determine a first device status feature corresponding to each set of first device status data;

[0064] A state analysis module 23 is used to analyze the state features of the plurality of first devices using a pre-trained server state detection model to obtain a first state analysis result of the target monitoring server, wherein the server state detection model at least includes: a long short-term memory network layer, an attention mechanism layer, and a residual connection layer;

[0065] The management module 24 is used to generate server abnormality alarm information when the first status analysis result indicates that the target monitoring server is in an abnormal state.

[0066] The functions of each module of the management device of the monitoring server are explained below in conjunction with a specific implementation process.

[0067] First, the acquisition module acquires multiple groups of first device status data of multiple target devices monitored by the target monitoring server within a first time period based on the first data acquisition frequency, wherein each target device corresponds to a group of first device status data at each data acquisition moment. The process may include the following steps:

[0068] Obtain the cookie information of the power environment monitoring platform to which the target monitoring server belongs, and access the monitoring device tree of the power environment monitoring platform based on the cookie information, wherein the monitoring device tree includes multiple monitoring servers under the power environment monitoring platform and the structural relationship and device information between multiple devices monitored by each monitoring server;

[0069] Determine multiple data collection moments in the first time period based on the first data collection frequency, and obtain device state code data corresponding to each target device at each data collection moment from the monitoring device tree;

[0070] A data decoding method corresponding to a data encoding method in a monitoring device tree is determined to decode a plurality of sets of device status encoding data to obtain a plurality of sets of first device status data.

[0071] After obtaining multiple groups of first device status data, the feature determination module determines the first device status feature corresponding to each group of first device status data. The process may include the following steps: for each group of first device status data, preprocessing the first device status data, wherein the preprocessing includes at least one of the following: data cleaning, data normalization; extracting at least one of the following target features in the preprocessed first device status data: data acquisition timestamp, device identification, ambient temperature, ambient humidity, ambient pressure, device voltage, device current; and using the multi-dimensional fusion feature vector formed by combining the various target features as the first device status feature corresponding to the first device status data.

[0072] After obtaining the above-mentioned first device state characteristics, the state analysis module uses a pre-trained server state detection model to analyze multiple first device state characteristics to obtain a first state analysis result of the target monitoring server, wherein the server state detection model includes at least: a long short-term memory network layer, an attention mechanism layer and a residual connection layer.

[0073] As an optional implementation, the training process of the server status detection model may include the following steps:

[0074] S1, constructing an initial prediction model, wherein the initial prediction model at least includes: an input layer, an embedding layer, a long short-term memory network layer, and an attention mechanism layer for feature analysis and weight determination, a first fully connected layer, a first residual connection layer, and a first output layer for predicting the state of the monitoring server based on the feature analysis and weight determination results, and a second fully connected layer, a second residual connection layer, and a second output layer for predicting the future state data of each device based on the feature analysis and weight determination results;

[0075] S2, acquiring multiple groups of fourth device status data of multiple target devices in multiple continuous historical time periods, and preprocessing each group of fourth device status data, wherein the preprocessing includes at least one of the following: data cleaning and data normalization processing;

[0076] S3, for the multiple groups of pre-processed fourth device status data corresponding to each historical time period, perform data enhancement processing on the fourth device status data, and determine the fourth device status features corresponding to each group of fourth device status data after the data enhancement processing, use the multiple fourth device status features as a training sample, and use the status information of the target monitoring server in the historical time period and the multiple groups of fourth device status data in the next historical time period of the historical time period as sample labels of the training samples;

[0077] S4, iteratively train the initial prediction model using multiple training samples and sample labels to obtain a server status detection model. This process can be carried out in the following steps:

[0078] Divide multiple training samples into training sets, validation sets and test sets, and determine the hyperparameters of the training model, wherein the hyperparameters include at least a learning rate and a decay rate; in each training cycle, based on the hyperparameters, use the training set to iteratively train the initial prediction model, wherein in each training batch, construct a target loss function based on the model output and the corresponding sample label, and update the model parameters based on the target loss function and the back propagation algorithm, and the target loss function includes a cross entropy loss function; at the end of each training cycle, use the validation set to verify the performance of the trained model, and adjust the hyperparameters based on the verification results; after the training is completed, use the test set to test the performance of the obtained server status detection model.

[0079] After the server status detection model is trained in the above manner, the model is used to analyze multiple first device status features to obtain a first status analysis result. When the first status analysis result indicates that the target monitoring server status is abnormal, the management module generates server abnormality alarm information.

[0080] As an optional implementation, after obtaining the first status analysis result of the target monitoring server, the method also includes: when the first status analysis result indicates that the status of the target monitoring server is abnormal, obtaining multiple groups of second device status data of multiple target devices in a second time period based on a second data acquisition frequency, wherein the second time period is a sub-time period within the first time period, the second data acquisition frequency is greater than the first data acquisition frequency, and each target device corresponds to a group of second device status data at each data acquisition moment; determining the second device status characteristics corresponding to each group of second device status data; analyzing the multiple second device status characteristics using a server status detection model to obtain the second status analysis result of the target monitoring server; when the second status analysis result indicates that the status of the target monitoring server is normal, determining that the status of the target monitoring server is normal; when the second status analysis result indicates that the status of the target monitoring server is abnormal, generating server abnormality alarm information.

[0081] According to the above process, the following steps can also be performed: taking multiple first device state features as a new training sample, and storing the second state analysis results as corresponding sample labels; taking multiple second device state features as a new training sample, and storing the second state analysis results as corresponding sample labels; periodically using all stored new training samples and sample labels to train the server state detection model and update the model parameters.

[0082] In addition to the above, the server status detection model can also be used to analyze multiple first device status characteristics to obtain third device status data for each target device within a third time period, wherein the third time period is the time period after the first time period; for each target device, when the third device status data corresponding to the target device is greater than a preset status threshold, a device abnormality alarm information is generated.

[0083] It should be noted that each module in the management device of the monitoring server in the embodiment of the present application corresponds one by one to each implementation step of the management method of the monitoring server in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be repeated here.

[0084] Example 3

[0085] According to an embodiment of the present application, a computer program product is also provided. The computer program product includes a computer program, wherein when the computer program is executed by a processor, the management method of the monitoring server in Example 1 is implemented.

[0086] According to an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the management method of the monitoring server in Example 1 by running the computer program.

[0087] According to an embodiment of the present application, a processor is further provided, which is used to run a computer program, wherein the management method of the monitoring server in Example 1 is executed when the computer program is running.

[0088] According to an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the management method of the monitoring server in Example 1 through the computer program.

[0089] Specifically, the following steps are executed when the computer program is running: based on a first data collection frequency, multiple groups of first device status data of each of the multiple target devices monitored by the target monitoring server within a first time period are obtained, wherein each target device corresponds to a group of first device status data at each data collection moment; first device status features corresponding to each group of first device status data are determined; multiple first device status features are analyzed using a pre-trained server status detection model to obtain a first status analysis result of the target monitoring server, wherein the server status detection model includes at least: a long short-term memory network layer, an attention mechanism layer, and a residual connection layer; when the first status analysis result indicates that the target monitoring server status is abnormal, server abnormality alarm information is generated.

[0090] As an optional implementation, the electronic device may be in the form of a mobile terminal, a computer terminal or a similar computing device. Figure 3 FIG. 1 shows a hardware structure block diagram of an electronic device for implementing a management method for a monitoring server. Figure 3 As shown, the electronic device 30 may include one or more (302a, 302b, ..., 302n are used to illustrate) processors 302 (the processor 302 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 3 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 3More or fewer components as shown, or with Figure 3 Different configurations are shown.

[0091] It should be noted that the one or more processors 302 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the electronic device 30. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0092] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the management method of the monitoring server in the embodiment of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, realizing the vulnerability detection method of the above-mentioned application. The memory 304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 304 may further include a memory remotely arranged relative to the processor 302, and these remote memories may be connected to the electronic device 30 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0093] The transmission device 306 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the electronic device 30. In one example, the transmission device 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 306 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0094] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 30 .

[0095] The serial numbers of the above embodiments are only for description and do not represent the advantages or disadvantages of the embodiments.

[0096] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0098] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0099] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0100] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.

[0101] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A management method for a monitoring server, characterized in that: include: Acquire multiple groups of first device status data of each of multiple target devices monitored by the target monitoring server within a first time period based on the first data collection frequency, wherein each target device corresponds to a group of first device status data at each data collection moment; Determine a first device status feature corresponding to each group of the first device status data; Analyze the state features of the plurality of first devices using a pre-trained server state detection model to obtain a first state analysis result of the target monitoring server, wherein the server state detection model includes at least: a long short-term memory network layer, an attention mechanism layer, and a residual connection layer; When the first status analysis result indicates that the target monitoring server is in an abnormal state, server abnormality alarm information is generated.

2. The method according to claim 1, characterized in that After obtaining the first status analysis result of the target monitoring server, the method further includes: In the case where the first status analysis result indicates that the target monitoring server is in an abnormal state, obtaining multiple groups of second device status data of each of the multiple target devices in a second time period based on a second data collection frequency, wherein the second time period is a sub-time period within the first time period, the second data collection frequency is greater than the first data collection frequency, and each target device corresponds to a group of second device status data at each data collection moment; Determine a second device status feature corresponding to each set of the second device status data; Analyze the state characteristics of the plurality of second devices using the server state detection model to obtain a second state analysis result of the target monitoring server; When the second status analysis result indicates that the target monitoring server is in a normal state, determining that the target monitoring server is in a normal state; When the second status analysis result indicates that the target monitoring server status is abnormal, the server abnormality alarm information is generated.

3. The method according to claim 1, characterized in that The method further comprises: Analyzing the state characteristics of the plurality of the first devices by using the server state detection model to obtain third device state data of each of the target devices in a third time period, wherein the third time period is a time period after the first time period; For each target device, when the third device status data corresponding to the target device is greater than a preset status threshold, device abnormality alarm information is generated.

4. The method according to claim 1, characterized in that: Acquiring multiple groups of first device status data of each of multiple target devices monitored by the target monitoring server within a first time period based on the first data collection frequency includes: Obtain cookie information of the power environment monitoring platform to which the target monitoring server belongs, and access the monitoring device tree of the power environment monitoring platform according to the cookie information, wherein the monitoring device tree includes multiple monitoring servers under the power environment monitoring platform and the structural relationship and device information between multiple devices monitored by each monitoring server; Determine a plurality of data collection moments in the first time period based on the first data collection frequency, and obtain device status code data corresponding to each of the target devices at each data collection moment from the monitoring device tree; Determine a data decoding method corresponding to the data encoding method in the monitoring device tree to decode the multiple groups of device status encoding data to obtain multiple groups of the first device status data.

5. The method according to claim 1, characterized in that: Determining a first device status feature corresponding to each set of the first device status data includes: For each group of the first device status data, preprocessing the first device status data, wherein the preprocessing includes at least one of the following: data cleaning and data normalization; Extracting at least one of the following target features from the preprocessed first device status data: data acquisition timestamp, device identification, ambient temperature, ambient humidity, ambient pressure, device voltage, and device current; A multi-dimensional fusion feature vector formed by combining the target features is used as a first device state feature corresponding to the first device state data.

6. The method according to claim 3, characterized in that The training process of the server status detection model includes: Constructing an initial prediction model, wherein the initial prediction model at least includes: an input layer, an embedding layer, a long short-term memory network layer, and an attention mechanism layer for feature analysis and weight determination, a first fully connected layer, a first residual connection layer, and a first output layer for predicting the state of the monitoring server based on the feature analysis and weight determination results, and a second fully connected layer, a second residual connection layer, and a second output layer for predicting the future state data of each device based on the feature analysis and weight determination results; Acquire multiple groups of fourth device status data of multiple target devices in multiple continuous historical time periods, and preprocess each group of the fourth device status data, wherein the preprocessing includes at least one of the following: data cleaning and data normalization processing; For the preprocessed multiple groups of fourth device status data corresponding to each historical time period, data enhancement processing is performed on the fourth device status data, and the fourth device status features corresponding to each group of fourth device status data after the data enhancement processing are determined, and the multiple fourth device status features are used as a training sample, and the status information of the target monitoring server in the historical time period and the multiple groups of fourth device status data in the next historical time period of the historical time period are used as sample labels of the training samples; The initial prediction model is iteratively trained using a plurality of the training samples and sample labels to obtain the server status detection model.

7. The method according to claim 6, characterized in that Iteratively training the initial prediction model using a plurality of the training samples and sample labels to obtain the server status detection model includes: Dividing the plurality of training samples into a training set, a validation set and a test set, and determining hyperparameters of the training model, wherein the hyperparameters include at least: a learning rate and a decay rate; In each training cycle, the initial prediction model is iteratively trained using the training set based on the hyperparameters, wherein in each training batch, a target loss function is constructed according to the model output and the corresponding sample label, and the model parameters are updated according to the target loss function and the back propagation algorithm, wherein the target loss function includes: a cross entropy loss function; At the end of each training cycle, the performance of the trained model is verified using the verification set, and the hyperparameters are adjusted based on the verification results; After the training is completed, the performance of the obtained server status detection model is tested using the test set.

8. The method according to claim 2, characterized in that: The method further comprises: Taking the plurality of first device state features as a new training sample, and storing the second state analysis result as a corresponding sample label; Taking the plurality of the second device state features as a new training sample, and storing the second state analysis result as a corresponding sample label; The server status detection model is trained periodically using all new stored training samples and sample labels to update model parameters.

9. A management device for monitoring a server, characterized in that: include: An acquisition module, configured to acquire, based on a first data acquisition frequency, a plurality of groups of first device status data of each of a plurality of target devices monitored by a target monitoring server within a first time period, wherein each target device corresponds to a group of first device status data at each data acquisition moment; a feature determination module, configured to determine a first device status feature corresponding to each set of the first device status data; A state analysis module, used to analyze the state features of the plurality of first devices using a pre-trained server state detection model to obtain a first state analysis result of the target monitoring server, wherein the server state detection model at least includes: a long short-term memory network layer, an attention mechanism layer, and a residual connection layer; The management module is used to generate server abnormality alarm information when the first status analysis result indicates that the target monitoring server status is abnormal.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the management method of the monitoring server according to any one of claims 1 to 8 through the computer program.