Intelligent PDU-based Data Center Monitoring Method and System

By obtaining the monitoring data and historical abnormal data of the intelligent PDU, generating model guidance information, using a large model with fine-tuned reinforcement learning to identify abnormalities, and combining infrared image analysis, the problem of lack of timeliness and accuracy in intelligent PDU monitoring is solved, and timely control of abnormal situations in the data center is achieved.

CN120046083BActive Publication Date: 2025-07-11GUANGDONG XINFENGLONG ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510518503.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-11
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the prior art, when monitoring abnormal situations in data centers, intelligent PDUs can only be discovered after serious problems, and lack timeliness and accuracy.

Method used

By obtaining the monitoring data of the intelligent PDU, determining the equipment attributes and historical anomaly data, generating model guidance information, using a large model with fine-tuned reinforcement learning to perform abnormal identification, and combining infrared image analysis to achieve accurate abnormal monitoring.

Benefits of technology

It realizes timely and precise monitoring of abnormal situations in the data center, can be timely regulated, and improves monitoring efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046083B_ABST
    Figure CN120046083B_ABST
Patent Text Reader

Abstract

The present application discloses a data center monitoring method and system based on an intelligent PDU. The method includes: obtaining the monitoring data of the intelligent PDU; determining the device attributes of the intelligent PDU in the data center and extracting attribute feature information from the device attributes; determining abnormal feature information based on the historical abnormal data of the intelligent PDU; generating model guidance information based on the attribute feature information and the abnormal feature information; and invoking an abnormal monitoring model based on the monitoring data and the model guidance information to perform abnormal identification, obtaining an abnormal identification result for regulating the data center. The abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning, so that corresponding guidance information can be generated according to the device attributes and historical abnormal data to guide the abnormal monitoring model obtained by fine-tuning the large model through reinforcement learning to efficiently and accurately analyze the monitoring data, and then timely and accurately monitor the abnormal conditions existing in the data center for corresponding regulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of anomaly monitoring, and in particular, to a data center monitoring method and system based on an intelligent PDU. Background Art

[0002] With the increase in the scale of data centers, it is necessary to conduct accurate real-time monitoring on them. And the intelligent PDU (Power Distribution Unit) as a power management device in the data center can detect the power data and power consumption of each electrical device in the data center. Therefore, by analyzing the data of the intelligent PDU, the power consumption details of the entire data center can be obtained, so as to monitor in real time whether there are faults in the data center.

[0003] In the related art, usually, each index value in the PDU monitoring data is directly compared with the corresponding index threshold value to monitor whether there are anomalies in the data center. However, when a certain index value no longer meets the requirements of the threshold value, the problems existing in the data center are often relatively serious. Summary of the Invention

[0004] To solve the above technical problems, the embodiments of this application propose a data center monitoring method and system based on an intelligent PDU, which can timely and accurately monitor the anomalies existing in the data center for corresponding regulation.

[0005] In a first aspect, the embodiments of this application provide a data center monitoring method based on an intelligent PDU, including:

[0006] Obtain the monitoring data of the intelligent PDU;

[0007] Determine the device attributes of the intelligent PDU in the data center, and extract attribute feature information from the device attributes;

[0008] Determine anomaly feature information based on the historical anomaly data of the intelligent PDU;

[0009] Generate model guidance information based on the attribute feature information and the anomaly feature information;

[0010] Based on the monitoring data and the model guidance information, call an anomaly monitoring model to perform anomaly recognition, and obtain an anomaly recognition result for the data center, so as to regulate the data center, where the anomaly monitoring model is obtained by fine-tuning a large model through reinforcement learning.

[0011] Optionally, the step of based on the monitoring data and the model guidance information, calling an anomaly monitoring model to perform anomaly recognition, and obtaining an anomaly recognition result for the data center includes:

[0012] Determine the data feature information that matches the monitored data;

[0013] Based on the data feature information, retrieve the corresponding anomaly recognition rules in the target knowledge base, where the target knowledge base includes multiple types of anomaly recognition rules;

[0014] Based on the retrieved anomaly recognition rules, the data feature information, and the model guidance information, call the anomaly monitoring model for anomaly recognition to obtain the anomaly recognition result.

[0015] Optionally, the monitored data includes multiple sub-data segments, and different sub-data have different collection times and / or different metric types. The data feature information includes multiple data sub-features corresponding one by one to the multiple sub-data segments;

[0016] The retrieved anomaly recognition rules are suitable for indicating the recognition processing order of the anomaly monitoring model for the multiple data sub-features.

[0017] Optionally, the data center includes multiple PDUs, and the intelligent PDU is at least one of the multiple PDUs. The fine-tuning method of the anomaly monitoring model includes:

[0018] Obtain sample data, where the sample data includes sample monitored data and sample guidance information. The sample monitored data includes the historical monitored data of each of the multiple PDUs, and the sample guidance information is generated based on sample attribute feature information and sample anomaly feature information. The sample attribute feature information is extracted from the device attributes of each of the multiple PDUs in the data center, and the sample anomaly feature information is determined based on the historical anomaly data of each of the multiple PDUs;

[0019] According to the sample data, use reinforcement learning to fine-tune the large model to be fine-tuned.

[0020] Optionally, before calling the anomaly monitoring model for anomaly recognition based on the monitored data and the model guidance information, the method further includes:

[0021] Obtain the inspection examples of the anomaly monitoring model;

[0022] Based on the inspection examples and the model guidance information, call the anomaly monitoring model for anomaly recognition to obtain the example recognition result;

[0023] Adjust the model guidance information according to the example recognition result, and use the adjusted guidance information as the new model guidance information.

[0024] Optionally, based on the inspection example and the model guidance information, the anomaly monitoring model is called for anomaly recognition to obtain an example recognition result, including:

[0025] Based on the inspection example and the model guidance information, the anomaly monitoring model is called multiple times for anomaly recognition respectively to obtain multiple example recognition results;

[0026] The adjustment of the model guidance information according to the example recognition result includes:

[0027] Feature extraction is performed on multiple example recognition results respectively to obtain multiple example recognition features;

[0028] Cluster the multiple example recognition features to obtain a clustering result;

[0029] At least based on the clustering result, the model guidance information is adjusted.

[0030] Optionally, the at least based on the clustering result to adjust the model guidance information includes:

[0031] Determine the example similarity between each of the multiple example recognition features and the clustering result;

[0032] Calculate the average similarity of each example similarity;

[0033] Adjust the model guidance information according to the average similarity.

[0034] Optionally, after calling the anomaly monitoring model for anomaly recognition based on the monitoring data and the model guidance information, the method further includes:

[0035] Obtain multiple infrared images taken of the intelligent PDU, wherein the shooting time of each of the multiple infrared images is included in the acquisition time period of the monitoring data;

[0036] Identify each frame of infrared image to obtain corresponding intelligent PDU infrared information and environmental infrared information;

[0037] Based on the intelligent PDU infrared information, environmental infrared information and shooting time corresponding to each of the multiple infrared images, generate image recognition information, wherein the image recognition information is suitable for indicating the state of the intelligent PDU;

[0038] Input the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly recognition result according to the image recognition information.

[0039] Optionally, generating image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multi-frame infrared images includes:

[0040] Sort the multi-frame infrared images according to the shooting time;

[0041] For any two adjacent infrared images in the sorted multi-frame infrared images, identify the first difference information between the intelligent PDU infrared information corresponding to the two infrared images, and identify the second difference information between the environmental infrared information corresponding to the two infrared images;

[0042] Generate image recognition information based on the shooting time, the first difference information, and the second difference information.

[0043] In a second aspect, an embodiment of the present application provides a data center monitoring system based on an intelligent PDU, including:

[0044] A monitoring module for obtaining monitoring data of the intelligent PDU;

[0045] An attribute feature extraction module for determining the device attributes of the intelligent PDU in the data center and extracting attribute feature information from the device attributes;

[0046] An abnormal feature extraction module for determining abnormal feature information based on historical abnormal data of the intelligent PDU;

[0047] A guidance information generation module for generating model guidance information based on the attribute feature information and the abnormal feature information;

[0048] An abnormal recognition module for performing abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information to obtain an abnormal recognition result for the data center, so as to regulate the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning.

[0049] In summary, the embodiments of the present application have at least the following beneficial effects:

[0050] By adopting the embodiment of the present application, the monitoring data of the intelligent PDU is obtained; the device attributes of the intelligent PDU in the data center are determined, and the attribute feature information is extracted from the device attributes; the abnormal feature information is determined based on the historical abnormal data of the intelligent PDU; the model guiding information is generated based on the attribute feature information and the abnormal feature information; and based on the monitoring data and the model guiding information, an abnormal monitoring model is called for abnormal identification to obtain an abnormal identification result for the data center, so as to be used for regulating the data center. Among them, the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning, so that corresponding guiding information can be generated according to the device attributes and historical abnormal data, so as to guide the abnormal monitoring model obtained by fine-tuning the large model through reinforcement learning to efficiently and accurately analyze the monitoring data, and then timely and accurately monitor the abnormal conditions existing in the data center for corresponding regulation. Description of the Drawings

[0051] Figure 1 It is a schematic flowchart of a data center monitoring method based on an intelligent PDU provided by an embodiment of the present application;

[0052] Figure 2 It is a schematic structural diagram of a data center monitoring system based on an intelligent PDU provided by an embodiment of the present application;

[0053] Figure 3 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed Embodiments

[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0055] In the description of the present application, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more. In the description of the present application, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "according to" means "at least partially according to". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments".

[0056] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0057] In the description of the present application, it should be noted that, unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0058] In a first aspect, referring to Figure 1 , a schematic flow diagram of a data center monitoring method based on an intelligent PDU provided by an embodiment of the present application is shown. The method includes steps S101 - S105, specifically as follows:

[0059] S101, obtaining the monitoring data of the intelligent PDU.

[0060] In one example, the PDU usually has a sensor interface and / or a network interface. The monitoring data may include sensor data collected via the sensor interface and / or network data collected via the network interface. Among them, the sensor data may include at least one of the following: current data, voltage data, power data, temperature data, humidity data.

[0061] S102. Determine the device attributes of the intelligent PDU in the data center, and extract attribute feature information from the device attributes.

[0062] In some cases, the device attributes may refer to the functions and / or characteristics that the intelligent PDU has / undertakes in the data center. The intelligent PDU in the data center is usually not just a simple power distribution device. It can also have one or more advanced functions and characteristics in the data center, thereby improving the management efficiency, reliability, and security of the data center. Thus, exemplarily, the device attributes may include at least one of the following: power distribution and management attributes (such as socket attributes, modular socket attributes, redundant power supply attributes), remote management and control attributes, security attributes (such as the physical isolation attributes it undertakes).

[0063] In one example, the device attributes may include the power distribution attributes corresponding to the intelligent PDU in the data center, which are used to indicate which devices in the data center the intelligent PDU can provide power distribution functions for.

[0064] In one example, attribute feature information can be extracted from the device attributes by performing vectorization processing on the device attributes.

[0065] In one example, a pre-trained feature extraction model (such as a model containing an encoding network) can be used to perform feature extraction on the device attributes to obtain attribute feature information.

[0066] S103. Based on the historical abnormal data of the intelligent PDU, determine abnormal feature information;

[0067] In one example, abnormal feature information can be determined by performing vectorization processing on the historical abnormal data.

[0068] In one example, the above-mentioned feature extraction model can be used to perform feature extraction on the historical abnormal data of the intelligent PDU to obtain abnormal feature information.

[0069] S104. Based on the attribute feature information and the abnormal feature information, generate model guidance information;

[0070] In one example, the attribute feature information and the abnormal feature information can be directly concatenated to obtain model guidance information.

[0071] S105. Based on the monitoring data and the model guidance information, call an abnormal monitoring model for abnormal identification to obtain an abnormal identification result for the data center, which is used to regulate the data center. Among them, the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning. Exemplarily, reinforcement learning may include at least one of the following: imitation learning, multi-agent system, policy optimization, and reinforcement learning.

[0072] In one example, the large model can be a currently common large language model. Generally, the large model can directly perform automatic recognition based on the input. However, in the case of having reliable guiding information, its recognition is often more accurate. Thus, in this embodiment, the model guiding information (including information related to device attributes and historical abnormal data) can be used as the guiding information for the large model, so that the large model can more accurately perform abnormal recognition on the monitoring data.

[0073] In addition, it should be understood that generally, a common large model needs to be fine-tuned to better apply to a specific field. And using reinforcement learning (such as the reward mechanism therein) to fine-tune the large model to make it more suitable for the intelligent PDU abnormal monitoring field can be achieved by using common fine-tuning methods in this field, which is not strictly limited herein.

[0074] In one example, corresponding regulation can be performed on the data center according to the abnormalities existing in the data center indicated by the abnormal recognition result. For example, if it is recognized that the data center has too high a load, the power of the data center can be reduced within a certain period of time in the future.

[0075] In an alternative embodiment, the calling of the abnormal monitoring model based on the monitoring data and the model guiding information to perform abnormal recognition and obtain the abnormal recognition result for the data center includes:

[0076] Determine the data feature information that matches the monitoring data;

[0077] Based on the data feature information, retrieve the corresponding abnormal recognition rule in the target knowledge base, where the target knowledge base includes multiple types of abnormal recognition rules;

[0078] Based on the retrieved abnormal recognition rule, the data feature information, and the model guiding information, call the abnormal monitoring model to perform abnormal recognition and obtain the abnormal recognition result.

[0079] In one example, the data feature information can be obtained by performing vectorization processing on the monitoring data. Or, the data feature information can be obtained by extracting features from the monitoring data through the above-mentioned feature extraction model.

[0080] In one example, multiple types of abnormal recognition rules can include at least one of the following:

[0081] Hardware failure rules: The temperature of the computing devices in the data center is too high, the memory occupancy rate is too large, etc.;

[0082] Network attack rules: The network devices in the data center have DDoS attack characteristics, abnormal traffic patterns, etc.;

[0083] Environmental anomaly rules: The temperature in the data center computer room is too high, the humidity exceeds the standard, etc.

[0084] In one example, based on the retrieved anomaly recognition rules, the data feature information, and the model guidance information, calling the anomaly monitoring model for anomaly recognition to obtain the anomaly recognition result may include: inputting the retrieved anomaly recognition rules, the data feature information, and the model guidance information into the anomaly monitoring model, so that the anomaly monitoring model, under the guidance of the model guidance information, uses the retrieved anomaly recognition rules for anomaly recognition according to the data feature information and outputs the anomaly recognition result.

[0085] In this embodiment, since the monitoring data can be used to reflect the general current situation of the monitoring data of the intelligent PDU, and different situations often correspond to different anomaly recognition rules, pre-matching the corresponding anomaly recognition rules for the current monitoring data can facilitate the subsequent anomaly monitoring model to directly select the corresponding anomaly recognition rules for anomaly recognition without having to analyze a suitable anomaly recognition scheme based on the input data, improving the efficiency of model operation and reducing the computational amount each time the model is called.

[0086] In an alternative embodiment, the monitoring data includes multiple sub-data segments, different sub-data segments have different acquisition times and / or different metric types, and the data feature information includes multiple data sub-features corresponding one-to-one to the multiple sub-data segments; exemplarily, each data sub-feature can be obtained by performing feature extraction on a corresponding sub-data segment, and this feature extraction can be implemented by vectorization processing / feature extraction model.

[0087] The retrieved anomaly recognition rules are suitable for indicating the recognition processing order of the anomaly monitoring model for the multiple data sub-features.

[0088] In this embodiment, since during the acquisition time of the monitoring data, more than one anomaly / fault may have occurred in the intelligent PDU and / or the data center, and different anomalies / faults may have a sequence (which can be characterized by the acquisition time) and / or a priority of importance level (which can be characterized by the metric type), this embodiment can specify the recognition processing order of the anomaly monitoring model for multiple data sub-features according to the above sequence and / or priority of importance level, so that the anomaly recognition processing process is more in line with the actual situation.

[0089] In an alternative embodiment, the data center includes multiple PDUs, the intelligent PDU is at least one of the multiple PDUs, and the fine-tuning method of the anomaly monitoring model includes:

[0090] Obtain sample data, where the sample data includes sample monitoring data and sample guidance information. The sample monitoring data includes the historical monitoring data of each of the multiple PDUs, and the sample guidance information is generated based on sample attribute feature information and sample anomaly feature information. The sample attribute feature information is extracted from the device attributes of each of the multiple PDUs in the data center, and the sample anomaly feature information is determined based on the historical anomaly data of each of the multiple PDUs.

[0091] According to the sample data, use reinforcement learning to fine-tune the large model to be fine-tuned.

[0092] In one example, the fine-tuning process may include the following steps:

[0093] Set the reward mechanism in the reinforcement learning algorithm (such as reinforcement learning, policy optimization, multi-agent system, imitation learning), where the reward mechanism includes positive rewards (positive rewards are given when the model correctly identifies and processes anomalies) and negative rewards (negative rewards are given when the model misreports or misses anomalies), and may also include long-term rewards (encouraging the model to maintain high accuracy and low false alarm rate in the long term);

[0094] Use the set reinforcement learning algorithm to fine-tune the large model to be fine-tuned according to the sample data, and give corresponding positive or negative rewards according to the output during the fine-tuning process, so as to complete the fine-tuning.

[0095] In an alternative implementation, before calling the anomaly monitoring model to perform anomaly identification based on the monitoring data and the model guidance information, the method further includes:

[0096] Obtain the inspection examples of the anomaly monitoring model;

[0097] Based on the inspection examples and the model guidance information, call the anomaly monitoring model to perform anomaly identification, and obtain the example identification results;

[0098] Adjust the model guidance information according to the example identification results, and use the adjusted guidance information as the new model guidance information.

[0099] In some cases, generally, the more accurate the guidance information provided to the large model is, the more accurate the output will be. However, the directly generated model guidance information may not necessarily ensure its sufficient accuracy. Therefore, in this embodiment, the pre-configured inspection examples are provided to the anomaly monitoring model together with the model guidance information for anomaly identification, and then based on the difference between the output example identification results and the standard identification results corresponding to the inspection examples, it is judged whether the model guidance information is sufficiently accurate, so as to adjust the model guidance information accordingly to make it more accurate, so as to improve the accuracy of subsequent anomaly identification.

[0100] In an alternative embodiment, calling the anomaly monitoring model for anomaly recognition based on the inspection example and the model guidance information to obtain an example recognition result, includes:

[0101] Based on the inspection example and the model guidance information, calling the anomaly monitoring model to perform multiple anomaly recognitions respectively to obtain multiple example recognition results; in one example, the multiple anomaly recognitions specifically refer to multiple independent anomaly recognitions, and each anomaly recognition uses the inspection example and the model guidance information as the recognition basis / model input, so as to obtain multiple independent example recognition results.

[0102] Adjusting the model guidance information according to the example recognition result, includes:

[0103] Performing feature extraction on multiple example recognition results respectively to obtain multiple example recognition features;

[0104] Clustering the multiple example recognition features to obtain a clustering result;

[0105] Adjusting the model guidance information at least according to the clustering result.

[0106] In this embodiment, multiple independent example recognition results obtained through multiple independent anomaly recognitions can more effectively judge whether the model guidance information is accurate enough.

[0107] In one example, the number of inspection examples can also be multiple. For each inspection example, the above multiple anomaly recognitions can be performed to obtain multiple independent example recognition results corresponding to each inspection example. Finally, clustering can be performed according to the example recognition features of the example recognition results corresponding to all inspection examples.

[0108] In an alternative embodiment, the adjusting the model guidance information at least according to the clustering result, includes:

[0109] Determining the example similarity between each of the multiple example recognition features and the clustering result;

[0110] Calculating the average similarity of the example similarities;

[0111] Adjusting the model guidance information according to the average similarity.

[0112] It should be noted that the similarity described in any one or more embodiments of the present application can be calculated by at least one of the following: cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc. It should be understood that the implementation manner of this similarity is only an example and does not limit the present application.

[0113] In an alternative implementation manner, after invoking the anomaly monitoring model to perform anomaly recognition based on the monitoring data and the model guidance information, the method further includes:

[0114] Obtain multiple infrared images taken of the intelligent PDU, where the shooting time of each of the multiple infrared images is included in the acquisition time period of the monitoring data;

[0115] Perform recognition on each frame of infrared image to obtain the corresponding intelligent PDU infrared information and environmental infrared information;

[0116] Generate image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images, where the image recognition information is suitable for indicating the state of the intelligent PDU;

[0117] Input the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly recognition result according to the image recognition information.

[0118] In one example, infrared images of the intelligent PDU can be obtained by an infrared camera.

[0119] In one example, generating image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images may include: splicing the intelligent PDU infrared information and environmental infrared information corresponding to the same frame of infrared image into infrared information, and splicing the infrared information corresponding to each of the multiple infrared images in the order of the shooting time to generate image recognition information.

[0120] In this embodiment, the image recognition information can reflect the state change of the intelligent PDU during the acquisition time period of the monitoring data, so that the anomaly monitoring model can update the anomaly recognition result on the premise of understanding the state change, and make the updated anomaly recognition result more in line with the actual state change.

[0121] In an alternative implementation manner, generating image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images includes:

[0122] Sort the multiple infrared images according to the shooting time;

[0123] For any two adjacent infrared images in the sorted multi-frame infrared images, identify the first difference information between the intelligent PDU infrared information corresponding to each of the two infrared images, and identify the second difference information between the ambient infrared information corresponding to each of the two infrared images;

[0124] Generate image recognition information based on the shooting time, the first difference information, and the second difference information.

[0125] In a second aspect, correspondingly, an embodiment of the present application further provides a data center monitoring system based on an intelligent PDU, which can implement all processes of the data center monitoring method based on an intelligent PDU provided in the above embodiment.

[0126] See Figure 2 , which shows a schematic structural diagram of a data center monitoring system based on an intelligent PDU provided in an embodiment of the present application. The data center monitoring system based on an intelligent PDU includes:

[0127] A monitoring module 201 for obtaining monitoring data of the intelligent PDU;

[0128] An attribute feature extraction module 202 for determining the device attributes of the intelligent PDU in the data center and extracting attribute feature information from the device attributes;

[0129] An abnormal feature extraction module 203 for determining abnormal feature information based on historical abnormal data of the intelligent PDU;

[0130] A guidance information generation module 204 for generating model guidance information based on the attribute feature information and the abnormal feature information;

[0131] An abnormal recognition module 205 for performing abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information to obtain an abnormal recognition result for the data center, so as to regulate the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning.

[0132] In an optional implementation manner, the performing abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information to obtain an abnormal recognition result for the data center includes:

[0133] Determine data feature information matching the monitoring data;

[0134] Retrieve corresponding abnormal recognition rules in a target knowledge base based on the data feature information, where the target knowledge base includes multiple types of abnormal recognition rules;

[0135] Based on the retrieved anomaly recognition rules, the data feature information, and the model guidance information, call the anomaly monitoring model for anomaly recognition to obtain the anomaly recognition result.

[0136] In an alternative embodiment, the monitoring data includes multiple sub-data segments, and different sub-data segments have different collection times and / or different metric types. The data feature information includes multiple data sub-features corresponding one-to-one to the multiple sub-data segments.

[0137] The retrieved anomaly recognition rules are adapted to indicate the recognition processing order of the anomaly monitoring model for the multiple data sub-features.

[0138] In an alternative embodiment, the data center includes multiple PDUs, and the intelligent PDU is at least one of the multiple PDUs. The fine-tuning method of the anomaly monitoring model includes:

[0139] Obtain sample data, where the sample data includes sample monitoring data and sample guidance information. The sample monitoring data includes the historical monitoring data of each of the multiple PDUs, and the sample guidance information is generated based on sample attribute feature information and sample anomaly feature information. The sample attribute feature information is extracted from the device attributes of each of the multiple PDUs in the data center, and the sample anomaly feature information is determined based on the historical anomaly data of each of the multiple PDUs.

[0140] According to the sample data, use reinforcement learning to fine-tune the large model to be fine-tuned.

[0141] In an alternative embodiment, the system further includes a guidance information adjustment module, and the guidance information adjustment module is used for:

[0142] Before calling the anomaly monitoring model for anomaly recognition based on the monitoring data and the model guidance information, obtain the inspection example of the anomaly monitoring model.

[0143] Based on the inspection example and the model guidance information, call the anomaly monitoring model for anomaly recognition to obtain the example recognition result.

[0144] Adjust the model guidance information according to the example recognition result, and use the adjusted guidance information as the new model guidance information.

[0145] In an alternative embodiment, the step of calling the anomaly monitoring model for anomaly recognition based on the inspection example and the model guidance information to obtain the example recognition result includes:

[0146] Based on the inspection examples and the model guidance information, call the anomaly monitoring model to perform multiple anomaly identifications respectively to obtain multiple example identification results;

[0147] The adjustment of the model guidance information according to the example identification results includes:

[0148] Extract features from the multiple example identification results respectively to obtain multiple example identification features;

[0149] Cluster the multiple example identification features to obtain a clustering result;

[0150] Adjust the model guidance information at least according to the clustering result.

[0151] In an alternative embodiment, the at least adjusting the model guidance information according to the clustering result includes:

[0152] Determine the example similarity between each of the multiple example identification features and the clustering result;

[0153] Calculate the average similarity of the example similarities;

[0154] Adjust the model guidance information according to the average similarity.

[0155] In an alternative embodiment, the system further includes an anomaly identification result update module, and the anomaly identification result update module is used for:

[0156] After calling the anomaly monitoring model to perform anomaly identification based on the monitoring data and the model guidance information, obtain multiple infrared images taken of the intelligent PDU, wherein the shooting time of each of the multiple infrared images is included in the acquisition time period of the monitoring data;

[0157] Identify each frame of infrared image to obtain corresponding intelligent PDU infrared information and environmental infrared information;

[0158] Generate image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images, wherein the image recognition information is suitable for indicating the state of the intelligent PDU;

[0159] Input the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly identification result according to the image recognition information.

[0160] In an alternative embodiment, the generating the image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images includes:

[0161] Sort the multiple infrared images according to the shooting time;

[0162] For any two adjacent infrared images among the sorted multiple infrared images, identify the first difference information between the intelligent PDU infrared information corresponding to the two infrared images respectively, and identify the second difference information between the environmental infrared information corresponding to the two infrared images respectively;

[0163] Generate image recognition information based on the shooting time, the first difference information and the second difference information.

[0164] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data center monitoring method based on an intelligent PDU described in any one of the above are implemented.

[0165] In a fourth aspect, an embodiment of the present application provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the data center monitoring method based on an intelligent PDU described in any one of the above are implemented.

[0166] In a fifth aspect, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the steps of the data center monitoring method based on an intelligent PDU described in any one of the above are implemented.

[0167] See Figure 3 , the computer device of this embodiment includes: a processor 301, a memory 302, and a computer program stored in the memory 302 and operable on the processor 301, such as a data center monitoring program based on an intelligent PDU. When the processor 301 executes the computer program, the steps in the above various embodiments of the data center monitoring method based on an intelligent PDU are implemented, such as Figure 1 the steps S101-S105 shown.

[0168] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 302 and executed by the processor 301 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device.

[0169] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that the schematic diagram is only an example of the computer device, and does not constitute a limitation on the computer device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, a bus, etc.

[0170] The processor 301 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 301 may also be any conventional processor, etc. The processor 301 is the control center of the computer device, and connects various parts of the entire computer device through various interfaces and lines.

[0171] The memory 302 can be used to store the computer program and / or module. The processor 301 realizes various functions of the computer device by running or executing the computer program and / or module stored in the memory 302, and by calling the data stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0172] Among them, if the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 301, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0173] In summary, the embodiments of this application have at least the following beneficial effects:

[0174] By adopting the embodiments of this application, by obtaining the monitoring data of the intelligent PDU; determining the device attributes of the intelligent PDU in the data center, and extracting attribute feature information from the device attributes; determining abnormal feature information based on the historical abnormal data of the intelligent PDU; generating model guiding information based on the attribute feature information and the abnormal feature information; and based on the monitoring data and the model guiding information, calling an abnormal monitoring model to perform abnormal identification to obtain an abnormal identification result for the data center, so as to be used for regulating the data center. Among them, the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning. Thus, corresponding guiding information can be generated according to the device attributes and historical abnormal data to guide the abnormal monitoring model obtained by fine-tuning a large model through reinforcement learning to efficiently and accurately analyze the monitoring data, and then timely and accurately monitor the abnormal conditions existing in the data center for corresponding regulation.

[0175] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary hardware platform. Of course, it can also be implemented entirely through hardware. Based on this understanding, all or part of the technical solution of this application that contributes to the background art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this application.

[0176] The above is the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of this application.

Claims

1. A data center monitoring method based on an intelligent PDU, characterized in that, Including: Obtaining the monitoring data of the intelligent PDU; Determining the device attributes of the intelligent PDU in the data center, and extracting attribute feature information from the device attributes; the device attributes include at least one of the following: power distribution and management attributes including socket attributes, modular socket attributes, and / or redundant power supply attributes, and security attributes including physical isolation attributes; Determining abnormal feature information based on the historical abnormal data of the intelligent PDU; Generating model guidance information based on the attribute feature information and the abnormal feature information; Based on the monitoring data and the model guidance information, calling an abnormal monitoring model to perform abnormal identification, obtaining an abnormal identification result for the data center, for regulating the data center, wherein the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning; Wherein, the calling the abnormal monitoring model to perform abnormal identification based on the monitoring data and the model guidance information, obtaining an abnormal identification result for the data center, includes: determining data feature information matching the monitoring data; based on the data feature information, retrieving corresponding abnormal identification rules in a target knowledge base, wherein the target knowledge base includes multiple types of abnormal identification rules, and the multiple types of abnormal identification rules include hardware failure rules and / or environmental abnormality rules; calling the abnormal monitoring model to perform abnormal identification based on the retrieved abnormal identification rules, the data feature information, and the model guidance information, obtaining the abnormal identification result; The monitoring data includes multiple sub-data, different sub-data having different collection times and / or different index types, the data feature information includes multiple data sub-features corresponding one-to-one to the multiple sub-data; the retrieved abnormal identification rules are adapted to indicate the identification processing order of the abnormal monitoring model for the multiple data sub-features, and the identification processing order is determined according to the importance priority characterized by the index type.

2. The method according to claim 1, wherein The data center includes multiple PDUs, the intelligent PDU is at least one of the multiple PDUs, and the fine-tuning method of the abnormal monitoring model includes: Obtaining sample data, wherein the sample data includes sample monitoring data and sample guidance information, the sample monitoring data includes the historical monitoring data of each of the multiple PDUs, the sample guidance information is generated based on sample attribute feature information and sample abnormal feature information, the sample attribute feature information is extracted from the device attributes of each of the multiple PDUs in the data center, and the sample abnormal feature information is determined based on the historical abnormal data of each of the multiple PDUs; Fine-tuning the large model to be fine-tuned using reinforcement learning according to the sample data.

3. The method according to any one of claims 1-2, characterized in that, Before the calling the abnormal monitoring model to perform abnormal identification based on the monitoring data and the model guidance information, the method further includes: Obtaining an inspection example of the abnormal monitoring model; Calling the abnormal monitoring model to perform abnormal identification based on the inspection example and the model guidance information, obtaining an example identification result; Adjust the model guidance information according to the example recognition result, and use the adjusted guidance information as the new model guidance information.

4. The method according to claim 3, wherein Based on the inspection example and the model guidance information, calling the anomaly monitoring model for anomaly recognition to obtain an example recognition result, including: Based on the inspection example and the model guidance information, calling the anomaly monitoring model to perform multiple anomaly recognitions respectively to obtain multiple example recognition results, wherein the multiple anomaly recognitions refer to multiple independent anomaly recognitions; Wherein, the number of inspection examples is multiple, and the adjusting the model guidance information according to the example recognition result includes: Performing feature extraction on multiple example recognition results respectively to obtain multiple example recognition features; Clustering the example recognition features corresponding to all inspection examples to obtain a clustering result; Adjusting the model guidance information at least according to the clustering result.

5. The method according to claim 4, wherein The adjusting the model guidance information at least according to the clustering result includes: Determining the example similarity between each of the multiple example recognition features and the clustering result; Calculating the average similarity of each example similarity; Adjusting the model guidance information according to the average similarity.

6. The method according to any one of claims 1-2, characterized in that, After calling the anomaly monitoring model for anomaly recognition based on the monitoring data and the model guidance information, the method further includes: Obtaining multiple infrared images taken of the intelligent PDU, wherein the shooting time of each of the multiple infrared images is included in the acquisition time period of the monitoring data; Recognizing each frame of infrared image to obtain corresponding intelligent PDU infrared information and environmental infrared information; Generating image recognition information based on the intelligent PDU infrared information, environmental infrared information and shooting time corresponding to each of the multiple infrared images, wherein the image recognition information is suitable for indicating the state of the intelligent PDU; Inputting the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly recognition result according to the image recognition information.

7. The method according to claim 6, wherein The generating image recognition information based on the intelligent PDU infrared information, environmental infrared information and shooting time corresponding to each of the multiple infrared images includes: Sorting the multiple infrared images according to the shooting time; For any two adjacent infrared images among the sorted multiple infrared images, recognizing the first difference information between the intelligent PDU infrared information corresponding to the two infrared images respectively, and recognizing the second difference information between the environmental infrared information corresponding to the two infrared images respectively; Generating image recognition information based on the shooting time, the first difference information and the second difference information.

8. A data center monitoring system based on an intelligent PDU, characterized in that, Including: A monitoring module for obtaining the monitoring data of the intelligent PDU; An attribute feature extraction module, configured to determine the device attributes of the intelligent PDU in the data center and extract attribute feature information from the device attributes; the device attributes include at least one of the following: power distribution and management attributes including socket attributes, modular socket attributes, and / or redundant power supply attributes, and security attributes including physical isolation attributes; An abnormal feature extraction module, configured to determine abnormal feature information based on the historical abnormal data of the intelligent PDU; A guidance information generation module, configured to generate model guidance information based on the attribute feature information and the abnormal feature information; An abnormal recognition module, configured to perform abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information, and obtain an abnormal recognition result for the data center, so as to regulate the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning; Among them, the step of performing abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information to obtain an abnormal recognition result for the data center includes: determining data feature information matching the monitoring data; based on the data feature information, retrieving corresponding abnormal recognition rules in a target knowledge base, where the target knowledge base includes multiple types of abnormal recognition rules, and the multiple types of abnormal recognition rules include hardware failure rules and / or environmental abnormality rules; invoking the abnormal monitoring model for abnormal recognition based on the retrieved abnormal recognition rules, the data feature information, and the model guidance information to obtain the abnormal recognition result; The monitoring data includes multiple sub-data segments, and different sub-data segments have different collection times and / or different index types. The data feature information includes multiple data sub-features corresponding one-to-one to the multiple sub-data segments. The retrieved abnormal recognition rules are adapted to indicate the recognition processing order of the abnormal monitoring model for the multiple data sub-features, and the recognition processing order is determined according to the priority of importance characterized by the index type.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

10. A computer device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Abnormality processing method and device of data center, electronic equipment and medium

    CN115827318A

  • Infrared thermal imaging abnormal scene monitoring method based on multi-modal large model and related device

    CN119723464A