Data center monitoring method and system based on intelligent PDU
Through the monitoring data and abnormal feature information of the intelligent PDU, model guidance information is generated, and a large model with fine-tuned reinforcement learning is called for abnormal identification, which solves the problem of data center abnormal monitoring in the existing technology, and realizes efficient and accurate abnormal monitoring and regulation.
Patent Information
- Application Number
- CN202510518503.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
When monitoring data centers, the prior art usually directly compares the index value with the threshold value, which is discovered only when the problem is serious, and abnormal situations cannot be monitored in a timely and accurate manner.
By obtaining the monitoring data of the intelligent PDU, determining the equipment attributes and abnormal characteristics information, generating model guidance information, and calling the large model with fine-tuning reinforcement learning to identify abnormalities, realizing timely and precise monitoring of abnormal situations in the data center.
By generating model guidance information, the abnormality monitoring model is guided to efficiently and accurately analyze the monitoring data, timely identify and regulate abnormal situations in the data center, and improve the accuracy and timeliness of monitoring.
Smart Images

Figure CN120046083A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of anomaly monitoring, and particularly to a data center monitoring method and system based on an intelligent PDU (Power Distribution Unit). Background Art
[0002] With the increase in the scale of data centers, it is necessary to perform accurate real-time monitoring on them. As a power management device in a data center, an intelligent PDU (Power Distribution Unit) can detect the power data and power consumption of each electrical device in the data center. Therefore, by analyzing the data of the intelligent PDU, the power consumption details of the entire data center can be obtained, so as to monitor in real time whether there are faults in the data center.
[0003] In related technologies, usually, each index value in the PDU monitoring data is directly compared with the corresponding index threshold value to monitor whether there are anomalies in the data center. However, when a certain index value no longer meets the requirements of the threshold value, the problems existing in the data center are often relatively serious. Summary of the Invention
[0004] To solve the above technical problems, embodiments of this application propose a data center monitoring method and system based on an intelligent PDU, which can timely and accurately monitor the anomalies existing in the data center for corresponding regulation.
[0005] In a first aspect, embodiments of this application provide a data center monitoring method based on an intelligent PDU, including: Obtain the monitoring data of the intelligent PDU; Determine the device attributes of the intelligent PDU in the data center, and extract attribute feature information from the device attributes; Based on the historical anomaly data of the intelligent PDU, determine anomaly feature information; Based on the attribute feature information and the anomaly feature information, generate model guiding information; Based on the monitoring data and the model guiding information, call an anomaly monitoring model to perform anomaly recognition, and obtain an anomaly recognition result for the data center for regulating the data center, where the anomaly monitoring model is obtained by fine-tuning a large model through reinforcement learning.
[0006] Optionally, the step of based on the monitoring data and the model guiding information, calling an anomaly monitoring model to perform anomaly recognition, and obtaining an anomaly recognition result for the data center includes: Determine data feature information matching the monitoring data; Retrieve the corresponding anomaly recognition rule in the target knowledge base based on the data feature information, where the target knowledge base includes multiple types of anomaly recognition rules; Based on the retrieved anomaly recognition rule, the data feature information, and the model guidance information, call the anomaly monitoring model to perform anomaly recognition to obtain the anomaly recognition result.
[0007] Optionally, the monitored data includes multiple sub-data segments, different sub-data segments have different collection times and / or different metric types, and the data feature information includes multiple data sub-features corresponding one by one to the multiple sub-data segments; The retrieved anomaly recognition rule is suitable for indicating the recognition processing order of the anomaly monitoring model for the multiple data sub-features.
[0008] Optionally, the data center includes multiple PDUs, the intelligent PDU is at least one of the multiple PDUs, and the fine-tuning method of the anomaly monitoring model includes: Obtain sample data, where the sample data includes sample monitored data and sample guidance information, the sample monitored data includes the historical monitored data of each of the multiple PDUs, the sample guidance information is generated based on sample attribute feature information and sample anomaly feature information, the sample attribute feature information is extracted from the device attributes of each of the multiple PDUs in the data center, and the sample anomaly feature information is determined based on the historical anomaly data of each of the multiple PDUs; According to the sample data, use reinforcement learning to fine-tune the large model to be fine-tuned.
[0009] Optionally, before calling the anomaly monitoring model to perform anomaly recognition based on the monitored data and the model guidance information, the method further includes: Obtain the inspection example of the anomaly monitoring model; Based on the inspection example and the model guidance information, call the anomaly monitoring model to perform anomaly recognition to obtain the example recognition result; Adjust the model guidance information according to the example recognition result, and use the adjusted guidance information as the new model guidance information.
[0010] Optionally, the step of obtaining the example recognition result by calling the anomaly monitoring model based on the inspection example and the model guidance information includes: Based on the inspection example and the model guidance information, call the anomaly monitoring model to perform multiple anomaly recognitions respectively to obtain multiple example recognition results; The step of adjusting the model guidance information according to the example recognition result includes: Feature extraction is performed on each of the multiple example recognition results to obtain multiple example recognition features; Cluster the multiple example recognition features to obtain a clustering result; Adjust the model guidance information at least according to the clustering result.
[0011] Optionally, the adjusting the model guidance information at least according to the clustering result includes: Determine the example similarity between each of the multiple example recognition features and the clustering result; Calculate the average similarity of the example similarities; Adjust the model guidance information according to the average similarity.
[0012] Optionally, after invoking the anomaly monitoring model to perform anomaly recognition based on the monitoring data and the model guidance information, the method further includes: Obtain multiple infrared images taken of the intelligent PDU, where the shooting time of each of the multiple infrared images is included in the acquisition time period of the monitoring data; Perform recognition on each frame of infrared image to obtain corresponding intelligent PDU infrared information and environmental infrared information; Generate image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images, where the image recognition information is suitable for indicating the state of the intelligent PDU; Input the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly recognition result according to the image recognition information.
[0013] Optionally, the generating the image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images includes: Sort the multiple infrared images according to the shooting time; For any two adjacent infrared images among the sorted multiple infrared images, identify the first difference information between the intelligent PDU infrared information corresponding to the two infrared images, and identify the second difference information between the environmental infrared information corresponding to the two infrared images; Generate image recognition information based on the shooting time, the first difference information, and the second difference information.
[0014] In a second aspect, an embodiment of the present application provides a data center monitoring system based on an intelligent PDU, including: A monitoring module for obtaining monitoring data of the intelligent PDU; An attribute feature extraction module, configured to determine the device attributes of the intelligent PDU in the data center, and extract attribute feature information from the device attributes; An abnormal feature extraction module, configured to determine abnormal feature information based on the historical abnormal data of the intelligent PDU; A guidance information generation module, configured to generate model guidance information based on the attribute feature information and the abnormal feature information; An abnormal recognition module, configured to perform abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information, so as to obtain an abnormal recognition result for the data center, for regulating the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning.
[0015] In summary, the embodiments of the present application at least have the following beneficial effects: By adopting the embodiments of the present application, by obtaining the monitoring data of the intelligent PDU; determining the device attributes of the intelligent PDU in the data center, and extracting attribute feature information from the device attributes; determining abnormal feature information based on the historical abnormal data of the intelligent PDU; generating model guidance information based on the attribute feature information and the abnormal feature information; performing abnormal recognition by invoking an abnormal monitoring model based on the monitoring data and the model guidance information, so as to obtain an abnormal recognition result for the data center, for regulating the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning, it is possible to generate corresponding guidance information according to the device attributes and historical abnormal data, so as to guide the abnormal monitoring model obtained by fine-tuning a large model through reinforcement learning to efficiently and accurately analyze the monitoring data, and then timely and accurately monitor the abnormal conditions existing in the data center for corresponding regulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flowchart of a data center monitoring method based on an intelligent PDU provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of a data center monitoring system based on an intelligent PDU provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts belong to the scope of protection of the present application.
[0018] In the description of the present application, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more. In the description of the present application, the term "comprising" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "according to" means "at least partially according to". The term "an embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments".
[0019] In the description of the present application, it should be noted that, unless otherwise clearly defined and limited, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0020] In the description of the present application, it should be noted that, unless otherwise defined, all the technical and scientific terms used in the present application have the same meanings as those commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0021] In a first aspect, referring to Figure 1 , a flowchart showing a data center monitoring method based on an intelligent PDU provided by an embodiment of the present application is shown. The method includes steps S101-S105, specifically as follows: S101, obtaining the monitoring data of the intelligent PDU.
[0022] In one example, a PDU typically has a sensor interface and / or a network interface, and the monitoring data may include sensor data collected via the sensor interface and / or network data collected via the network interface. Among them, the sensor data may include at least one of the following: current data, voltage data, power data, temperature data, humidity data.
[0023] S102. Determine the device attributes of the intelligent PDU in the data center, and extract attribute feature information from the device attributes.
[0024] In some cases, the device attributes may refer to the functions and / or characteristics that the intelligent PDU has / undertakes in the data center. The intelligent PDU in the data center is usually not just a simple power distribution device, and it can also have one or more advanced functions and characteristics in the data center, thereby improving the management efficiency, reliability, and security of the data center. Thus, exemplarily, the device attributes may include at least one of the following: power distribution and management attributes (such as socket attributes, modular socket attributes, redundant power supply attributes), remote management and control attributes, security attributes (such as the physical isolation attributes it undertakes).
[0025] In one example, the device attributes may include the power distribution attributes corresponding to the intelligent PDU in the data center, which are used to indicate which devices in the data center the intelligent PDU can provide power distribution functions for.
[0026] In one example, attribute feature information can be extracted from the device attributes by performing vectorization processing on the device attributes.
[0027] In one example, a pre-trained feature extraction model (such as a model containing an encoding network) can be used to extract features from the device attributes to obtain attribute feature information.
[0028] S103. Determine abnormal feature information based on the historical abnormal data of the intelligent PDU; In one example, abnormal feature information can be determined by performing vectorization processing on the historical abnormal data.
[0029] In one example, the above feature extraction model can be used to extract features from the historical abnormal data of the intelligent PDU to obtain abnormal feature information.
[0030] S104. Generate model guiding information based on the attribute feature information and the abnormal feature information; In one example, the attribute feature information and the abnormal feature information can be directly concatenated to obtain model guiding information.
[0031] S105. Based on the monitoring data and the model guidance information, call the anomaly monitoring model for anomaly recognition to obtain the anomaly recognition result for the data center, which is used to regulate the data center. The anomaly monitoring model is obtained by fine-tuning a large model through reinforcement learning. Exemplarily, reinforcement learning may include at least one of the following: imitation learning, multi-agent system, policy optimization, and reinforcement learning.
[0032] In one example, the large model can be a currently common large language model. Generally, the large model can directly perform automatic recognition based on the input. However, in the case of having reliable guidance information, its recognition is often more accurate. Thus, in this embodiment, the model guidance information (including information related to device attributes and historical anomaly data) can be used as the guidance information for the large model, so that the large model can more accurately perform anomaly recognition on the monitoring data.
[0033] In addition, it should be understood that a general large model often needs to be fine-tuned to better apply to a specific field. And using reinforcement learning (such as the reward mechanism therein) to fine-tune the large model to make the large model more suitable for the intelligent PDU anomaly monitoring field can be achieved by using common fine-tuning methods in this field, which is not strictly limited here.
[0034] In one example, the data center can be regulated accordingly based on the anomalies existing in the data center indicated by the anomaly recognition result. For example, if it is recognized that the data center has a too high load, the power of the data center can be reduced within a certain period of time in the future.
[0035] In an alternative implementation manner, the step of calling the anomaly monitoring model for anomaly recognition based on the monitoring data and the model guidance information to obtain the anomaly recognition result for the data center includes: Determine the data feature information that matches the monitoring data; Based on the data feature information, retrieve the corresponding anomaly recognition rule in the target knowledge base, where the target knowledge base includes multiple types of anomaly recognition rules; Based on the retrieved anomaly recognition rule, the data feature information, and the model guidance information, call the anomaly monitoring model for anomaly recognition to obtain the anomaly recognition result.
[0036] In one example, the data feature information can be obtained by vectorizing the monitoring data. Or, the data feature information can be obtained by extracting features from the monitoring data through the above-mentioned feature extraction model.
[0037] In one example, multiple types of anomaly recognition rules may include at least one of the following: Hardware failure rules: The temperature of the computing devices in the data center is too high, the memory occupancy rate is too large, etc.; Network attack rules: The network devices in the data center have DDoS attack characteristics, abnormal traffic patterns, etc.; Environmental anomaly rules: The temperature in the data center computer room is too high, the humidity exceeds the standard, etc.
[0038] In one example, based on the retrieved anomaly recognition rules, the data feature information, and the model guidance information, calling the anomaly monitoring model for anomaly recognition to obtain the anomaly recognition result may include: inputting the retrieved anomaly recognition rules, the data feature information, and the model guidance information into the anomaly monitoring model, so that the anomaly monitoring model, under the guidance of the model guidance information, uses the retrieved anomaly recognition rules for anomaly recognition according to the data feature information and outputs the anomaly recognition result.
[0039] In this embodiment, since the monitoring data can be used to reflect the general current situation of the monitoring data of the intelligent PDU, and different situations often correspond to different anomaly recognition rules, pre-matching the corresponding anomaly recognition rules for the current monitoring data can facilitate the subsequent anomaly monitoring model to directly select the corresponding anomaly recognition rules for anomaly recognition, without having to analyze a suitable anomaly recognition scheme based on the input data, improving the efficiency of model operation and reducing the computational amount each time the model is called.
[0040] In an alternative embodiment, the monitoring data includes multiple sub-data segments, and different sub-data segments have different collection times and / or different index types. The data feature information includes multiple data sub-features corresponding one-to-one to the multiple sub-data segments; exemplarily, each data sub-feature can be obtained by performing feature extraction on a corresponding sub-data segment, and this feature extraction can be implemented by vectorization processing / feature extraction model.
[0041] The retrieved anomaly recognition rules are suitable for indicating the recognition processing order of the anomaly monitoring model for the multiple data sub-features.
[0042] In this embodiment, since during the collection time of the monitoring data, more than one anomaly / failure may have occurred in the intelligent PDU and / or the data center, and different anomalies / failures may have a sequence (which can be characterized by the collection time) and / or a priority of importance level (which can be characterized by the index type), this embodiment can specify the recognition processing order of the anomaly monitoring model for multiple data sub-features according to the above sequence and / or priority of importance level, so that the anomaly recognition processing process is more in line with the actual situation.
[0043] In an alternative embodiment, the data center includes a plurality of PDUs, and the intelligent PDU is at least one of the plurality of PDUs. The fine-tuning method of the anomaly monitoring model includes: Obtain sample data, where the sample data includes sample monitoring data and sample guidance information. The sample monitoring data includes the historical monitoring data of each of the plurality of PDUs. The sample guidance information is generated based on sample attribute feature information and sample anomaly feature information. The sample attribute feature information is extracted from the device attributes of each of the plurality of PDUs in the data center, and the sample anomaly feature information is determined based on the historical anomaly data of each of the plurality of PDUs; According to the sample data, use reinforcement learning to fine-tune the large model to be fine-tuned.
[0044] In one example, the fine-tuning process may include the following steps: Set the reward mechanism in the reinforcement learning algorithm (such as reinforcement learning, policy optimization, multi-agent system, imitation learning), where the reward mechanism includes positive rewards (given positive rewards when the model correctly identifies and processes anomalies) and negative rewards (given negative rewards when the model gives false alarms or misses anomalies), and may also include long-term rewards (encouraging the model to maintain high accuracy and low false alarm rate in the long term); Use the set reinforcement learning algorithm to fine-tune the large model to be fine-tuned according to the sample data, and give corresponding positive or negative rewards according to the output during the fine-tuning process, so as to complete the fine-tuning.
[0045] In an alternative embodiment, before calling the anomaly monitoring model to perform anomaly identification based on the monitoring data and the model guidance information, the method further includes: Obtain the inspection example of the anomaly monitoring model; Based on the inspection example and the model guidance information, call the anomaly monitoring model to perform anomaly identification to obtain an example identification result; Adjust the model guidance information according to the example identification result, and use the adjusted guidance information as the new model guidance information.
[0046] In some cases, generally, the more accurate the guidance information provided to the large model, the more accurate the output will be. However, the directly generated model guidance information may not be guaranteed to be accurate enough. Therefore, in this embodiment, the pre-configured inspection example and the model guidance information are provided to the anomaly monitoring model for anomaly identification together, and then based on the difference between the output example identification result and the standard identification result corresponding to the inspection example, it is judged whether the model guidance information is accurate enough, so as to adjust the model guidance information accordingly to make it more accurate, so as to improve the accuracy of subsequent anomaly identification.
[0047] In an alternative embodiment, calling the anomaly monitoring model for anomaly recognition based on the inspection example and the model guidance information to obtain an example recognition result, includes: Based on the inspection example and the model guidance information, calling the anomaly monitoring model to perform multiple anomaly recognitions respectively to obtain multiple example recognition results; in one example, the multiple anomaly recognitions specifically refer to multiple independent anomaly recognitions, and each anomaly recognition uses the inspection example and the model guidance information as the recognition basis / model input, so as to obtain multiple independent example recognition results.
[0048] Adjusting the model guidance information according to the example recognition result, includes: Performing feature extraction on multiple example recognition results respectively to obtain multiple example recognition features; Clustering the multiple example recognition features to obtain a clustering result; Adjusting the model guidance information at least according to the clustering result.
[0049] In this embodiment, multiple independent example recognition results obtained through multiple independent anomaly recognitions can more effectively judge whether the model guidance information is accurate enough.
[0050] In one example, the number of inspection examples can also be multiple. For each inspection example, the above-mentioned multiple anomaly recognitions can be performed to obtain multiple independent example recognition results corresponding to each inspection example. Finally, clustering can be performed according to the example recognition features of the example recognition results corresponding to all inspection examples.
[0051] In an alternative embodiment, the adjusting the model guidance information at least according to the clustering result, includes: Determining the example similarity between each of the multiple example recognition features and the clustering result; Calculating the average similarity of the example similarities; Adjusting the model guidance information according to the average similarity.
[0052] It should be noted that the similarity described in any one or more embodiments of the present application can be calculated by at least one of the following: cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc. It should be understood that the implementation manner of this similarity is only an example and does not constitute a limitation to the present application.
[0053] In an alternative embodiment, after calling the anomaly monitoring model for anomaly recognition based on the monitoring data and the model guidance information, the method further includes: Obtain multiple infrared images captured for the intelligent PDU, where the capture time of each of the multiple infrared images is included in the acquisition time period of the monitoring data; Identify each frame of the infrared image to obtain the corresponding intelligent PDU infrared information and environmental infrared information; Generate image recognition information based on the intelligent PDU infrared information, environmental infrared information, and capture time corresponding to each of the multiple infrared images, where the image recognition information is suitable for indicating the state of the intelligent PDU; Input the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly recognition result according to the image recognition information.
[0054] In one example, an infrared camera can be used to capture infrared images of the intelligent PDU.
[0055] In one example, generating image recognition information based on the intelligent PDU infrared information, environmental infrared information, and capture time corresponding to each of the multiple infrared images may include: splicing the intelligent PDU infrared information and environmental infrared information corresponding to the same frame of infrared image into infrared information, and splicing the infrared information corresponding to each of the multiple infrared images in the order of the capture time to generate image recognition information.
[0056] In this embodiment, the image recognition information can reflect the state change of the intelligent PDU during the acquisition time period of the monitoring data, so that the anomaly monitoring model can update the anomaly recognition result on the premise of understanding the state change, and make the updated anomaly recognition result more in line with the actual state change.
[0057] In an alternative implementation, generating image recognition information based on the intelligent PDU infrared information, environmental infrared information, and capture time corresponding to each of the multiple infrared images includes: Sort the multiple infrared images according to the capture time; For any two adjacent infrared images among the sorted multiple infrared images, identify the first difference information between the intelligent PDU infrared information corresponding to the two infrared images, and identify the second difference information between the environmental infrared information corresponding to the two infrared images; Generate image recognition information based on the capture time, the first difference information, and the second difference information.
[0058] Second aspect, correspondingly, the embodiments of the present application further provide a data center monitoring system based on an intelligent PDU, which can implement all processes of the data center monitoring method based on an intelligent PDU provided in the above embodiments.
[0059] See Figure 2 , which shows a schematic structural diagram of the data center monitoring system based on an intelligent PDU provided in the embodiments of the present application. The data center monitoring system based on an intelligent PDU includes: A monitoring module 201, configured to obtain monitoring data of the intelligent PDU; An attribute feature extraction module 202, configured to determine the device attributes of the intelligent PDU in the data center and extract attribute feature information from the device attributes; An abnormal feature extraction module 203, configured to determine abnormal feature information based on historical abnormal data of the intelligent PDU; A guidance information generation module 204, configured to generate model guidance information based on the attribute feature information and the abnormal feature information; An abnormal identification module 205, configured to perform abnormal identification by invoking an abnormal monitoring model based on the monitoring data and the model guidance information to obtain an abnormal identification result for the data center, so as to be used to regulate the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning.
[0060] In an optional implementation manner, the performing abnormal identification by invoking an abnormal monitoring model based on the monitoring data and the model guidance information to obtain an abnormal identification result for the data center includes: Determining data feature information matching the monitoring data; Based on the data feature information, retrieving corresponding abnormal identification rules in a target knowledge base, where the target knowledge base includes multiple types of abnormal identification rules; Based on the retrieved abnormal identification rules, the data feature information, and the model guidance information, invoking the abnormal monitoring model to perform abnormal identification to obtain the abnormal identification result.
[0061] In an optional implementation manner, the monitoring data includes multiple sub-data segments, and different sub-data segments have different collection times and / or different index types. The data feature information includes multiple data sub-features corresponding one by one to the multiple sub-data segments; The retrieved abnormal identification rules are adapted to indicate the identification processing order of the abnormal monitoring model for the multiple data sub-features.
[0062] In an alternative embodiment, the data center includes a plurality of PDUs, and the intelligent PDU is at least one of the plurality of PDUs. The fine-tuning method of the anomaly monitoring model includes: Obtain sample data, where the sample data includes sample monitoring data and sample guidance information. The sample monitoring data includes the historical monitoring data of each of the plurality of PDUs. The sample guidance information is generated based on sample attribute feature information and sample anomaly feature information. The sample attribute feature information is extracted from the device attributes of each of the plurality of PDUs in the data center, and the sample anomaly feature information is determined based on the historical anomaly data of each of the plurality of PDUs; According to the sample data, use reinforcement learning to fine-tune the large model to be fine-tuned.
[0063] In an alternative embodiment, the system further includes a guidance information adjustment module, and the guidance information adjustment module is used for: Before calling the anomaly monitoring model to perform anomaly recognition based on the monitoring data and the model guidance information, obtain the inspection examples of the anomaly monitoring model; Based on the inspection examples and the model guidance information, call the anomaly monitoring model to perform anomaly recognition to obtain example recognition results; Adjust the model guidance information according to the example recognition results, and use the adjusted guidance information as the new model guidance information.
[0064] In an alternative embodiment, the step of calling the anomaly monitoring model to perform anomaly recognition based on the inspection examples and the model guidance information to obtain example recognition results includes: Based on the inspection examples and the model guidance information, call the anomaly monitoring model to perform anomaly recognition multiple times to obtain a plurality of the example recognition results; The step of adjusting the model guidance information according to the example recognition results includes: Extract features from each of the plurality of example recognition results to obtain a plurality of example recognition features; Cluster the plurality of example recognition features to obtain a clustering result; Adjust the model guidance information at least according to the clustering result.
[0065] In an alternative embodiment, the step of adjusting the model guidance information at least according to the clustering result includes: Determine the example similarity between each of the plurality of example recognition features and the clustering result; Calculate the average similarity of the example similarities; Adjust the model guidance information according to the average similarity.
[0066] In an alternative embodiment, the system further includes an anomaly recognition result update module, and the anomaly recognition result update module is configured to: After calling an anomaly monitoring model to perform anomaly recognition based on the monitoring data and the model guidance information, obtain multiple infrared images taken of the intelligent PDU, where the shooting time of each of the multiple infrared images is included in the acquisition time period of the monitoring data; Identify corresponding intelligent PDU infrared information and environmental infrared information for each frame of infrared image; Generate image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images, where the image recognition information is suitable for indicating the state of the intelligent PDU; Input the image recognition information into the anomaly monitoring model, so that the anomaly monitoring model updates the anomaly recognition result according to the image recognition information.
[0067] In an alternative embodiment, the generating image recognition information based on the intelligent PDU infrared information, environmental infrared information, and shooting time corresponding to each of the multiple infrared images includes: Sort the multiple infrared images according to the shooting time; For any two adjacent infrared images among the sorted multiple infrared images, identify first difference information between the intelligent PDU infrared information corresponding to the two infrared images, and identify second difference information between the environmental infrared information corresponding to the two infrared images; Generate image recognition information based on the shooting time, the first difference information, and the second difference information.
[0068] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data center monitoring method based on an intelligent PDU described in any one of the above are implemented.
[0069] In a fourth aspect, an embodiment of the present application provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the data center monitoring method based on an intelligent PDU described in any one of the above are implemented.
[0070] In a fifth aspect, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the data center monitoring method based on an intelligent PDU described in any one of the above are implemented.
[0071] Referring to Figure 3 , the computer device of this embodiment includes: a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301, such as a data center monitoring program based on an intelligent PDU. When the processor 301 executes the computer program, the steps in each of the above-described embodiments of the data center monitoring method based on an intelligent PDU are implemented, such as Figure 1 the steps S101 - S105 shown.
[0072] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 302 and executed by the processor 301 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device.
[0073] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that the schematic diagram is only an example of the computer device, and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.
[0074] The processor 301 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor 301 may also be any conventional processor, etc. The processor 301 is the control center of the computer device, and connects various parts of the entire computer device through various interfaces and lines.
[0075] The memory 302 can be used to store the computer programs and / or modules. The processor 301 realizes various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 302, and by calling the data stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0076] Among them, if the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 301, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0077] In summary, the embodiments of the present application have at least the following beneficial effects: By adopting the embodiments of the present application, by obtaining the monitoring data of the intelligent PDU; determining the device attributes of the intelligent PDU in the data center, and extracting attribute feature information from the device attributes; determining abnormal feature information based on the historical abnormal data of the intelligent PDU; generating model guidance information based on the attribute feature information and the abnormal feature information; and based on the monitoring data and the model guidance information, calling an abnormal monitoring model to perform abnormal recognition to obtain an abnormal recognition result for the data center, so as to be used to regulate the data center, where the abnormal monitoring model is obtained by fine-tuning a large model through reinforcement learning, it is thus possible to generate corresponding guidance information according to the device attributes and historical abnormal data to guide the abnormal monitoring model obtained by fine-tuning the large model through reinforcement learning to efficiently and accurately analyze the monitoring data, and then timely and accurately monitor the abnormalities existing in the data center for corresponding regulation.
[0078] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary hardware platform, and of course, it can also be implemented entirely by hardware. Based on such an understanding, all or part of the technical solution of the present application that contributes to the background art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0079] The above is the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. A data center monitoring method based on intelligent PDU, characterized in that: include: Obtain monitoring data of the intelligent PDU; Determine the device attributes of the intelligent PDU in the data center, and extract attribute feature information from the device attributes; Determine abnormal characteristic information based on historical abnormal data of the intelligent PDU; Generate model guidance information based on the attribute feature information and the abnormal feature information; Based on the monitoring data and the model guidance information, the abnormal monitoring model is called to perform abnormality identification to obtain an abnormality identification result for the data center for use in regulating the data center, wherein the abnormal monitoring model is a large model obtained by fine-tuning through reinforcement learning.
2. The method according to claim 1, characterized in that: The calling of the abnormal monitoring model to perform abnormality identification based on the monitoring data and the model guidance information to obtain an abnormality identification result for the data center includes: Determining data feature information that matches the monitoring data; Based on the data feature information, a corresponding anomaly identification rule is retrieved in a target knowledge base, wherein the target knowledge base includes multiple types of anomaly identification rules; Based on the retrieved anomaly identification rule, the data feature information and the model guidance information, the anomaly monitoring model is called to perform anomaly identification to obtain the anomaly identification result.
3. The method according to claim 2, characterized in that The monitoring data includes multiple segments of sub-data, different sub-data have different collection times and / or different indicator types, and the data feature information includes multiple data sub-features corresponding to the multiple segments of sub-data one by one; The retrieved anomaly identification rules are adapted to indicate an identification processing order of the anomaly monitoring model for the plurality of data sub-features.
4. The method according to claim 1, characterized in that: The data center includes a plurality of PDUs, the intelligent PDU is at least one of the plurality of PDUs, and the fine-tuning method of the abnormal monitoring model includes: Acquire sample data, wherein the sample data includes sample monitoring data and sample guidance information, the sample monitoring data includes historical monitoring data of each of the multiple PDUs, the sample guidance information is generated based on sample attribute feature information and sample abnormality feature information, the sample attribute feature information is extracted from device attributes of each of the multiple PDUs in the data center, and the sample abnormality feature information is determined based on historical abnormality data of each of the multiple PDUs; Based on the sample data, the large model to be fine-tuned is fine-tuned using reinforcement learning.
5. The method according to any one of claims 1 to 4, characterized in that: Before calling the abnormal monitoring model to perform abnormality identification based on the monitoring data and the model guidance information, the method further includes: Obtaining a check example of the abnormal monitoring model; Based on the inspection example and the model guidance information, calling the abnormality monitoring model to perform abnormality identification and obtain an example identification result; The model guidance information is adjusted according to the example recognition result, and the adjusted guidance information is used as new model guidance information.
6. The method according to claim 5, characterized in that The method of calling the abnormality monitoring model to perform abnormality identification based on the inspection example and the model guidance information to obtain an example identification result includes: Based on the inspection example and the model guidance information, calling the abnormal monitoring model to perform multiple abnormality identifications respectively to obtain multiple example identification results; The adjusting the model guidance information according to the example recognition result includes: Extracting features from the plurality of example recognition results respectively to obtain a plurality of example recognition features; Clustering the multiple example recognition features to obtain a clustering result; The model guidance information is adjusted at least according to the clustering result.
7. The method according to claim 6, characterized in that The adjusting the model guidance information at least according to the clustering result includes: Determining an example similarity between each of the plurality of example identification features and the clustering result; Calculate the average similarity of each of the example similarities; The model guidance information is adjusted according to the average similarity.
8. The method according to any one of claims 1 to 4, characterized in that: After calling the abnormal monitoring model to perform abnormality identification based on the monitoring data and the model guidance information, the method further includes: Acquire multiple frames of infrared images taken for the smart PDU, wherein the shooting time of each of the multiple frames of infrared images is included in the collection time period of the monitoring data; Identify each frame of infrared image to obtain the corresponding intelligent PDU infrared information and environmental infrared information; Generate image recognition information based on the intelligent PDU infrared information, environmental infrared information and shooting time corresponding to each of the multiple frames of infrared images, wherein the image recognition information is suitable for indicating the state of the intelligent PDU; The image recognition information is input into the abnormality monitoring model, so that the abnormality monitoring model updates the abnormality recognition result according to the image recognition information.
9. The method according to claim 8, characterized in that The generating of image recognition information based on the intelligent PDU infrared information, the environment infrared information and the shooting time corresponding to each of the multiple frames of infrared images includes: sorting the multiple frames of infrared images according to the shooting time; For any two adjacent infrared image frames in the sorted multiple infrared image frames, identify first difference information between the intelligent PDU infrared information corresponding to the arbitrary two infrared image frames, and identify second difference information between the environmental infrared information corresponding to the arbitrary two infrared image frames; Image recognition information is generated based on the photographing time, the first difference information, and the second difference information.
10. A data center monitoring system based on intelligent PDU, characterized in that: include: A monitoring module, used to obtain monitoring data of the intelligent PDU; An attribute feature extraction module, used to determine the device attributes of the intelligent PDU in the data center and extract attribute feature information from the device attributes; An abnormal feature extraction module, used to determine abnormal feature information based on historical abnormal data of the intelligent PDU; A guidance information generation module, used to generate model guidance information based on the attribute feature information and the abnormal feature information; An anomaly identification module is used to call the anomaly monitoring model to perform anomaly identification based on the monitoring data and the model guidance information, and obtain an anomaly identification result for the data center for use in regulating the data center, wherein the anomaly monitoring model is a large model obtained by fine-tuning through reinforcement learning.
Citation Information
Patent Citations
Data center power environment monitoring system and method
CN110046074A
Power production abnormity monitoring method and device, computer equipment and storage medium
CN112669316A
Abnormality processing method and device of data center, electronic equipment and medium
CN115827318A
Intelligent power monitoring method and system
CN116882804A
System equipment state monitoring method and related device
CN117994618A